Speech-to-text (Whisper large v3 turbo)
ai_transcribeTranscribes an audio file with Whisper large v3 turbo: pass an https URL (mp3, wav, flac, m4a, ogg, up to 10 MB) or base64 audio in a POST. Returns the text, the detected language with its probability, the duration, and timed segments; add words=true for word-level timestamps. Around 100 languages. You are only charged if the transcript is delivered. No API key, no account. $0.02 per call, paid over x402 (USDC).
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | https URL of the audio file, up to 10 MB. | |
| audio | No | Alternative to url: the audio as base64 (POST only). | |
| words | No | Include word-level timestamps. Default false. | |
| language | No | Optional ISO code to skip detection, e.g. es. |