transcribe_audio
Transcribe audio to text with WORD-LEVEL timestamps (timestamps:'word' returns per-word start/end times — subtitle alignment, karaoke captions, cutting video to speech) or segment timestamps. Uses Mistral Transcription — high-accuracy speech recognition that handles accents, background noise, and overlapping speakers. 13 languages: en, zh, hi, es, ar, fr, pt, ru, de, ja, ko, it, nl. Up to 500 MB / 60 minutes per file. Async — returns requestId, poll with check_job_status(jobType='transcription'), then get_job_result. 10 sats/min. Privacy: audio and transcripts are ephemeral — processed, returned, and discarded. Never persisted. Pay per request with Bitcoin Lightning — no API key or signup needed. Requires create_payment with toolName='transcribe_audio'.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| diarize | No | Identify different speakers (default false). Forces segment granularity upstream. If the speaker-label request fails, you still get the text, and result.degraded includes 'diarization'. | |
| language | No | Language code (e.g., 'en', 'es') | |
| paymentId | Yes | Valid payment ID (must be paid) | |
| timestamps | No | Timestamp granularity in result.segments. 'word' returns per-word start/end times (subtitle alignment, karaoke captions, cutting video to speech). Default 'segment'. If the timestamp request fails, you still get the text, and result.degraded includes 'timestamps' (no segments, srt or vtt). | |
| audioBase64 | Yes | Base64 encoded audio file | |
| callback_id | No | Optional correlation string echoed back in the webhook body. Max 128 chars. | |
| callback_url | No | Optional HTTPS webhook we POST when the job finishes (HMAC-signed). Polling still works. |