Transcribe audio
transcribe_audioConvert audio to transcript text for Poly-Glot workflows. Use for supplied speech recordings, not for translating text. Provide exactly one source: a publicly reachable HTTPS audioUrl or base64-encoded audioBase64 (without a data-URL prefix). For base64 audio, supply mimeType and filename when known; languageHint can improve recognition. Set detectLanguage=true to identify the resulting text's language. Requires an active trial or Pro. Returns JSON containing transcript text, optional detectedLanguage, provider, and model. Audio is submitted to a transcription provider; do not send material without permission.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | No | Optional recognition context, such as expected names or specialist vocabulary; not the text to transcribe. | |
| audioUrl | No | Public HTTPS URL of an audio recording to transcribe. Use this OR audioBase64, never both; private or local URLs are not supported. | |
| filename | No | Original audio filename with extension, such as meeting.m4a, to help determine the format. | |
| mimeType | No | Audio MIME type, such as audio/mpeg, audio/mp4, or audio/wav. Recommended for base64 input. | |
| audioBase64 | No | Base64-encoded audio bytes, without a data: prefix. Alternative to audioUrl; provide exactly one audio source. | |
| languageHint | No | Optional language hint for speech recognition, such as en, es, or fr. Leave blank if unknown. | |
| detectLanguage | No | When true, also infer a supported Poly-Glot language from the resulting transcript; defaults to false. |