transcribe
Convert spoken audio into written text from a file path or base64 input, with optional language hint. Returns recognized text, detected language, and duration.
Instructions
Transcribe spoken audio to text.
Pass exactly one of audio_base64 or audio_path.
Args:
audio_base64: Base64-encoded audio bytes (wav/mp3/webm/m4a).
audio_path: Path to an audio file under OMNIVOICE_MCP_BASE_PATH
(relative to it, or absolute inside it). The base path is the
security boundary: with none configured, paths are refused.
Prefer this lane for LLM agents - the audio never enters the
agent's context.
language: Optional language hint; omit for auto-detect.
Returns:
JSON with the recognized text, language, and duration.
For long operations use voicestudio_start_job.Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| language | No | ||
| audio_path | No | ||
| audio_base64 | No |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| result | Yes |