generate_speech
Convert text into speech audio using saved voice profiles, with control over language, style, speed, and output format.
Instructions
Generate speech audio from text.
Args:
text: The text to synthesize into speech.
language: Target language (ISO code or 'Auto'). 646 languages
supported. Omit to use the voice profile's saved language;
an explicit 'Auto' overrides it.
profile_id: ID of a saved voice profile to clone. Omit to use this
agent's bound voice (Settings → MCP), else the default voice.
instruct: Style instruction (e.g. 'whisper', 'excited', 'narrator').
speed: Speech speed multiplier (0.5–2.0, default 1.0).
steps: Diffusion steps (8=fast/draft, 16=balanced, 32=quality).
format: File and URL format: wav (default), ogg or opus. Both
ogg and opus carry Opus in Ogg; requires files/both mode and ffmpeg.
Returns:
JSON with audio_id, generation_time_s, audio_duration_s and the
audio shaped by OMNIVOICE_MCP_OUTPUT_MODE: base64 WAV data
('resources', the default), a URL plus an optional file ('files'),
or both ('both'). Prefer 'files' for LLM agents.
For long operations use voicestudio_start_job.Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| speed | No | ||
| steps | No | ||
| format | No | wav | |
| instruct | No | ||
| language | No | ||
| profile_id | No |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| result | Yes |