Skip to main content
Glama
AIM-IT4
by AIM-IT4

generate_speech

Convert text into speech audio using saved voice profiles, with control over language, style, speed, and output format.

Instructions

Generate speech audio from text.

    Args:
        text: The text to synthesize into speech.
        language: Target language (ISO code or 'Auto'). 646 languages
            supported. Omit to use the voice profile's saved language;
            an explicit 'Auto' overrides it.
        profile_id: ID of a saved voice profile to clone. Omit to use this
            agent's bound voice (Settings → MCP), else the default voice.
        instruct: Style instruction (e.g. 'whisper', 'excited', 'narrator').
        speed: Speech speed multiplier (0.5–2.0, default 1.0).
        steps: Diffusion steps (8=fast/draft, 16=balanced, 32=quality).
        format: File and URL format: wav (default), ogg or opus. Both
            ogg and opus carry Opus in Ogg; requires files/both mode and ffmpeg.

    Returns:
        JSON with audio_id, generation_time_s, audio_duration_s and the
        audio shaped by OMNIVOICE_MCP_OUTPUT_MODE: base64 WAV data
        ('resources', the default), a URL plus an optional file ('files'),
        or both ('both'). Prefer 'files' for LLM agents.
     For long operations use voicestudio_start_job.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
textYes
speedNo
stepsNo
formatNowav
instructNo
languageNo
profile_idNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the output-mode behavior (resources/files/both via OMNIVOICE_MCP_OUTPUT_MODE), the dependency that ogg/opus 'requires files/both mode and ffmpeg', and the fallback chain for voice selection. It does not mention auth requirements, rate limits, or cost, which is a minor gap for a generation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action, then a tidy Args list where each line earns its place, followed by Returns and the sibling pointer. Slightly verbose and the dangling 'For long operations' sentence sits awkwardly after the Returns block, but nothing is padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a seven-parameter generation tool with output modes, the description covers every decision point: parameter meanings, defaults, format constraints with their prerequisite (ffmpeg, files/both mode), output-mode selection, and long-operation escalation. An output schema exists, so the Returns detail is bonus rather than a necessity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the prose must compensate, and it does: all seven parameters get semantics beyond their names — language as ISO code or 'Auto' with 646 languages, steps as 8=draft/16=balanced/32=quality, speed range 0.5–2.0, and valid format values. This is far richer than the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Generate speech audio from text'), making it unmistakable against the reverse-direction sibling transcribe and the voice-creation siblings (clone_voice, design_voice). No ambiguity about what the tool produces.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit alternative routing: 'For long operations use voicestudio_start_job' gives the when-not condition, and 'Prefer files for LLM agents' guides mode selection. Default-resolution rules for language and profile_id (omit to inherit binding, explicit 'Auto' overrides) tell the agent exactly when to pass or omit each optional argument.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.