Skip to main content
Glama

generate_speech

Convert text into speech audio on your machine, returning a download link or a job id to poll. Handles up to 3000 characters per call.

Instructions

Turn text into speech on this machine with Guaardvark's Audio Foundry. Waits up to about a minute for the file (usually seconds) and returns its name, library document id, length and a download link. A longer read, such as a script near the 3000-character limit or a cold start, answers with a job_id instead: poll get_generation_status with it, and do not call again, since the file lands in the library either way. Up to 3000 characters per call; split longer scripts. Voices: naming a voice (a Kokoro id such as 'af_heart' or 'bm_george') always speaks with Kokoro; with no voice, engine 'auto' uses Chatterbox's single stock voice when Chatterbox is installed and Kokoro's default voice otherwise. Chatterbox has no voice ids, so engine 'chatterbox' with a voice is refused. A model or voice pack that is not installed is refused with a pointer to Audio Studio → Manage models; nothing is downloaded. Needs the Audio Foundry plugin running. Cloning a real person's voice (a Chatterbox reference clip) is not available here; it needs the person to confirm in Audio Studio that they have the right to clone that voice. For a song use generate_music.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
textYesWhat to say, as it should be spoken (spell out numbers and names the way they sound).
voiceNoA Kokoro voice id: accent and gender prefix plus a name, e.g. 'af_heart' (American female, the default), 'am_michael', 'bf_emma', 'bm_george', 'ef_dora' (Spanish). Naming one selects Kokoro. An unknown id is refused with the full list. Omit for the engine's default voice.
engineNo'auto' (default): Kokoro when voice is set; otherwise Chatterbox when it is installed, falling back to Kokoro. 'kokoro': the named voice or af_heart. 'chatterbox': its one stock voice; do not combine with voice.auto
idempotency_keyNoOptional. Send the same key when retrying this call: the retry waits for or returns the first run instead of starting a second one.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Added

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=false and destructiveHint=false; the description goes well beyond, disclosing the ~1 minute wait, the returned fields, the job_id long-read fallback, and the refusal cases (uninstalled model/voice pack, chatterbox+voice combination) with a pointer to Manage models. This is exactly the mutating-tool context the annotations cannot carry.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but front-loaded: purpose and return behavior come first, then constraints and refusal cases. It is long, but almost every sentence carries an actionable constraint rather than filler, so it stays on the right side of verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by describing the return shape (file name, library document id, length, download link) and the job_id alternative. For a complex multi-engine, mutating tool this is complete enough to call correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3 and the schema already documents all four parameters. The description adds value on top by explaining cross-parameter interaction and refusal behavior (e.g. 'engine chatterbox with a voice is refused', voice selection forcing Kokoro), which goes slightly beyond the field-level descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Turn text into speech on this machine') and names the engine/plugin backing it. It also explicitly distinguishes itself from the sibling generate_music ('For a song use generate_music'), so an agent can route without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete when-to-use and when-not guidance: split scripts over 3000 chars, poll get_generation_status on a job_id and 'do not call again', and use generate_music for songs. It also states the prerequisite (Audio Foundry plugin running) and the cloning restriction, leaving little to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.