Skip to main content
Glama

Change Voice

change_voice

Re-voice existing speech into one of the user's cloned voices (Voice Changer, speech-to-speech): keeps the timing and delivery of the source, swaps the voice. The source can be audio or video, up to 5 minutes. Costs about 10 credits per minute. voiceId must be a cloned voice id from list_voices (create one with create_voice_clone). Poll with wait_for_voiceover. Renders a live audio player in app-capable hosts.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
voiceIdYesCloned voice id (list_voices → clones).
sourceUrlYesPublic URL of the speech audio or video to convert.
outputFormatNoDefault mp3.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does so well: it discloses cost (~10 credits/minute), a hard 5-minute source limit, accepted input types, the async/polling nature of the operation, and the output behavior ('renders a live audio player'). These are exactly the traits an agent needs before invoking and typically missing from structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense paragraph, front-loaded with the core action and transformation, then constraints, cost, id sourcing, polling, and output. Every sentence adds a distinct, non-redundant fact with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-param async mutation tool with no annotations and no output schema, the description covers purpose, input constraints, cost, id source, polling path, and output rendering. An agent has everything needed to call it correctly and know what to expect.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters including the cloned-voice constraint on voiceId and video acceptance for sourceUrl. The description largely reinforces the schema (cloned voice id, audio-or-video source) rather than adding new syntax or format detail, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Re-voice existing speech into one of the user's cloned voices', with 'Voice Changer, speech-to-speech' as a gloss) and clarifies the transformation ('keeps the timing and delivery of the source, swaps the voice'). This clearly separates it from sibling create_voiceover (text-to-speech) and create_voice_clone (voice creation), which an agent can distinguish without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear usage context: source must be audio or video up to 5 minutes, voiceId must come from list_voices, and results should be polled with wait_for_voiceover. It routes the agent to the right supporting tools, though it never states an explicit 'do not use this for X' exclusion (e.g., versus create_voiceover for text input).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources