Skip to main content
Glama
AIWerk

@aiwerk/mcp-server-elevenlabs

by AIWerk

text_to_voice_design

Generate a synthetic voice from a text description with ElevenLabs. Configure model, quality, and format to design and preview custom voices.

Instructions

Design A Voice. Spends ElevenLabs credits.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
seedNo
textNo
qualityNo
loudnessNoControls the volume level of the generated voice. -1 is quietest, 1 is loudest, 0 corresponds to roughly -24 LUFS.
model_idNoModel to use for the voice generation. Possible values: eleven_multilingual_ttv_v2, eleven_ttv_v3.
output_formatNoOutput format of the generated audio. Formatted as codec_sample_rate_bitrate. So an mp3 with 22.05kHz sample rate at 32kbs is represented as mp3_22050_32. MP3 with 192kbps bitrate requires you to be subscribed to Creator tier or above. PCM with 44.1kHz sample rate requires you to be subscribed to Pr
guidance_scaleNoControls how closely the AI follows the prompt. Lower numbers give the AI more freedom to be creative, while higher numbers force it to stick more to the prompt. High numbers can cause voice to sound artificial or robotic. We recommend to use longer, more detailed prompts at lower Guidance Scale.
should_enhanceNoWhether to enhance the voice description using AI to add more detail and improve voice generation quality. When enabled, the system will automatically expand simple prompts into more detailed voice descriptions. Defaults to False
prompt_strengthNo
stream_previewsNoDetermines whether the Text to Voice previews should be included in the response. If true, only the generated IDs will be returned which can then be streamed via the /v1/text-to-voice/:generated_voice_id/stream endpoint.
voice_descriptionYesDescription to use for the created voice.
auto_generate_textNoWhether to automatically generate a text suitable for the voice description.
remixing_session_idNo
reference_audio_base64No
remixing_session_iteration_idNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

C2.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare non-read-only, non-idempotent, open-world behavior, so mutation and repeat-call semantics are covered. The description does add one genuine behavioral fact beyond annotations — that the call consumes ElevenLabs credits — which is real value for a billable operation. However, it says nothing about whether design is reversible, whether permission tiers matter, or what is returned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded, with no filler — the credit-cost warning is the one sentence that genuinely earns its place. But the extreme brevity is under-specification rather than conciseness for a 15-parameter, credit-consuming generation tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 15-parameter, non-idempotent, credit-spending voice-design tool with no output schema and only 53% schema coverage, this description is far too thin. It omits required inputs, model/format choices, preview-streaming behavior, and cost magnitude — none of which the structured fields fully carry.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema documents only about 53% of the 15 parameters, so the description would need to compensate for the remaining gap — yet it mentions no parameter at all, not even the required voice_description. Parameters like guidance_scale, prompt_strength, and stream_previews are left to be inferred from partial schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

"Design A Voice" is essentially a restatement of the tool name text_to_voice_design plus the title "Text To Voice Design" — it adds no specific verb+resource detail beyond what the identifier already conveys. It also fails to distinguish this tool from close siblings such as text_to_voice, text_to_voice_remix, text_to_voice_preview_stream, or create_voice.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance and no alternatives: an agent cannot tell from this text why it would call text_to_voice_design rather than text_to_voice or text_to_voice_remix. The only implicit cue is cost awareness, which is not usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools