Skip to main content
Glama
AIWerk

@aiwerk/mcp-server-elevenlabs

by AIWerk

text_to_voice

Generate a voice preview from a written description using ElevenLabs, spending credits. Deprecated upstream.

Instructions

[Deprecated] Generate A Voice Preview From Description Spends ElevenLabs credits. Deprecated upstream.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
seedNo
textNo
qualityNoHigher quality results in better voice output but less variety.
loudnessNoControls the volume level of the generated voice. -1 is quietest, 1 is loudest, 0 corresponds to roughly -24 LUFS.
output_formatNoOutput format of the generated audio. Formatted as codec_sample_rate_bitrate. So an mp3 with 22.05kHz sample rate at 32kbs is represented as mp3_22050_32. MP3 with 192kbps bitrate requires you to be subscribed to Creator tier or above. PCM with 44.1kHz sample rate requires you to be subscribed to Pr
guidance_scaleNoControls how closely the AI follows the prompt. Lower numbers give the AI more freedom to be creative, while higher numbers force it to stick more to the prompt. High numbers can cause voice to sound artificial or robotic. We recommend to use longer, more detailed prompts at lower Guidance Scale.
should_enhanceNoWhether to enhance the voice description using AI to add more detail and improve voice generation quality. When enabled, the system will automatically expand simple prompts into more detailed voice descriptions. Defaults to False
voice_descriptionYesDescription to use for the created voice.
auto_generate_textNoWhether to automatically generate a text suitable for the voice description.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

C2.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already disclose readOnlyHint=false, idempotentHint=false and openWorldHint=true, so the safety profile is covered. The description adds a genuinely useful trait not in the annotations: the call spends ElevenLabs credits, and it is deprecated upstream, both of which materially affect tool selection. It still omits whether a persistent voice is created versus a throwaway preview.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is short and front-loads the deprecation notice, which is the right priority. But the sentence runs together into an ambiguous fragment ('Generate A Voice Preview From Description Spends ElevenLabs credits'), making the cost statement read as part of the purpose rather than a separate fact.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter, credit-consuming generation tool with no output schema and no return-value description, the definition is thin. It never clarifies that the result is an audio preview rather than a saved voice, nor does it quantify the credit cost or point at a replacement, leaving key decision inputs missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 78%, close to the high-coverage baseline, so the schema already carries most parameter meaning (quality, loudness, guidance_scale, output_format, should_enhance are documented inline). The description adds no parameter detail whatsoever, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a concrete verb and resource ('Generate A Voice Preview From Description'), which is more specific than a tautology. However, it does not distinguish this tool from close siblings such as text_to_voice_design, text_to_voice_remix, or text_to_voice_preview_stream, so an agent cannot tell which voice-generation variant to pick from the text alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The '[Deprecated]' and 'Deprecated upstream' markers signal that this tool should be avoided, which is weak negative guidance, but no alternative is named and no condition for legitimate use is given. An agent is left to infer that a sibling deprecated-adjacent tool should be used instead.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools