Skip to main content
Glama
AIWerk

@aiwerk/mcp-server-elevenlabs

by AIWerk

text_to_voice_remix

Modify an existing voice by applying described changes to its characteristics. Produces a new voice variant from a voice ID and a change description.

Instructions

Remix A Voice. Spends ElevenLabs credits.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
seedNo
textNo
loudnessNoControls the volume level of the generated voice. -1 is quietest, 1 is loudest, 0 corresponds to roughly -24 LUFS.
voice_idYesVoice ID to be used, you can use https://api.elevenlabs.io/v1/voices to list all the available voices.
output_formatNoOutput format of the generated audio. Formatted as codec_sample_rate_bitrate. So an mp3 with 22.05kHz sample rate at 32kbs is represented as mp3_22050_32. MP3 with 192kbps bitrate requires you to be subscribed to Creator tier or above. PCM with 44.1kHz sample rate requires you to be subscribed to Pr
guidance_scaleNoControls how closely the AI follows the prompt. Lower numbers give the AI more freedom to be creative, while higher numbers force it to stick more to the prompt. High numbers can cause voice to sound artificial or robotic. We recommend to use longer, more detailed prompts at lower Guidance Scale.
prompt_strengthNo
stream_previewsNoDetermines whether the Text to Voice previews should be included in the response. If true, only the generated IDs will be returned which can then be streamed via the /v1/text-to-voice/:generated_voice_id/stream endpoint.
voice_descriptionYesDescription of the changes to make to the voice.
auto_generate_textNoWhether to automatically generate a text suitable for the voice description.
remixing_session_idNo
remixing_session_iteration_idNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

C2.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, and openWorldHint=true. The description adds the useful behavioral fact that the operation spends ElevenLabs credits, but says nothing about required permissions, rate limits, or what happens to the original voice.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The two short sentences are concise and front-load a rough purpose and a cost warning, but the first sentence adds little beyond the name, and the overall structure is too sparse to guide correct invocation of a 12-parameter tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex mutation tool with 12 parameters, no output schema, and only partial schema descriptions, the description is severely incomplete. It omits usage context, parameter semantics, and nearly all behavioral details beyond a generic cost note.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description provides no parameter-level guidance despite 12 parameters and only 58% schema description coverage. Several nullable parameters (seed, text, prompt_strength, remixing_session_id, remixing_session_iteration_id) are undocumented in both the schema and the description, so the description fails to compensate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description 'Remix A Voice' restates the tool name almost verbatim, and adds only a billing note ('Spends ElevenLabs credits'). It does not explain what remixing a voice actually does or how it differs from siblings like text_to_voice, text_to_voice_design, or text_to_voice_preview_stream.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool instead of alternatives such as text_to_voice or text_to_voice_design. The only usage-relevant information is the cost warning, which is a constraint but not selection guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools