Skip to main content
Glama

Why This Package?

@runapi.ai/elevenlabs-mcp is a focused Model Context Protocol server for the ElevenLabs model line on RunAPI. It gives MCP-compatible assistants direct access to 5 endpoints and 6 model variants without loading the full RunAPI catalog.

Use this per-model server when an agent should stay scoped to ElevenLabs. Use @runapi.ai/mcp when one assistant should discover every RunAPI model line.


Related MCP server: ElevenLabs MCP Server

Install

Add it to Claude Code:

claude mcp add elevenlabs -s user -- npx -y @runapi.ai/elevenlabs-mcp

Use project scope when the server should be shared with a repository:

claude mcp add elevenlabs -s project -- npx -y @runapi.ai/elevenlabs-mcp

Codex, Cursor, Windsurf, VS Code, Roo Code, and other MCP hosts can use the same stdio command:

{
  "mcpServers": {
    "elevenlabs": {
      "command": "npx",
      "args": ["-y", "@runapi.ai/elevenlabs-mcp"]
    }
  }
}

check_pricing works before sign-in. For task creation and status polling, ask your assistant to call the login tool. It opens a browser login and saves credentials to ~/.config/runapi/config.json, the same file used by runapi login. Headless and CI hosts can still set RUNAPI_API_KEY before starting the MCP host.

Ready-made examples are in examples/ for Claude, Cursor, Windsurf, VS Code, and Roo Code.


Tools

Tool

Auth

Purpose

isolate_audio

Yes

Create an ElevenLabs isolate audio task and optionally wait for a terminal status. Returns the task id, status, and output URLs.

speech_to_text

Yes

Create an ElevenLabs speech to text task and optionally wait for a terminal status. Returns the task id, status, and output URLs.

text_to_dialogue

Yes

Create an ElevenLabs text to dialogue task and optionally wait for a terminal status. Returns the task id, status, and output URLs.

text_to_sound

Yes

Create an ElevenLabs text to sound task and optionally wait for a terminal status. Returns the task id, status, and output URLs.

text_to_speech

Yes

Create an ElevenLabs text to speech task and optionally wait for a terminal status. Returns the task id, status, and output URLs.

get_task

Yes

Fetch the current status and latest payload for an existing task.

check_pricing

No

Look up current pricing for a ElevenLabs model and endpoint.


Models

ElevenLabs covers 6 model variants across 5 endpoints. Each tool accepts the models listed for it:

Tool

Models

isolate_audio

audio-isolation

speech_to_text

speech-to-text

text_to_dialogue

text-to-dialogue-v3

text_to_sound

sound-effect-v2

text_to_speech

text-to-speech-multilingual-v2, text-to-speech-turbo-v2.5

Model availability can change between releases. Use check_pricing or the ElevenLabs model page for the current catalog view.


Agent Prompts

Ask your assistant in natural language; it can inspect pricing, create the task, and return the task id plus output URLs.

Create a task

Run an ElevenLabs isolate audio task with RunAPI.

The assistant can call check_pricing, then isolate_audio, and return the task id, status, and output URLs.

Submit without waiting

Create the task but don't wait for it to finish.

The assistant calls the create tool with wait: false and returns the task id. Check on it later with get_task.

Check pricing before creating

Check current ElevenLabs pricing, then create the task if it matches my request.

The assistant calls check_pricing and can link to the ElevenLabs model page for the canonical catalog entry.


Configuration

The server resolves auth in this order:

  1. RUNAPI_API_KEY environment variable, useful for headless and CI hosts

  2. ~/.config/runapi/config.json, created by the MCP login tool or runapi login

  3. No key, which still allows check_pricing

The config file is normally managed by login. A pre-provisioned headless config can use:

{
  "apiKey": "your_runapi_key"
}

Do not commit real API keys.


Resource

URL

ElevenLabs model page

https://runapi.ai/models/elevenlabs

npm package

@runapi.ai/elevenlabs-mcp

GitHub repository

runapi-ai/elevenlabs-mcp

RunAPI MCP overview

runapi.ai/mcp

RunAPI docs

runapi.ai/docs


License

Licensed under the Apache License, Version 2.0.

Available Tools

8 tools
check_pricingB

Look up RunAPI pricing for the elevenlabs model line.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoModel slug. Defaults to the line's primary model.
actionNoEndpoint name. Defaults to the endpoint that offers the model.

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. 'Look up' implies a non-destructive read and the 'elevenlabs model line' restriction usefully bounds scope, but nothing is said about authentication, whether prices are live or cached, units/currency, or what a lookup costs in rate limits. It discloses a little beyond the structured fields but not enough for a no-annotation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with zero filler, front-loaded with the verb and the scope constraint. Nothing needs to be trimmed and nothing is buried.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter, zero-required lookup tool with full schema coverage, the description covers the essential purpose but omits the return semantics (currency, unit basis, per-endpoint breakdown) that matter when no output schema exists. It is adequate but leaves the agent guessing about what a 'pricing' result contains.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (model slug, action endpoint) and their defaults are already documented in the schema. The description adds no syntax, format, or defaulting nuance beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('look up') and resource ('pricing') and scopes it to the 'elevenlabs model line', which distinguishes it from the sibling action tools (isolate_audio, text_to_speech, etc.) that perform work rather than quote prices. It is clear, though it does not explicitly say 'this returns cost information and performs no action', leaving the read-only nature implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use or when-not-to-use guidance, no mention of prerequisites, and no indication of when an agent should consult pricing versus simply invoking an endpoint. The scope phrase 'elevenlabs model line' hints at applicability but is not framed as usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_taskB

Fetch the current status and latest result payload for a elevenlabs task.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionYesAsynchronous endpoint the task was created on.
task_idYesTask id returned when the task was created.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses the response shape at a high level (status plus latest result payload), which hints that results may be partial. However, it says nothing about behavior when the task is still pending, error cases, or any auth requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single efficient sentence with the key information front-loaded and no filler. It is arguably undersized for a polling tool with state-dependent behavior, but there is no wasted text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description must cover the polling semantics the agent needs. It identifies what is returned but omits pending/retry behavior and error handling, leaving meaningful gaps for a task-status tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: both task_id and the action enum are documented directly in the schema, including the meaning of the enum values. The description adds no additional parameter meaning (e.g. where task_id comes from), so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (fetch) and resource (task status and latest result payload), so it is immediately distinguishable from the sibling endpoints (isolate_audio, speech_to_text, etc.) which create tasks rather than read them. It stops short of explicitly contrasting itself with those siblings, but the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this tool: nothing states that it is the polling endpoint to use after creating an async task, nor how to handle non-terminal states or how often to poll. The agent must infer the entire usage pattern from the resource name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

isolate_audioC

Create a ElevenLabs task on RunAPI (isolate audio). Returns a task id, status, and output URLs.

ParametersJSON Schema
NameRequiredDescriptionDefault
waitNoPoll until the task reaches a terminal status.
modelNoRunAPI model slug for this model line.
timeout_msNo
callback_urlNoDeclared type: string.
poll_interval_msNo
source_audio_urlYesDeclared type: string.

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does add useful context by disclosing the return shape (task id, status, output URLs) even though no output schema exists. It omits auth requirements, cost implications, provider throttling, and whether polling is synchronous by default (that detail lives only in the schema).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight, front-loaded sentences with no filler; the operation and return values come first. Brevity comes at the cost of the details the other dimensions penalize, but structurally it is clean.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter async task tool with no annotations and no output schema, the description leaves the operational workflow (wait/poll behavior, get_task relationship, callback usage, timeouts) unexplained. It only covers purpose and return fields.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 67%, with several parameters documented only as placeholder prose ('Declared type: string'). The description contributes nothing about source_audio_url format, model slug selection, timeout, or callback semantics, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Create a ElevenLabs task... isolate audio') and clarifies the async task-creation nature, which separates it from synchronous siblings like text_to_speech. It does not, however, distinguish itself from other task-creating model tools (text_to_sound, text_to_dialogue) beyond the parenthetical 'isolate audio'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance on when to use this versus the sibling audio/model tools, and no mention of the natural follow-up relationship with get_task for polling a created task. The agent must infer all routing from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

loginA

Authenticate RunAPI by opening a browser PKCE login flow and saving the API key to ~/.config/runapi/config.json.

ParametersJSON Schema
NameRequiredDescriptionDefault
forceNoRe-run browser login when the current credential comes from the local config file.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It discloses the interactive browser flow and the file write side effect (config.json). However, it does not mention that it may overwrite existing credentials or that it could block waiting for user input, though these are implied.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that front-loads the action ('Authenticate RunAPI') and provides necessary details without extraneous information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple login tool with one optional parameter and no output schema, the description covers the core purpose and side effect. It lacks an explicit statement that this is a prerequisite for other tools, but that is implied.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% (the only parameter 'force' has a description). The tool description adds no additional meaning about parameters beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with a specific verb ('Authenticate'), target resource ('RunAPI'), method ('browser PKCE login flow'), and side effect (saving to config.json). It is distinct from sibling tools, none of which relate to authentication.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage (to authenticate RunAPI) but does not explicitly say when to run it (e.g., before other RunAPI tools) or when to use the 'force' parameter. Since there are no alternative auth tools among siblings, 'vs alternatives' is not applicable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

speech_to_textB

Create a ElevenLabs task on RunAPI (speech to text). Returns a task id, status, and output URLs.

ParametersJSON Schema
NameRequiredDescriptionDefault
waitNoPoll until the task reaches a terminal status.
modelNoRunAPI model slug for this model line.
diarizeNoDeclared type: boolean.
timeout_msNo
callback_urlNoDeclared type: string.
language_codeNoDeclared type: string.
poll_interval_msNo
source_audio_urlYesDeclared type: string.
tag_audio_eventsNoDeclared type: boolean.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses that this is an asynchronous task-creation call and that it returns a task id, status, and output URLs, which is meaningful behavioral context. It omits auth/permission needs, cost implications, and how the 'wait' polling mode changes the call's blocking behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action and resource, with no wasted text. Minor grammar slip ('a ElevenLabs') but the structure is efficient and readable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description compensates for the absent output schema by naming the return fields (task id, status, output URLs) and signaling the async task model. But for a 9-parameter mutation tool with no annotations, it lacks usage conditions and parameter meaning, leaving gaps an agent must fill elsewhere.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 78%, so most parameters are documented in the schema itself. The description adds no parameter-level detail (e.g., what model slugs are valid, what diarize or tag_audio_events do), leaving it at the baseline for a high-coverage schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a clear verb (Create) plus resource (task) and disambiguates direction with '(speech to text)', which separates it from the sibling text_to_speech and text_to_sound. It does not, however, name an alternative or explain how it relates to the get_task sibling used for polling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no stated preconditions, and no mention of alternatives such as get_task (for retrieving an existing task) or text_to_speech. The agent must infer context purely from the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

text_to_dialogueB

Create a ElevenLabs task on RunAPI (text to dialogue). Returns a task id, status, and output URLs.

ParametersJSON Schema
NameRequiredDescriptionDefault
waitNoPoll until the task reaches a terminal status.
modelNoRunAPI model slug for this model line.
dialogueYesDeclared type: array.
stabilityNoDeclared type: number. Known values: 0, 0.5, 1.
timeout_msNo
callback_urlNoDeclared type: string.
language_codeNoDeclared type: string.
poll_interval_msNo

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose the async task contract by stating it returns a task id, status, and output URLs. However, it says nothing about cost/billing, auth requirements, failure modes, or how `wait`/polling affects behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action and resource, with the return contract appended. Only minor issue is the awkward 'a ElevenLabs' phrasing; no wasted content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter, no-annotation, no-output-schema tool, the description partially compensates by describing the return payload. It leaves opaque parameters (dialogue structure, language_code, callback_url semantics) and async/timeout behavior unaddressed, so it is only minimally complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, below the high-coverage baseline, and the description adds no parameter detail. The single required parameter `dialogue` is documented only as 'Declared type: array' with no item shape, so an agent gets no help on the most critical input.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb ('Create') plus resource ('ElevenLabs task on RunAPI') and qualifies it as text-to-dialogue, which loosely distinguishes it from text_to_speech and text_to_sound. It does not explicitly name those siblings, so differentiation is left implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to choose dialogue synthesis over text_to_speech or text_to_sound, no prerequisites, and no mention of the related get_task follow-up for retrieving results. Usage must be inferred entirely from the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

text_to_soundC

Create a ElevenLabs task on RunAPI (text to sound). Returns a task id, status, and output URLs.

ParametersJSON Schema
NameRequiredDescriptionDefault
loopNoDeclared type: boolean.
textYesDeclared type: string.
waitNoPoll until the task reaches a terminal status.
modelNoRunAPI model slug for this model line.
timeout_msNo
callback_urlNoDeclared type: string.
output_formatNoDeclared type: string. Known values: "mp3_22050_32", "mp3_44100_32", "mp3_44100_64", "mp3_44100_96", "mp3_44100_128", "mp3_44100_192", "pcm_8000", "pcm_16000", "pcm_22050", "pcm_24000", "pcm_44100", "pcm_48000", "ulaw_8000", "alaw_8000", "opus_48000_32", "opus_48000_64", "opus_48000_96", "opus_48000_128", "opus_48000_192".
duration_secondsYesDeclared type: number.
poll_interval_msNo
prompt_influenceNoDeclared type: number.

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose that this is an asynchronous task submission returning a task id, status, and output URLs, plus the implicit polling behavior of the 'wait' parameter. It says nothing about authentication, rate limits, cost, or what state a task is in when 'wait' is false (the get_task sibling is never referenced).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the action front-loaded and the return shape following. Every sentence carries information, though the second sentence overlaps with what a reader would infer from the task-based naming.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter mutation tool with no annotations and no output schema, the description does useful work by stating the return values. It remains incomplete on how the many optional parameters interact, whether the call blocks by default, and how this differs from text_to_speech or when to fall back to get_task.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 80%, so the schema already documents most parameters (including the 'wait' polling semantics and enum values for output_format). The description adds no parameter-level detail beyond what the schema provides, which is the expected baseline of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource ('Create a ElevenLabs task on RunAPI (text to sound)'), so the core action is clear. However, it does not distinguish this tool from the near-identical sibling 'text_to_speech', and the parenthetical 'text to sound' only deepens the ambiguity rather than resolving it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no mention of any alternative. With siblings like text_to_speech, text_to_dialogue, and get_task in the same namespace, the agent is left to guess which one to call.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

text_to_speechB

Create a ElevenLabs task on RunAPI (text to speech). Returns a task id, status, and output URLs.

ParametersJSON Schema
NameRequiredDescriptionDefault
textYesDeclared type: string.
waitNoPoll until the task reaches a terminal status.
modelNoRunAPI model slug for this model line.
speedNoDeclared type: number.
styleNoDeclared type: number.
voiceNoDeclared type: string.
next_textNoDeclared type: string.
stabilityNoDeclared type: number.
timeout_msNo
timestampsNoDeclared type: boolean.
callback_urlNoDeclared type: string.
language_codeNoDeclared type: string.
previous_textNoDeclared type: string.
poll_interval_msNo
similarity_boostNoDeclared type: number.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses that the call is asynchronous and returns a task id, status, and output URLs, which implies a task-based workflow. However, it omits auth/permission needs, cost implications, whether the returned task may still be pending, and how wait/poll_interval_ms/timeout_ms interact.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the action and return shape front-loaded, and no filler. It is slightly terse given the 15-parameter surface, but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 15-parameter, no-output-schema tool, the description covers the return contract (task id, status, output URLs) but says nothing about the model/voice/audio-tuning parameters or the polling behavior. With 87% schema coverage the structured fields do most of the work, but the description leaves real gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 87%, so the schema already documents most parameters (e.g., wait as 'Poll until the task reaches a terminal status'). The description adds no parameter-level meaning beyond that, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (Create), resource (ElevenLabs task on RunAPI), and clarifies the modality as text-to-speech, which distinguishes it from siblings like speech_to_text, text_to_sound, and text_to_dialogue. It does not explicitly name which sibling to use instead, but the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus text_to_sound, text_to_dialogue, or speech_to_text, nor any mention of when to set wait=true versus polling get_task manually. The agent is left to infer usage entirely from the name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 7 tool updatesv0.2.0
    • Changedcheck_pricing2 fields changed
      • removedInput schema / additionalProperties
        Removed value: -false
      • removedInput schema / properties / model / enum
        Removed value: -[
        -  "audio-isolation",
        -  "speech-to-text",
        -  "text-to-dialogue-v3",
        -  "sound-effect-v2",
        -  "text-to-speech-multilingual-v2",
        -  "text-to-speech-turbo-v2.5"
        -]
    • Changedget_task1 field changed
      • removedInput schema / additionalProperties
        Removed value: -false
    • Changedisolate_audio6 fields changed
      • changedInput schema / additionalProperties
        Previous value: -falseNew value: +{}
      • addedInput schema / properties / callback_url / description
        Added value: +"Declared type: string."
      • removedInput schema / properties / model / enum
        Removed value: -[
        -  "audio-isolation"
        -]
      • addedInput schema / properties / poll_interval_ms / maximum
        Added value: +9007199254740991
      • addedInput schema / properties / source_audio_url / description
        Added value: +"Declared type: string."
      • addedInput schema / properties / timeout_ms / maximum
        Added value: +9007199254740991
    • Changedspeech_to_text9 fields changed
      • changedInput schema / additionalProperties
        Previous value: -falseNew value: +{}
      • addedInput schema / properties / callback_url / description
        Added value: +"Declared type: string."
      • addedInput schema / properties / diarize / description
        Added value: +"Declared type: boolean."
      • addedInput schema / properties / language_code / description
        Added value: +"Declared type: string."
      • removedInput schema / properties / model / enum
        Removed value: -[
        -  "speech-to-text"
        -]
      • addedInput schema / properties / poll_interval_ms / maximum
        Added value: +9007199254740991
      • addedInput schema / properties / source_audio_url / description
        Added value: +"Declared type: string."
      • addedInput schema / properties / tag_audio_events / description
        Added value: +"Declared type: boolean."
      • addedInput schema / properties / timeout_ms / maximum
        Added value: +9007199254740991
    • Changedtext_to_dialogue9 fields changed
      • changedInput schema / additionalProperties
        Previous value: -falseNew value: +{}
      • addedInput schema / properties / callback_url / description
        Added value: +"Declared type: string."
      • addedInput schema / properties / dialogue / description
        Added value: +"Declared type: array."
      • addedInput schema / properties / language_code / description
        Added value: +"Declared type: string."
      • removedInput schema / properties / model / enum
        Removed value: -[
        -  "text-to-dialogue-v3"
        -]
      • addedInput schema / properties / poll_interval_ms / maximum
        Added value: +9007199254740991
      • addedInput schema / properties / stability / description
        Added value: +"Declared type: number. Known values: 0, 0.5, 1."
      • removedInput schema / properties / stability / enum
        Removed value: -[
        -  0,
        -  0.5,
        -  1
        -]
      • addedInput schema / properties / timeout_ms / maximum
        Added value: +9007199254740991
    • Changedtext_to_sound12 fields changed
      • changedInput schema / additionalProperties
        Previous value: -falseNew value: +{}
      • addedInput schema / properties / callback_url / description
        Added value: +"Declared type: string."
      • addedInput schema / properties / duration_seconds / description
        Added value: +"Declared type: number."
      • addedInput schema / properties / loop / description
        Added value: +"Declared type: boolean."
      • removedInput schema / properties / model / enum
        Removed value: -[
        -  "sound-effect-v2"
        -]
      • addedInput schema / properties / output_format / description
        Added value: +"Declared type: string. Known values: \"mp3_22050_32\", \"mp3_44100_32\", \"mp3_44100_64\", \"mp3_44100_96\", \"mp3_44100_128\", \"mp3_44100_192\", \"pcm_8000\", \"pcm_16000\", \"pcm_22050\", \"pcm_24000\", \"pcm_44100\", \"pcm_48000\", \"ulaw_8000\", \"alaw_8000\", \"opus_48000_32\", \"opus_48000_64\", \"opus_48000_96\", \"opus_48000_128\", \"opus_48000_192\"."
      • removedInput schema / properties / output_format / enum
        Removed value: -[
        -  "mp3_22050_32",
        -  "mp3_44100_32",
        -  "mp3_44100_64",
        -  "mp3_44100_96",
        -  "mp3_44100_128",
        -  "mp3_44100_192",
        -  "pcm_8000",
        -  "pcm_16000",
        -  "pcm_22050",
        -  "pcm_24000",
        -  "pcm_44100",
        -  "pcm_48000",
        -  "ulaw_8000",
        -  "alaw_8000",
        -  "opus_48000_32",
        -  "opus_48000_64",
        -  "opus_48000_96",
        -  "opus_48000_128",
        -  "opus_48000_192"
        -]
      • addedInput schema / properties / poll_interval_ms / maximum
        Added value: +9007199254740991
      • addedInput schema / properties / prompt_influence / description
        Added value: +"Declared type: number."
      • addedInput schema / properties / text / description
        Added value: +"Declared type: string."
      • addedInput schema / properties / timeout_ms / maximum
        Added value: +9007199254740991
      • changedInput schema / required
        Previous value: -[
        -  "text"
        -]New value: +[
        +  "text",
        +  "duration_seconds"
        +]
    • Changedtext_to_speech15 fields changed
      • changedInput schema / additionalProperties
        Previous value: -falseNew value: +{}
      • addedInput schema / properties / callback_url / description
        Added value: +"Declared type: string."
      • addedInput schema / properties / language_code / description
        Added value: +"Declared type: string."
      • removedInput schema / properties / model / enum
        Removed value: -[
        -  "text-to-speech-multilingual-v2",
        -  "text-to-speech-turbo-v2.5"
        -]
      • addedInput schema / properties / next_text / description
        Added value: +"Declared type: string."
      • addedInput schema / properties / poll_interval_ms / maximum
        Added value: +9007199254740991
      • addedInput schema / properties / previous_text / description
        Added value: +"Declared type: string."
      • addedInput schema / properties / similarity_boost / description
        Added value: +"Declared type: number."
      • addedInput schema / properties / speed / description
        Added value: +"Declared type: number."
      • addedInput schema / properties / stability / description
        Added value: +"Declared type: number."
      • addedInput schema / properties / style / description
        Added value: +"Declared type: number."
      • addedInput schema / properties / text / description
        Added value: +"Declared type: string."
      • addedInput schema / properties / timeout_ms / maximum
        Added value: +9007199254740991
      • addedInput schema / properties / timestamps / description
        Added value: +"Declared type: boolean."
      • addedInput schema / properties / voice / description
        Added value: +"Declared type: string."
  2. 6 tool updatesv0.1.7
    • Changedget_task1 field changed
      • changedInput schema / properties / action / description
        Previous value: -"Endpoint the task was created on."New value: +"Asynchronous endpoint the task was created on."
    • Changedisolate_audio3 fields changed
      • addedInput schema / properties / callback_url
        Added value: +{
        +  "type": "string"
        +}
      • addedInput schema / properties / source_audio_url / type
        Added value: +"string"
      • addedInput schema / required
        Added value: +[
        +  "source_audio_url"
        +]
    • Changedspeech_to_text6 fields changed
      • addedInput schema / properties / callback_url
        Added value: +{
        +  "type": "string"
        +}
      • addedInput schema / properties / diarize
        Added value: +{
        +  "type": "boolean"
        +}
      • addedInput schema / properties / language_code
        Added value: +{
        +  "type": "string"
        +}
      • addedInput schema / properties / source_audio_url / type
        Added value: +"string"
      • addedInput schema / properties / tag_audio_events
        Added value: +{
        +  "type": "boolean"
        +}
      • addedInput schema / required
        Added value: +[
        +  "source_audio_url"
        +]
    • Changedtext_to_dialogue5 fields changed
      • addedInput schema / properties / callback_url
        Added value: +{
        +  "type": "string"
        +}
      • addedInput schema / properties / dialogue / items
        Added value: +{}
      • addedInput schema / properties / dialogue / type
        Added value: +"array"
      • addedInput schema / properties / language_code
        Added value: +{
        +  "type": "string"
        +}
      • addedInput schema / required
        Added value: +[
        +  "dialogue"
        +]
    • Changedtext_to_sound6 fields changed
      • addedInput schema / properties / callback_url
        Added value: +{
        +  "type": "string"
        +}
      • addedInput schema / properties / duration_seconds
        Added value: +{
        +  "type": "number"
        +}
      • addedInput schema / properties / loop
        Added value: +{
        +  "type": "boolean"
        +}
      • addedInput schema / properties / prompt_influence
        Added value: +{
        +  "type": "number"
        +}
      • addedInput schema / properties / text / type
        Added value: +"string"
      • addedInput schema / required
        Added value: +[
        +  "text"
        +]
    • Changedtext_to_speech12 fields changed
      • addedInput schema / properties / callback_url
        Added value: +{
        +  "type": "string"
        +}
      • addedInput schema / properties / language_code
        Added value: +{
        +  "type": "string"
        +}
      • addedInput schema / properties / next_text
        Added value: +{
        +  "type": "string"
        +}
      • addedInput schema / properties / previous_text
        Added value: +{
        +  "type": "string"
        +}
      • addedInput schema / properties / similarity_boost
        Added value: +{
        +  "type": "number"
        +}
      • addedInput schema / properties / speed
        Added value: +{
        +  "type": "number"
        +}
      • addedInput schema / properties / stability
        Added value: +{
        +  "type": "number"
        +}
      • addedInput schema / properties / style
        Added value: +{
        +  "type": "number"
        +}
      • addedInput schema / properties / text / type
        Added value: +"string"
      • addedInput schema / properties / timestamps
        Added value: +{
        +  "type": "boolean"
        +}
      • addedInput schema / properties / voice
        Added value: +{
        +  "type": "string"
        +}
      • addedInput schema / required
        Added value: +[
        +  "text"
        +]
  3. 1 tool updatev0.1.6
    • Addedlogin
  4. 7 tool updatesv0.1.0
    • First observedcheck_pricing
    • First observedget_task
    • First observedisolate_audio
    • First observedspeech_to_text
    • First observedtext_to_dialogue
    • First observedtext_to_sound
    • First observedtext_to_speech

TDQS

B3.3/5.0

Scored across 8 tools

Disambiguation3/5

The create-task tools share identical boilerplate descriptions, and text_to_speech vs text_to_sound (speech vs sound effects) are easily confused without further detail. login, get_task, and check_pricing are clearly distinct, but the audio-generation cluster has fuzzy boundaries.

Naming Consistency4/5

All names use consistent snake_case, and the *_to_* / verb_noun pattern is largely predictable. The lone bare verb 'login' is a minor deviation from the otherwise uniform resource-action style.

Tool Count5/5

Eight tools is a well-scoped set for an audio task API covering generation, transcription, isolation, task polling, auth, and pricing. Each tool earns its place with no redundant entries.

Completeness3/5

Core generation and retrieval are covered, but the async task surface is incomplete: there is no list_tasks, cancel_task, or delete_task, which agents commonly need to manage submitted jobs. Voice/model enumeration is also absent.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers