Skip to main content
Glama
AIWerk

@aiwerk/mcp-server-elevenlabs

by AIWerk

create_dubbing

Dubs video or audio files from a source language into a target language via ElevenLabs, consuming credits.

Instructions

Dub A Video Or An Audio File Spends ElevenLabs credits.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
modeNoThe mode in which to run this Dubbing job. Defaults to automatic, use manual if specifically providing a CSV transcript to use. Note that manual mode is experimental and production use is strongly discouraged.
nameNoName of the dubbing project.
csv_fpsNoFrames per second to use when parsing a CSV file for dubbing. If not provided, FPS will be inferred from timecodes.
end_timeNoEnd time of the source video/audio file.
file_pathNoA list of file paths to audio recordings intended for voice cloning Local path.
watermarkNoWhether to apply watermark to the output video.
source_urlNoURL of the source video/audio file.
start_timeNoStart time of the source video/audio file.
file_base64NoBase64 contents for "file". Use this when the server cannot read your local disk.
source_langNoSource language. Expects a valid iso639-1 or iso639-3 language code.
target_langNoThe Target language to dub the content into. Expects a valid iso639-1 or iso639-3 language code.
num_speakersNoNumber of speakers to use for the dubbing. Set to 0 to automatically detect the number of speakers
csv_file_pathNoCSV file containing transcription/translation metadata Local path.
file_filenameNoFilename to send for "file". Some endpoints infer the audio format from it.
target_accentNo[Experimental] An accent to apply when selecting voices from the library and to use to inform translation of the dialect to prefer.
dubbing_studioNoWhether to prepare dub for edits in dubbing studio or edits as a dubbing resource.
csv_file_base64NoBase64 contents for "csv_file". Use this when the server cannot read your local disk.
csv_file_filenameNoFilename to send for "csv_file". Some endpoints infer the audio format from it.
highest_resolutionNoWhether to use the highest resolution available.
use_profanity_filterNo[BETA] Whether transcripts should have profanities censored with the words '[censored]'
disable_voice_cloningNoInstead of using a voice clone in dubbing, use a similar voice from the ElevenLabs Voice Library. Voices used from the library will contribute towards a workspace's custom voices limit, and if there aren't enough available slots the dub will fail. Using this feature requires the caller to have the '
drop_background_audioNoAn advanced setting. Whether to drop background audio from the final dub. This can improve dub quality where it's known that audio shouldn't have a background track such as for speeches or monologues.
background_audio_file_pathNoFor use only with csv input Local path.
foreground_audio_file_pathNoFor use only with csv input Local path.
background_audio_file_base64NoBase64 contents for "background_audio_file". Use this when the server cannot read your local disk.
foreground_audio_file_base64NoBase64 contents for "foreground_audio_file". Use this when the server cannot read your local disk.
background_audio_file_filenameNoFilename to send for "background_audio_file". Some endpoints infer the audio format from it.
foreground_audio_file_filenameNoFilename to send for "foreground_audio_file". Some endpoints infer the audio format from it.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, openWorldHint=true, so the safety profile is covered. The description adds one genuinely useful behavior trait beyond annotations — that the call consumes ElevenLabs credits — but says nothing about the async nature of the job, required languages, or what happens on failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is very short and the credit-cost clause carries real information, but the phrasing is a run-on fragment with inconsistent capitalization. It is tersely front-loaded yet reads more like a title than a structured statement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 28-parameter dubbing job with zero required fields and no output schema, the description is far too thin: it omits how to supply input (URL/base64/local path), the language-code requirements, the automatic vs manual mode trade-off, and the credit-cost magnitude. Only the credit-consumption warning is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 28 parameters are already documented in the schema. The description adds no parameter meaning (no explanation of source_url vs file_path vs file_base64 selection, no language-code requirements), so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource ("Dub a video or an audio file"), so an agent can tell what the tool performs. However, it does not differentiate from close siblings like `dub` or `dubbing_project_create`, and the garbled Title Case phrasing with the run-on "Spends ElevenLabs credits." makes the scope slightly ambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no routing between this and the sibling `dub`/`dubbing_project_create` tools. The credit note hints at cost but does not tell the agent when invoking this is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools