Skip to main content
Glama
AIWerk

@aiwerk/mcp-server-elevenlabs

by AIWerk

video_to_music

Generate custom music from one or more video files using optional style tags or a description, then save the returned audio as a ZIP.

Instructions

Video To Music Spends ElevenLabs credits. Returns application/zip bytes; pass output_path to save them.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
tagsNoOptional list of style tags (e.g. ['upbeat', 'cinematic']). A maximum of 10 tags is allowed.
model_idNo
descriptionNoOptional text description of the music you want. A maximum of 1000 characters is allowed.
output_pathNoWhere to write the returned bytes. Relative paths resolve against ELEVENLABS_OUTPUT_DIR. Omit it to get the data inline as base64 (small files only).
videos_pathsNoOne or more video files sent via FormData array (multipart/form-data). They will be combined into one codec in order. A maximum of 10 videos is allowed, where the total size of the combined video is limited to 200MB. In total, the video can be up to 600 seconds long. Note that combining multiple vid
output_formatNoOutput format of the generated audio. Formatted as codec_sample_rate_bitrate. So an mp3 with 22.05kHz sample rate at 32kbs is represented as mp3_22050_32. MP3 with 192kbps bitrate requires you to be subscribed to Creator tier or above. PCM with 44.1kHz sample rate requires you to be subscribed to Pr
sign_with_c2paNoWhether to sign the generated song with C2PA. Applicable only for mp3 files.
videos_filenamesNoFilenames to send for "videos". Some endpoints infer the audio format from them.
videos_base64_listNoBase64 contents for "videos", one entry per file. Use this when the server cannot read your local disk.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

C2.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Despite sparse annotations (readOnlyHint=false, destructiveHint=false, openWorldHint=true), the description adds genuinely useful behavioral context: that it spends ElevenLabs credits and that it returns application/zip bytes with output_path controlling persistence. This is real value beyond the annotation set, though it omits failure modes and size/time limits already noted in the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the credit-cost caveat front-loaded, which is the most decision-relevant fact. No filler, though the extreme terseness leaves the core operation unexplained.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly covers the return type (zip bytes) and how to persist it, which is important for a 9-parameter tool. However, it omits any statement of the actual transformation performed and any credit-amount or size guidance, leaving the agent without enough to confidently select the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 89%, so the schema already documents parameters like output_path, tags, and output_format. The description only echoes output_path, adding no syntax or format detail beyond what the structured schema provides; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description never states the actual operation – that it generates music from video input – it merely restates the tool name ('Video To Music') and then talks about credits and output bytes. No verb+resource explanation of what the tool produces distinguishes it from siblings like sound_generation or create_video_generation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no mention of when to prefer this over sound_generation, create_video_generation, or separate_song_stems, and no prerequisites described. The only implicit constraint is the credit cost, which is not framed as a selection criterion.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools