Skip to main content
Glama
AIWerk

@aiwerk/mcp-server-elevenlabs

by AIWerk

create_video_generation

Generate videos from prompts with optional image, audio, or video references; control resolution, aspect ratio, duration, and audio.

Instructions

Create Video Generation

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
seedNo
audioNo
imageNo
audiosNoUp to 10 reference audios, e.g. for lipsync. Cannot be combined with `start_frame`/`end_frame`.
imagesNoUp to 30 reference images to draw subjects from. Cannot be combined with `start_frame`/`end_frame`.
promptNoA text description of the video to generate.
videosNoUp to 10 reference videos to draw subjects or motion from. Cannot be combined with `start_frame`/`end_frame`.
webhookNo
model_idNoThe model to use for the generation.
end_frameNo
resolutionNoThe resolution of the output video.
start_frameNo
aspect_ratioNoThe aspect ratio of the output video. With `auto`, the model picks an aspect ratio based on the inputs. First-frame / first-and-last-frame tasks always use `auto`.
duration_secsNoThe duration of the output video in seconds.
enhance_promptNoWhether the model may rewrite the prompt to improve results.
generate_audioNoWhether to generate audio with the video.
guidance_scaleNo
negative_promptNo
audio_guidance_scaleNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

D1.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, and openWorldHint=true. The description adds nothing about asynchronous generation behavior, asset prerequisites, webhook delivery, long-running job handling, or prompt/model requirements, so it provides no behavioral context beyond the structured annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a two-word title with no front-loaded information or structure. It is under-specified rather than concise, since it omits every useful detail an agent would need.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 19 parameters, nested media objects, no output schema, and no required parameters, the description provides no context about what generation entails or how to invoke it correctly. It is completely inadequate for the complexity of the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has 19 parameters and only 53% schema description coverage, leaving fields such as seed, type discriminators, guidance_scale, negative_prompt, and audio_guidance_scale without documented meaning. The description mentions no parameters and adds no semantics beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description exactly restates the tool name and title, 'Create Video Generation,' giving no scope or detail beyond the name. It does not distinguish this tool from siblings such as create_image_generation, get_video_generation, or list_video_generations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites, and no indication of when to choose this tool over alternatives like create_image_generation or list_video_generations. Usage is left entirely to inference from the name and schema.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools