Skip to main content
Glama

generate_video

Generate a video. Default model: seedance_1_5_pro at 480p, 8s, no audio (cheapest reasonable quality). Before calling, if the user has not specified them, briefly ASK for: (1) aspect ratio (9:16 vertical, 16:9 horizontal, 1:1 square), (2) duration (4/8/12s), (3) whether they want audio, (4) resolution (480p/720p/1080p) — higher costs more. Skip the questions only when the user already gave you those details or asked for "the cheapest". For character consistency across shots (e.g. a recurring person), pass extra face/body anchor photos in reference_image_urls — Seedance 2 / 2-fast use up to 7 references. MOTION CONTROL: model kling_2_6_motion transfers the motion of a driving video onto the person in your image — pass the person photo in image_url AND the driving clip URL in motion_video_url (both required); set duration to the driving clip length (billed ~11 coins/s, min 3s). Returns job_id; poll check_job until status=done.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
modelNo
promptYes
durationNo
image_urlNoFirst-frame / primary input image for image-to-video. Optional. For multi-ref, also pass reference_image_urls.
resolutionNo480p cheapest; 1080p ~5x cost. Default 480p.
with_audioNoGenerate sound. SeeDance/Kling support audio — typically doubles cost. Default false.
aspect_ratioNo
motion_video_urlNoDriving/reference VIDEO URL for Motion Control (model kling_2_6_motion only). The motion of this clip is applied to the person in image_url. Required for kling_2_6_motion; ignored by other models. Must be a publicly reachable mp4/mov.
reference_image_urlsNoExtra reference images (Seedance 2 / 2-fast accept up to 7 total including image_url). Use for character face/body anchors, multiple people, style references. Order matters — most important reference first.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden, and it does most of it: it states default model/resolution/duration/audio, cost multipliers for audio and resolution, a concrete billing rate for motion control (~11 coins/s, min 3s), the required parameter pairing for kling_2_6_motion, reference-count limits per model, and the async contract (returns job_id, poll check_job until status=done). It omits auth requirements, failure/retry behavior, and quota limits, so it falls short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The default configuration is front-loaded, which is the right priority, and most sentences carry actionable content (costs, required pairs, polling). It is dense and long, with a few parenthetical asides that could be trimmed, but little is pure filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter async generation tool with no annotations and no output schema, the description covers the essentials an agent needs: defaults for unspecified inputs, the clarifying-question protocol, model-specific parameter requirements, cost signals, and the return/polling contract. It stops short of error handling, authorization, and output shape beyond job_id, so a 5 is not warranted.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 56%, so the description must compensate, and it does add real meaning: it names the default model, explains that duration options are 4/8/12s, that aspect ratio maps to vertical/horizontal/square, that audio roughly doubles cost, and that resolution tiers scale cost to 1080p at ~5x. It does not explain the remaining enum members (grok_*, veo_3, kling_2_6) or the full duration enum, leaving some gaps against the 56% coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource ("Generate a video") and immediately scopes it with defaults and model family. It clearly separates this from generate_image and check_job, the latter named explicitly as the follow-up poll tool. An agent knows exactly what this tool produces.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-do-what rules: ask for the four unspecified options before calling, skip the questions if the user already supplied them or asked for "the cheapest", and use kling_2_6_motion only under a specific stated condition. Alternatives and their trigger conditions are spelled out rather than implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.