Skip to main content
Glama

ModelsLab

Text to Video

text-to-video

Generate videos from text descriptions. Creates AI-generated videos based on your text prompt. Returns a request ID that can be used with fetch-video to retrieve results.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
fpsNoFrames per second for the output video.
widthNoVideo width in pixels (512-1024).
heightNoVideo height in pixels (512-1024).
promptYesText description of the video to generate.
webhookNoURL to receive webhook notification when generation completes.
durationNoVideo duration in seconds (minimum 4).
model_idYesThe model ID to use for video generation.
portraitNoGenerate in portrait orientation.
track_idNoCustom tracking ID for the request.
init_audioNoURL of audio to sync with the video.
resolutionNoOutput resolution preset.
aspect_ratioNoAspect ratio for the video (e.g., "16:9", "9:16").
camera_fixedNoKeep camera position fixed during generation.
enhance_promptNoUse AI to enhance the prompt.
generate_audioNoGenerate audio for the video.
negative_promptNoThings to avoid in the generated video.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only supply openWorldHint=true, so the description carries most of the burden. It usefully discloses that generation is asynchronous and yields a request ID rather than a video directly, which is the single most important behavioral fact here. It omits cost, latency, and rate-limit behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The first two sentences are near-duplicates ('Generate videos from text descriptions' vs 'Creates AI-generated videos based on your text prompt'), wasting a line. The return-value/next-step sentence is the only one carrying new information, so the text is not front-loaded efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 16 parameters, no output schema, and only a weak openWorld annotation, the description does cover the one thing the schema cannot: that the call returns a request ID to be polled via fetch-video. But it says nothing about model choice, defaults for the 14 optional params (fps, resolution, aspect_ratio), or generation constraints, leaving real gaps for a high-parameter tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% across all 16 parameters, so the schema already documents fps, resolution, aspect_ratio, negative_prompt, etc. The description adds no parameter-level meaning beyond that, which is the expected baseline when the schema does the work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Generate videos from text descriptions') and confirms it creates AI-generated video from a prompt, which cleanly separates it from siblings like image-to-video and video-to-video. It never explicitly names those alternatives, but the resource is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the async workflow by noting the result is a request ID consumed by fetch-video, which is genuinely useful context. However, it gives no guidance on when to pick this tool over image-to-video or video-to-video, nor any prerequisites (e.g., model selection via list-models).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources