Skip to main content
Glama

Generate video clip

generate_video_clip

Create a video clip from a text prompt, source images, videos, spoken dialogue, or reference audio, with optional quality and aspect ratio.

Instructions

Generate a video clip from a text prompt, source images, source videos, spokenDialogue, or reference audio. quality is optional (LOW, STANDARD, HIGH, or MAX). Typically takes 1–3 minutes (HIGH/MAX can be longer). Tell the user that wait up front and keep polling calmly.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
promptNoDescription of the video to generate.
qualityNoVideo generation quality. Omit to use workspace settings.
aspectRatioNoOutput aspect ratio as width:height units (not pixels). Example: { width: 16, height: 9 }.
audioFileIdsNoOptional reference audio file ids used to lip-sync from a recording. Use spokenDialogue to have the model speak a line it generates itself.
imageFileIdsNoOptional reference image file ids.
videoFileIdsNoOptional reference video file ids.
generateAudioNoWhether the result must include generated audio.
spokenDialogueNoExact line the subject should speak as native lip-synced speech. The model synthesizes the voice from this text.
durationSecondsNoOptional clip length in whole seconds (1 to 30). Omit or pass null for Auto.
startFrameFileIdNoOptional opening-frame still file id. Used as the first frame of the clip.
voiceDescriptionNoNatural-language description of the voice that speaks spokenDialogue. Used when spokenDialogue is set.
suppressBackgroundMusicNoWhen true, the clip will not include a musical soundtrack. Spoken dialogue and environmental sound are still allowed. Use this when background music will be added separately.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
errorNoError details when status is failed; otherwise null.
statusNoJob status: pending, running, succeeded, failed, or cancelled.
resultsNoGenerated results with download URLs when succeeded.
toolTypeNoTool name (e.g. GENERATE_IMAGE).
attemptIndexNoCurrent or latest attempt index.
toolExecutionIdYesTool execution id (vg_tool_...).
progressPercentageNoCompletion progress 0-100 (present after polling).

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv2.2.1

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover the safety profile (not read-only, not destructive, closed-world), so the bar is lower. The description adds genuinely non-structured behavior: expected latency of 1-3 minutes with HIGH/MAX running longer, and the polling expectation. It could go further on failure/retry behavior, but this is solid added context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core action, then optionality, then operational timing. The final 'keep polling calmly' sentence is informal but earns its place as behavioral guidance. No significant padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 12-parameter, all-optional generation tool with a full schema and an output schema, the description supplies the missing non-schema essentials: input modalities and expected latency. Sibling disambiguation is the one notable gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter (quality enum, aspectRatio units, audioFileIds vs spokenDialogue, durationSeconds, suppressBackgroundMusic) is already documented in the schema. The description repeats the quality enum values and names some input types without adding format or interaction detail, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (generate) and resource (video clip) and enumerates the accepted input modalities (text prompt, images, videos, spokenDialogue, reference audio). It does not, however, distinguish itself from the sibling prompt_to_video_clip or storyboard_to_video, which appear to overlap in purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives practical operating guidance (expect 1-3 minutes, keep polling calmly, warn the user up front), which is useful when-context. But it never states when to pick this tool over prompt_to_video_clip, storyboard_to_video, or script_to_video, so routing between siblings is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.