Skip to main content
Glama

gemini_video_generate

Generate short MP4 videos from text prompts, reference images, or two still images, and edit or extend previous clips.

Instructions

Generate a short video (~10s) via the Gemini omni model: text→video, image→video / reference→video (supply reference image[s]), interpolate between two stills (pass first frame then last frame as images), or continue a prior video (task: "edit" or "extend" + previous_interaction_id / continue_last; extensions add ~3-10s each, to ~40s total). Cost scales with resolution — draft at 360p, keep at 1080p/4k. Output is written to disk as MP4 (video has no inline MCP block). Video runs long — give it a max_wait_ms budget (or async: true + gemini_get_result on a local install), or raise timeout_ms. Needs a funded account. Local file inputs are confirmed first: a confirmation prompt where the client supports one; otherwise the first call returns a preview and a confirmToken, and only a repeat call with that token proceeds (see MCP_CONFIRM_MODE).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
taskNotext_to_video (default), image_to_video / reference_to_video (need image input), or edit / extend (need previous_interaction_id)
asyncNoReturn a job_id immediately instead of the result, so a long generation cannot hit the host tools/call timeout (-32001); poll gemini_get_result. On the hosted connector prefer max_wait_ms — the executor only lives while a request is open, so async is served there as a bounded wait.
modelNoModel id override (default: gemini-omni-1.1-flash)
imagesNoReference image path(s) for image_to_video / reference_to_video
promptYesDescription of the video to generate (or the edit instruction when task=edit)
deliveryNoHow the clip comes back: "uri" (default — a Files API link the server downloads, no size ceiling) or "inline" (base64, capped ~4MB)
filenameNoBase filename for the output video (extension stripped; default: slugified prompt)
backgroundNoRun the generation on Google's side and poll it, so a killed job can be recovered by gemini_get_result. Trade-off: retrieving a backgrounded interaction is unreliable today (some become permanently unreadable), so the default is off
images_urlNoReference stills as public https URLs — the server downloads them, so no image bytes cross the conversation. Preferred over images_base64, which costs ~14k tokens per photo. Max 15MB each, Content-Type image/*.
output_dirNoDirectory to write the video to (default: $GEMINI_OUTPUT_DIR or cwd)
resolutionNoOutput resolution (default 720p). Video is billed per output token, so 360p costs roughly a third of 720p — use it for drafts
timeout_msNoUpstream timeout in ms for this call (default $GEMINI_TIMEOUT_MS, else 60000 — 120000 at 4K, which runs past 60s)
max_wait_msNoWait up to this many ms in-band, then hand back { job_id, status: "running" } to poll with gemini_get_result (e.g. 20000 for multi-image sets). Keeps fast results inline and slow ones off the host timeout (-32001). Ignored when async is set.
orientationNoOutput shape in plain terms: landscape (16:9), portrait (9:16) or square (1:1). For any other proportion — 3:2, 4:3, 4:5, 21:9 — use aspect_ratio, which wins if both are given.
aspect_ratioNoExact output aspect ratio (omni: 16:9 or 9:16). `orientation` is the plain-language shorthand; this wins if both are given.
confirmTokenNoONLY for the two-step confirmation fallback (a client without MCP elicitation). The confirmToken from this same tool's phase-1 "confirmation-required" response, passed back ONLY after the user has seen that preview and explicitly approved it in chat — never on the first call, never invented, never reused. Call again with the same arguments. Ignored when the client supports elicitation.
continue_lastNoContinue from the most recent video interaction this server created (explicit previous_interaction_id wins)
images_base64NoReference images as base64 strings or data URIs. Last resort — about 14k tokens per photo; prefer images_url or images_file_uris. The server uploads each one and reports a file_uri under image_inputs: pass that to images_file_uris next time instead of re-sending the bytes.
from_clipboardNoUse the image currently on the macOS clipboard as a reference
idempotency_keyNoRepeat calls with this key return the recorded result (reused: true) instead of billing a new generation. Set it when retrying after a host timeout (-32001).
images_file_urisNoReference stills as Files API references ("files/<id>" or the full uri) from gemini_upload_file. Upload once and reuse across calls with no bytes in the conversation; retained ~48h, after which the reference stops resolving.
previous_interaction_idNoInteraction id to edit/continue (with task: "edit")

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed2 schema fields changedv2.3.0
    • removedInput schema / properties / confirm
      Removed value: -{
      -  "description": "Must be true to proceed. Without this, the tool returns a preview.",
      -  "type": "boolean"
      -}
    • addedInput schema / properties / confirmToken
      Added value: +{
      +  "description": "ONLY for the two-step confirmation fallback (a client without MCP elicitation). The confirmToken from this same tool's phase-1 \"confirmation-required\" response, passed back ONLY after the user has seen that preview and explicitly approved it in chat — never on the first call, never invented, never reused. Call again with the same arguments. Ignored when the client supports elicitation.",
      +  "type": "string"
      +}
  2. Changed1 schema field changedv2.0.0
    • changedInput schema / $schema
      Previous value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
  3. Changed12 schema fields changedv1.14.0
    • changedInput schema / properties / async / description
      Previous value: -"Run in the background and return a job_id immediately instead of the image, so a long (Pro/4K) generation cannot hit the host tools/call timeout (-32001). Poll gemini_get_result with the job_id to fetch the result. PREFER `max_wait_ms` on the hosted connector: it runs where the executor is only guaranteed to stay alive while the request is open, so this option is served there as a bounded wait rather than an immediate hand-off."New value: +"Return a job_id immediately instead of the result, so a long generation cannot hit the host tools/call timeout (-32001); poll gemini_get_result. On the hosted connector prefer max_wait_ms — the executor only lives while a request is open, so async is served there as a bounded wait."
    • changedInput schema / properties / idempotency_key / description
      Previous value: -"Opaque idempotency key: a repeat call with the same key returns the recorded result (reused: true) instead of billing a new generation. Set it when retrying after a host timeout (-32001) to avoid a duplicate charge."New value: +"Repeat calls with this key return the recorded result (reused: true) instead of billing a new generation. Set it when retrying after a host timeout (-32001)."
    • changedInput schema / properties / images_base64 / description
      Previous value: -"Reference images as base64 strings or data URIs. Last resort: prefer images_url or images_file_uris, which keep image bytes out of the conversation"New value: +"Reference images as base64 strings or data URIs. Last resort — about 14k tokens per photo; prefer images_url or images_file_uris. The server uploads each one and reports a file_uri under image_inputs: pass that to images_file_uris next time instead of re-sending the bytes."
    • changedInput schema / properties / images_file_uris / description
      Previous value: -"Reference stills by Gemini Files API reference (\"files/<id>\", or the full uri) from gemini_upload_file or POST /upload. Upload once, then reference it across as many calls as you like — no bytes are re-sent and none enter the conversation. Files are retained ~48h, after which the reference stops resolving."New value: +"Reference stills as Files API references (\"files/<id>\" or the full uri) from gemini_upload_file. Upload once and reuse across calls with no bytes in the conversation; retained ~48h, after which the reference stops resolving."
    • changedInput schema / properties / images_url / description
      Previous value: -"Reference stills as public https URLs — the SERVER downloads them, so no image bytes travel through the conversation. Preferred over images_base64, which costs ~14k tokens per photo and breaks if a file read was truncated. Max 15MB each; must be a directly-linked image (Content-Type image/*)."New value: +"Reference stills as public https URLs — the server downloads them, so no image bytes cross the conversation. Preferred over images_base64, which costs ~14k tokens per photo. Max 15MB each, Content-Type image/*."
    • changedInput schema / properties / max_wait_ms / description
      Previous value: -"Wait up to this many ms for the result; if generation is still running when the budget expires, return { job_id, status: \"running\" } immediately instead (poll gemini_get_result). Keeps fast results in-band while a slow batch can never trip the host tools/call timeout (-32001) — e.g. 20000 for multi-image sets. Ignored when async is set."New value: +"Wait up to this many ms in-band, then hand back { job_id, status: \"running\" } to poll with gemini_get_result (e.g. 20000 for multi-image sets). Keeps fast results inline and slow ones off the host timeout (-32001). Ignored when async is set."
    • changedInput schema / properties / model / description
      Previous value: -"Model id override (default: gemini-omni-flash-preview)"New value: +"Model id override (default: gemini-omni-1.1-flash)"
    • changedInput schema / properties / orientation / description
      Previous value: -"Shape of the output, in plain terms: \"landscape\" (wide, 16:9), \"portrait\" (tall, 9:16) or \"square\" (1:1). Use this for a request phrased as landscape/portrait/vertical/horizontal. For any other proportion — 35mm photo (3:2), print (4:3), social (4:5), cinematic (21:9) — name it with aspect_ratio instead, which overrides this when both are given."New value: +"Output shape in plain terms: landscape (16:9), portrait (9:16) or square (1:1). For any other proportion — 3:2, 4:3, 4:5, 21:9 — use aspect_ratio, which wins if both are given."
    • addedInput schema / properties / resolution
      Added value: +{
      +  "description": "Output resolution (default 720p). Video is billed per output token, so 360p costs roughly a third of 720p — use it for drafts",
      +  "enum": [
      +    "360p",
      +    "720p",
      +    "1080p",
      +    "4k"
      +  ],
      +  "type": "string"
      +}
    • changedInput schema / properties / task / description
      Previous value: -"text_to_video (default), image_to_video / reference_to_video (need image input), or edit (needs previous_interaction_id)"New value: +"text_to_video (default), image_to_video / reference_to_video (need image input), or edit / extend (need previous_interaction_id)"
    • changedInput schema / properties / task / enum
      Previous value: -[
      -  "text_to_video",
      -  "image_to_video",
      -  "reference_to_video",
      -  "edit"
      -]New value: +[
      +  "text_to_video",
      +  "image_to_video",
      +  "reference_to_video",
      +  "edit",
      +  "extend"
      +]
    • changedInput schema / properties / timeout_ms / description
      Previous value: -"Upstream request timeout in ms for this call (default: $GEMINI_TIMEOUT_MS, else 60000 — or 120000 when image_size is 4K, which routinely runs past 60s)"New value: +"Upstream timeout in ms for this call (default $GEMINI_TIMEOUT_MS, else 60000 — 120000 at 4K, which runs past 60s)"
  4. Changed4 schema fields changedv1.10.0
    • changedInput schema / properties / aspect_ratio / description
      Previous value: -"Output aspect ratio (omni: 16:9 or 9:16)"New value: +"Exact output aspect ratio (omni: 16:9 or 9:16). `orientation` is the plain-language shorthand; this wins if both are given."
    • addedInput schema / properties / background
      Added value: +{
      +  "description": "Run the generation on Google's side and poll it, so a killed job can be recovered by gemini_get_result. Trade-off: retrieving a backgrounded interaction is unreliable today (some become permanently unreadable), so the default is off",
      +  "type": "boolean"
      +}
    • addedInput schema / properties / delivery
      Added value: +{
      +  "description": "How the clip comes back: \"uri\" (default — a Files API link the server downloads, no size ceiling) or \"inline\" (base64, capped ~4MB)",
      +  "enum": [
      +    "inline",
      +    "uri"
      +  ],
      +  "type": "string"
      +}
    • addedInput schema / properties / orientation
      Added value: +{
      +  "description": "Shape of the output, in plain terms: \"landscape\" (wide, 16:9), \"portrait\" (tall, 9:16) or \"square\" (1:1). Use this for a request phrased as landscape/portrait/vertical/horizontal. For any other proportion — 35mm photo (3:2), print (4:3), social (4:5), cinematic (21:9) — name it with aspect_ratio instead, which overrides this when both are given.",
      +  "enum": [
      +    "landscape",
      +    "portrait",
      +    "square"
      +  ],
      +  "type": "string"
      +}
  5. Changed2 schema fields changedv1.7.0
    • changedInput schema / properties / async / description
      Previous value: -"Run in the background and return a job_id immediately instead of the image, so a long (Pro/4K) generation cannot hit the host tools/call timeout (-32001). Poll gemini_get_result with the job_id to fetch the result (jobs are per-process and expire ~10 min after completion)."New value: +"Run in the background and return a job_id immediately instead of the image, so a long (Pro/4K) generation cannot hit the host tools/call timeout (-32001). Poll gemini_get_result with the job_id to fetch the result. PREFER `max_wait_ms` on the hosted connector: it runs where the executor is only guaranteed to stay alive while the request is open, so this option is served there as a bounded wait rather than an immediate hand-off."
    • addedInput schema / properties / max_wait_ms
      Added value: +{
      +  "description": "Wait up to this many ms for the result; if generation is still running when the budget expires, return { job_id, status: \"running\" } immediately instead (poll gemini_get_result). Keeps fast results in-band while a slow batch can never trip the host tools/call timeout (-32001) — e.g. 20000 for multi-image sets. Ignored when async is set.",
      +  "exclusiveMinimum": 0,
      +  "maximum": 600000,
      +  "type": "integer"
      +}
  6. First observedv1.2.0

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are minimal (readOnlyHint: false, openWorldHint: true), so the description carries the full behavioral burden — and it does impressively: cost scaling with resolution ('draft at 360p'), disk side effects ('Output is written to disk as MP4 (video has no inline MCP block)'), long-run timing behavior, the funded-account prerequisite, and the two-step confirmation/confirmToken fallback are all disclosed. The write-to-disk behavior is consistent with readOnlyHint: false, so there is no annotation contradiction.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

At roughly 200 words for a 22-parameter, 5-mode tool, the description is dense but every sentence carries load-bearing information, and the core function is front-loaded before any caveats. The progression (function → modes → cost → output → timing → access → confirmation) keeps the length navigable rather than rambling. Minor redundancy exists with the schema's per-parameter notes, but the prose is tighter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a high-complexity tool with no output schema, the description covers the critical runtime surfaces: each mode's input combinations, disk output, duration scaling, the async/job_id path, the max_wait_ms hand-back shape ({ job_id, status: "running" }), and the confirmToken flow. The remaining gap is the exact success-return shape of the normal synchronous path (what the MP4 reference looks like), which is only alluded to. Near-complete given the complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds genuinely non-schema semantics: interpolation requires passing first frame then last frame as images, extensions add ~3-10s each up to ~40s total, and each task value maps to a specific parameter set. The schema already covers most per-parameter semantics (transport trade-offs, defaults), so the added value is moderate rather than extensive.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource+scope ('Generate a short video (~10s) via the Gemini omni model') and immediately enumerates all five modes (text→video, image→video / reference→video, interpolation, edit/extend), making the tool's function unambiguous. This clearly differentiates it from siblings such as gemini_image_generate and gemini_music_generate, which produce other media types.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit mode routing with the parameters each mode requires ('supply reference image[s]', 'pass first frame then last frame as images', 'task: "edit" or "extend" + previous_interaction_id / continue_last'), and names the polling companion gemini_get_result plus the MCP_CONFIRM_MODE fallback. It also tells the agent how to handle long-running calls (max_wait_ms, async, timeout_ms). It lacks explicit when-not-to-use statements against sibling generators (e.g., when to prefer gemini_image_generate instead), but internal branching guidance is concrete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.