gemini_video_generate
Generate short MP4 videos from text prompts, reference images, or two still images, and edit or extend previous clips.
Instructions
Generate a short video (~10s) via the Gemini omni model: text→video, image→video / reference→video (supply reference image[s]), interpolate between two stills (pass first frame then last frame as images), or continue a prior video (task: "edit" or "extend" + previous_interaction_id / continue_last; extensions add ~3-10s each, to ~40s total). Cost scales with resolution — draft at 360p, keep at 1080p/4k. Output is written to disk as MP4 (video has no inline MCP block). Video runs long — give it a max_wait_ms budget (or async: true + gemini_get_result on a local install), or raise timeout_ms. Needs a funded account. Local file inputs are confirmed first: a confirmation prompt where the client supports one; otherwise the first call returns a preview and a confirmToken, and only a repeat call with that token proceeds (see MCP_CONFIRM_MODE).
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| task | No | text_to_video (default), image_to_video / reference_to_video (need image input), or edit / extend (need previous_interaction_id) | |
| async | No | Return a job_id immediately instead of the result, so a long generation cannot hit the host tools/call timeout (-32001); poll gemini_get_result. On the hosted connector prefer max_wait_ms — the executor only lives while a request is open, so async is served there as a bounded wait. | |
| model | No | Model id override (default: gemini-omni-1.1-flash) | |
| images | No | Reference image path(s) for image_to_video / reference_to_video | |
| prompt | Yes | Description of the video to generate (or the edit instruction when task=edit) | |
| delivery | No | How the clip comes back: "uri" (default — a Files API link the server downloads, no size ceiling) or "inline" (base64, capped ~4MB) | |
| filename | No | Base filename for the output video (extension stripped; default: slugified prompt) | |
| background | No | Run the generation on Google's side and poll it, so a killed job can be recovered by gemini_get_result. Trade-off: retrieving a backgrounded interaction is unreliable today (some become permanently unreadable), so the default is off | |
| images_url | No | Reference stills as public https URLs — the server downloads them, so no image bytes cross the conversation. Preferred over images_base64, which costs ~14k tokens per photo. Max 15MB each, Content-Type image/*. | |
| output_dir | No | Directory to write the video to (default: $GEMINI_OUTPUT_DIR or cwd) | |
| resolution | No | Output resolution (default 720p). Video is billed per output token, so 360p costs roughly a third of 720p — use it for drafts | |
| timeout_ms | No | Upstream timeout in ms for this call (default $GEMINI_TIMEOUT_MS, else 60000 — 120000 at 4K, which runs past 60s) | |
| max_wait_ms | No | Wait up to this many ms in-band, then hand back { job_id, status: "running" } to poll with gemini_get_result (e.g. 20000 for multi-image sets). Keeps fast results inline and slow ones off the host timeout (-32001). Ignored when async is set. | |
| orientation | No | Output shape in plain terms: landscape (16:9), portrait (9:16) or square (1:1). For any other proportion — 3:2, 4:3, 4:5, 21:9 — use aspect_ratio, which wins if both are given. | |
| aspect_ratio | No | Exact output aspect ratio (omni: 16:9 or 9:16). `orientation` is the plain-language shorthand; this wins if both are given. | |
| confirmToken | No | ONLY for the two-step confirmation fallback (a client without MCP elicitation). The confirmToken from this same tool's phase-1 "confirmation-required" response, passed back ONLY after the user has seen that preview and explicitly approved it in chat — never on the first call, never invented, never reused. Call again with the same arguments. Ignored when the client supports elicitation. | |
| continue_last | No | Continue from the most recent video interaction this server created (explicit previous_interaction_id wins) | |
| images_base64 | No | Reference images as base64 strings or data URIs. Last resort — about 14k tokens per photo; prefer images_url or images_file_uris. The server uploads each one and reports a file_uri under image_inputs: pass that to images_file_uris next time instead of re-sending the bytes. | |
| from_clipboard | No | Use the image currently on the macOS clipboard as a reference | |
| idempotency_key | No | Repeat calls with this key return the recorded result (reused: true) instead of billing a new generation. Set it when retrying after a host timeout (-32001). | |
| images_file_uris | No | Reference stills as Files API references ("files/<id>" or the full uri) from gemini_upload_file. Upload once and reuse across calls with no bytes in the conversation; retained ~48h, after which the reference stops resolving. | |
| previous_interaction_id | No | Interaction id to edit/continue (with task: "edit") |