gemini_interact
Create and refine images through multi-turn interactions; pass the returned interaction ID to continue editing the same image iteratively.
Instructions
Preferred tool for iterative refinement of a single image — multi-turn generation and editing via the Interactions API. To refine, pass the returned interaction id as previous_interaction_id on the next call; do NOT start a new interaction or re-send the image for each tweak. continue_last: true chains from this session's most recent interaction without threading the id. If a call times out client-side the generation usually still completes: the image and an <image>.json sidecar holding its interaction id land in the output dir, and continue_last: true still resumes it — look there before re-issuing, which bills a second generation. A chained 404 is re-anchored on the prior output and re-issued un-chained (chain_recovered); a second 404 means the interaction id was not the cause — check the model id / files uri. Output is JPEG. Local file inputs are confirmed first: a confirmation prompt where the client supports one; otherwise the first call returns a preview and a confirmToken, and only a repeat call with that token proceeds (see MCP_CONFIRM_MODE).
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| async | No | Return a job_id immediately instead of the result, so a long generation cannot hit the host tools/call timeout (-32001); poll gemini_get_result. On the hosted connector prefer max_wait_ms — the executor only lives while a request is open, so async is served there as a bounded wait. | |
| input | Yes | Text prompt or editing instruction | |
| model | No | Model id override (default: server default; see gemini_list_models). gemini-3.1-flash-image (Nano Banana 2, the default) is the generalist: fast, 4K, reliable text, strong multi-reference consistency. gemini-3-pro-image (Pro) is for the hardest work — best world knowledge, brand precision. gemini-3.1-flash-lite-image (Lite) is cheapest: 1K only, no search grounding. | |
| images | No | Paths to reference input images. NEW reference images only (e.g. a style or target photo). When chaining with previous_interaction_id, do NOT re-attach the prior turn's output — the interaction already contains it, and re-sending it anchors the model against your edit. | |
| inline | No | Return base64 images inline instead of writing to disk | |
| filename | No | Base filename for the output image (extension stripped; default: slugified input) | |
| video_url | No | Public YouTube URL (or a previously uploaded Files API uri) as a video reference (video→image; use a Flash model e.g. gemini-3.1-flash-image) | |
| image_size | No | Output resolution (512 = 0.5K, Flash-only) | |
| images_url | No | Reference images. NEW reference images only (e.g. a style or target photo). When chaining with previous_interaction_id, do NOT re-attach the prior turn's output — the interaction already contains it, and re-sending it anchors the model against your edit. Given as public https URLs — the server downloads them, so no image bytes cross the conversation. Preferred over images_base64, which costs ~14k tokens per photo. Max 15MB each, Content-Type image/*. | |
| output_dir | No | Directory to write images to (default: $GEMINI_OUTPUT_DIR or cwd) | |
| timeout_ms | No | Upstream timeout in ms for this call (default $GEMINI_TIMEOUT_MS, else 60000 — 120000 at 4K, which runs past 60s) | |
| video_path | No | Path to a local video file — uploaded to the Gemini Files API (~48h retention, 2 GB max) and used as the video reference. Alternative to video_url. | |
| max_wait_ms | No | Wait up to this many ms in-band, then hand back { job_id, status: "running" } to poll with gemini_get_result (e.g. 20000 for multi-image sets). Keeps fast results inline and slow ones off the host timeout (-32001). Ignored when async is set. | |
| orientation | No | Output shape in plain terms: landscape (16:9), portrait (9:16) or square (1:1). For any other proportion — 3:2, 4:3, 4:5, 21:9 — use aspect_ratio, which wins if both are given. | |
| aspect_ratio | No | Exact output aspect ratio. For a plain landscape/portrait/square request, `orientation` is the shorthand; this wins if both are given. | |
| confirmToken | No | ONLY for the two-step confirmation fallback (a client without MCP elicitation). The confirmToken from this same tool's phase-1 "confirmation-required" response, passed back ONLY after the user has seen that preview and explicitly approved it in chat — never on the first call, never invented, never reused. Call again with the same arguments. Ignored when the client supports elicitation. | |
| search_types | No | Grounding search types (implies google_search). image_search (gemini-3.1-flash-image only) uses Google Image Search results as visual references; per Google ToS the returned grounding.search_suggestions HTML must then be displayed to the user. Cannot depict real people from web images. | |
| continue_last | No | Continue from the most recent interaction this server created (convenience for previous_interaction_id; an explicit id wins). Survives a server restart by falling back to the newest <image>.json sidecar in the output dir. | |
| google_search | No | Ground the image in live Google Search results (current events, weather, data) | |
| images_base64 | No | Reference images. NEW reference images only (e.g. a style or target photo). When chaining with previous_interaction_id, do NOT re-attach the prior turn's output — the interaction already contains it, and re-sending it anchors the model against your edit. Given as base64 strings or data URIs. Last resort — about 14k tokens per photo; prefer images_url or images_file_uris. The server uploads each one and reports a file_uri under image_inputs: pass that to images_file_uris next time instead of re-sending the bytes. | |
| from_clipboard | No | Use the image currently on the macOS system clipboard as an input (downscaled to JPEG) | |
| thinking_level | No | Reasoning depth; higher can help complex/structural edits | |
| idempotency_key | No | Repeat calls with this key return the recorded result (reused: true) instead of billing a new generation. Set it when retrying after a host timeout (-32001). | |
| images_file_uris | No | Reference images. NEW reference images only (e.g. a style or target photo). When chaining with previous_interaction_id, do NOT re-attach the prior turn's output — the interaction already contains it, and re-sending it anchors the model against your edit. Given as Files API references ("files/<id>" or the full uri) from gemini_upload_file. Upload once and reuse across calls with no bytes in the conversation; retained ~48h, after which the reference stops resolving. | |
| previous_interaction_id | No | ID from a prior gemini_interact call — continues that multi-turn conversation |