gemini_image_edit
Edit or compose images by providing input images and a text instruction. Use for single edits or combining multiple images.
Instructions
Edit or compose images: provide one or more input images (paths or base64), plus a text instruction. For a SERIES of successive edits to the same image, prefer gemini_interact (multi-turn) — it keeps edit context and avoids re-processing the full image each round; use gemini_image_edit for one-off edits or composing multiple distinct inputs. Gemini over-preserves the input; there is no edit-strength control — for large structural changes, reroll with a different seed or more forceful wording. Local file inputs are confirmed first: a confirmation prompt where the client supports one; otherwise the first call returns a preview and a confirmToken, and only a repeat call with that token proceeds (see MCP_CONFIRM_MODE).
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | Seed for reproducible generation; random if omitted | |
| async | No | Return a job_id immediately instead of the result, so a long generation cannot hit the host tools/call timeout (-32001); poll gemini_get_result. On the hosted connector prefer max_wait_ms — the executor only lives while a request is open, so async is served there as a bounded wait. | |
| model | No | Model id override (default: server default; see gemini_list_models). gemini-3.1-flash-image (Nano Banana 2, the default) is the generalist: fast, 4K, reliable text, strong multi-reference consistency. gemini-3-pro-image (Pro) is for the hardest work — best world knowledge, brand precision. gemini-3.1-flash-lite-image (Lite) is cheapest: 1K only, no search grounding. | |
| style | No | Name of a saved style preset (gemini_list_styles): its prompt fragment, and reference image if it has one, are applied automatically. Hosted connector only. | |
| images | No | Paths to input image file(s) (1 = edit, 2+ = compose) | |
| inline | No | Return the image as an inline image block you can SEE, instead of a path (stdio) or link (hosted). The default costs nothing to carry and hands back a reference you can reuse; use this when you need to check the result yourself. On the hosted connector gemini_view_media does the same for an image you already have | |
| prompt | Yes | Instruction describing the edit or composition | |
| filename | No | Base filename for the output image (extension stripped; default: slugified prompt) | |
| characters | No | Names of saved characters (gemini_list_characters): each one's reference image and description are attached automatically, keeping recurring subjects consistent without re-sending anything. Hosted connector only. | |
| image_size | No | Output resolution (512 = 0.5K, Flash-only) | |
| images_url | No | Input images as public https URLs — the server downloads them, so no image bytes cross the conversation. Preferred over images_base64, which costs ~14k tokens per photo. Max 15MB each, Content-Type image/*. | |
| output_dir | No | Directory to write images to (default: $GEMINI_OUTPUT_DIR or cwd) | |
| timeout_ms | No | Upstream timeout in ms for this call (default $GEMINI_TIMEOUT_MS, else 60000 — 120000 at 4K, which runs past 60s) | |
| max_wait_ms | No | Wait up to this many ms in-band, then hand back { job_id, status: "running" } to poll with gemini_get_result (e.g. 20000 for multi-image sets). Keeps fast results inline and slow ones off the host timeout (-32001). Ignored when async is set. | |
| orientation | No | Output shape in plain terms: landscape (16:9), portrait (9:16) or square (1:1). For any other proportion — 3:2, 4:3, 4:5, 21:9 — use aspect_ratio, which wins if both are given. | |
| aspect_ratio | No | Exact output aspect ratio. For a plain landscape/portrait/square request, `orientation` is the shorthand; this wins if both are given. | |
| confirmToken | No | ONLY for the two-step confirmation fallback (a client without MCP elicitation). The confirmToken from this same tool's phase-1 "confirmation-required" response, passed back ONLY after the user has seen that preview and explicitly approved it in chat — never on the first call, never invented, never reused. Call again with the same arguments. Ignored when the client supports elicitation. | |
| google_search | No | Ground the image in live Google Search results (current events, weather, data) | |
| images_base64 | No | Input images as base64 strings or data URIs. Last resort — about 14k tokens per photo; prefer images_url or images_file_uris. The server uploads each one and reports a file_uri under image_inputs: pass that to images_file_uris next time instead of re-sending the bytes. | |
| from_clipboard | No | Use the image currently on the macOS system clipboard as an input (downscaled to JPEG) | |
| images_r2_keys | No | Input images by r2_key from this connector's store: a signed upload (gemini_get_upload_url → curl PUT) or an earlier generation's media[].r2_key. The server reads its own bucket — no bytes in the conversation, no ~48h expiry. Hosted connector only. | |
| thinking_level | No | Reasoning depth (Gemini 3 models); higher can help complex/structural edits | |
| idempotency_key | No | Repeat calls with this key return the recorded result (reused: true) instead of billing a new generation. Set it when retrying after a host timeout (-32001). | |
| images_file_uris | No | Input images as Files API references ("files/<id>" or the full uri) from gemini_upload_file. Upload once and reuse across calls with no bytes in the conversation; retained ~48h, after which the reference stops resolving. |