Skip to main content
Glama

gemini_image_edit

Edit or compose images by providing input images and a text instruction. Use for single edits or combining multiple images.

Instructions

Edit or compose images: provide one or more input images (paths or base64), plus a text instruction. For a SERIES of successive edits to the same image, prefer gemini_interact (multi-turn) — it keeps edit context and avoids re-processing the full image each round; use gemini_image_edit for one-off edits or composing multiple distinct inputs. Gemini over-preserves the input; there is no edit-strength control — for large structural changes, reroll with a different seed or more forceful wording. Local file inputs are confirmed first: a confirmation prompt where the client supports one; otherwise the first call returns a preview and a confirmToken, and only a repeat call with that token proceeds (see MCP_CONFIRM_MODE).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
seedNoSeed for reproducible generation; random if omitted
asyncNoReturn a job_id immediately instead of the result, so a long generation cannot hit the host tools/call timeout (-32001); poll gemini_get_result. On the hosted connector prefer max_wait_ms — the executor only lives while a request is open, so async is served there as a bounded wait.
modelNoModel id override (default: server default; see gemini_list_models). gemini-3.1-flash-image (Nano Banana 2, the default) is the generalist: fast, 4K, reliable text, strong multi-reference consistency. gemini-3-pro-image (Pro) is for the hardest work — best world knowledge, brand precision. gemini-3.1-flash-lite-image (Lite) is cheapest: 1K only, no search grounding.
styleNoName of a saved style preset (gemini_list_styles): its prompt fragment, and reference image if it has one, are applied automatically. Hosted connector only.
imagesNoPaths to input image file(s) (1 = edit, 2+ = compose)
inlineNoReturn the image as an inline image block you can SEE, instead of a path (stdio) or link (hosted). The default costs nothing to carry and hands back a reference you can reuse; use this when you need to check the result yourself. On the hosted connector gemini_view_media does the same for an image you already have
promptYesInstruction describing the edit or composition
filenameNoBase filename for the output image (extension stripped; default: slugified prompt)
charactersNoNames of saved characters (gemini_list_characters): each one's reference image and description are attached automatically, keeping recurring subjects consistent without re-sending anything. Hosted connector only.
image_sizeNoOutput resolution (512 = 0.5K, Flash-only)
images_urlNoInput images as public https URLs — the server downloads them, so no image bytes cross the conversation. Preferred over images_base64, which costs ~14k tokens per photo. Max 15MB each, Content-Type image/*.
output_dirNoDirectory to write images to (default: $GEMINI_OUTPUT_DIR or cwd)
timeout_msNoUpstream timeout in ms for this call (default $GEMINI_TIMEOUT_MS, else 60000 — 120000 at 4K, which runs past 60s)
max_wait_msNoWait up to this many ms in-band, then hand back { job_id, status: "running" } to poll with gemini_get_result (e.g. 20000 for multi-image sets). Keeps fast results inline and slow ones off the host timeout (-32001). Ignored when async is set.
orientationNoOutput shape in plain terms: landscape (16:9), portrait (9:16) or square (1:1). For any other proportion — 3:2, 4:3, 4:5, 21:9 — use aspect_ratio, which wins if both are given.
aspect_ratioNoExact output aspect ratio. For a plain landscape/portrait/square request, `orientation` is the shorthand; this wins if both are given.
confirmTokenNoONLY for the two-step confirmation fallback (a client without MCP elicitation). The confirmToken from this same tool's phase-1 "confirmation-required" response, passed back ONLY after the user has seen that preview and explicitly approved it in chat — never on the first call, never invented, never reused. Call again with the same arguments. Ignored when the client supports elicitation.
google_searchNoGround the image in live Google Search results (current events, weather, data)
images_base64NoInput images as base64 strings or data URIs. Last resort — about 14k tokens per photo; prefer images_url or images_file_uris. The server uploads each one and reports a file_uri under image_inputs: pass that to images_file_uris next time instead of re-sending the bytes.
from_clipboardNoUse the image currently on the macOS system clipboard as an input (downscaled to JPEG)
images_r2_keysNoInput images by r2_key from this connector's store: a signed upload (gemini_get_upload_url → curl PUT) or an earlier generation's media[].r2_key. The server reads its own bucket — no bytes in the conversation, no ~48h expiry. Hosted connector only.
thinking_levelNoReasoning depth (Gemini 3 models); higher can help complex/structural edits
idempotency_keyNoRepeat calls with this key return the recorded result (reused: true) instead of billing a new generation. Set it when retrying after a host timeout (-32001).
images_file_urisNoInput images as Files API references ("files/<id>" or the full uri) from gemini_upload_file. Upload once and reuse across calls with no bytes in the conversation; retained ~48h, after which the reference stops resolving.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed2 schema fields changedv2.3.0
    • removedInput schema / properties / confirm
      Removed value: -{
      -  "description": "Must be true to proceed. Without this, the tool returns a preview.",
      -  "type": "boolean"
      -}
    • addedInput schema / properties / confirmToken
      Added value: +{
      +  "description": "ONLY for the two-step confirmation fallback (a client without MCP elicitation). The confirmToken from this same tool's phase-1 \"confirmation-required\" response, passed back ONLY after the user has seen that preview and explicitly approved it in chat — never on the first call, never invented, never reused. Call again with the same arguments. Ignored when the client supports elicitation.",
      +  "type": "string"
      +}
  2. Changed1 schema field changedv2.0.0
    • changedInput schema / $schema
      Previous value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
  3. Changed13 schema fields changedv1.14.0
    • changedInput schema / properties / async / description
      Previous value: -"Run in the background and return a job_id immediately instead of the image, so a long (Pro/4K) generation cannot hit the host tools/call timeout (-32001). Poll gemini_get_result with the job_id to fetch the result. PREFER `max_wait_ms` on the hosted connector: it runs where the executor is only guaranteed to stay alive while the request is open, so this option is served there as a bounded wait rather than an immediate hand-off."New value: +"Return a job_id immediately instead of the result, so a long generation cannot hit the host tools/call timeout (-32001); poll gemini_get_result. On the hosted connector prefer max_wait_ms — the executor only lives while a request is open, so async is served there as a bounded wait."
    • changedInput schema / properties / characters / description
      Previous value: -"Names of saved characters (see gemini_list_characters / gemini_save_character): each one's reference image and description are attached to the request automatically, keeping recurring subjects consistent without re-sending anything. Hosted connector only."New value: +"Names of saved characters (gemini_list_characters): each one's reference image and description are attached automatically, keeping recurring subjects consistent without re-sending anything. Hosted connector only."
    • changedInput schema / properties / idempotency_key / description
      Previous value: -"Opaque idempotency key: a repeat call with the same key returns the recorded result (reused: true) instead of billing a new generation. Set it when retrying after a host timeout (-32001) to avoid a duplicate charge."New value: +"Repeat calls with this key return the recorded result (reused: true) instead of billing a new generation. Set it when retrying after a host timeout (-32001)."
    • changedInput schema / properties / images_base64 / description
      Previous value: -"Input images as base64 strings or data URIs. Last resort: prefer images_url or images_file_uris, which keep image bytes out of the conversation"New value: +"Input images as base64 strings or data URIs. Last resort — about 14k tokens per photo; prefer images_url or images_file_uris. The server uploads each one and reports a file_uri under image_inputs: pass that to images_file_uris next time instead of re-sending the bytes."
    • changedInput schema / properties / images_file_uris / description
      Previous value: -"Input images by Gemini Files API reference (\"files/<id>\", or the full uri) from gemini_upload_file or POST /upload. Upload once, then reference it across as many calls as you like — no bytes are re-sent and none enter the conversation. Files are retained ~48h, after which the reference stops resolving."New value: +"Input images as Files API references (\"files/<id>\" or the full uri) from gemini_upload_file. Upload once and reuse across calls with no bytes in the conversation; retained ~48h, after which the reference stops resolving."
    • changedInput schema / properties / images_r2_keys / description
      Previous value: -"Input images by r2_key from THIS connector's store: a signed upload (gemini_get_upload_url → curl PUT) or an earlier generation's media[].r2_key. The server reads its own bucket directly — no bytes in the conversation, no signed URL, no ~48h Files API expiry. Hosted connector only."New value: +"Input images by r2_key from this connector's store: a signed upload (gemini_get_upload_url → curl PUT) or an earlier generation's media[].r2_key. The server reads its own bucket — no bytes in the conversation, no ~48h expiry. Hosted connector only."
    • changedInput schema / properties / images_url / description
      Previous value: -"Input images as public https URLs — the SERVER downloads them, so no image bytes travel through the conversation. Preferred over images_base64, which costs ~14k tokens per photo and breaks if a file read was truncated. Max 15MB each; must be a directly-linked image (Content-Type image/*)."New value: +"Input images as public https URLs — the server downloads them, so no image bytes cross the conversation. Preferred over images_base64, which costs ~14k tokens per photo. Max 15MB each, Content-Type image/*."
    • changedInput schema / properties / inline / description
      Previous value: -"Return base64 images inline instead of writing to disk"New value: +"Return the image as an inline image block you can SEE, instead of a path (stdio) or link (hosted). The default costs nothing to carry and hands back a reference you can reuse; use this when you need to check the result yourself. On the hosted connector gemini_view_media does the same for an image you already have"
    • changedInput schema / properties / max_wait_ms / description
      Previous value: -"Wait up to this many ms for the result; if generation is still running when the budget expires, return { job_id, status: \"running\" } immediately instead (poll gemini_get_result). Keeps fast results in-band while a slow batch can never trip the host tools/call timeout (-32001) — e.g. 20000 for multi-image sets. Ignored when async is set."New value: +"Wait up to this many ms in-band, then hand back { job_id, status: \"running\" } to poll with gemini_get_result (e.g. 20000 for multi-image sets). Keeps fast results inline and slow ones off the host timeout (-32001). Ignored when async is set."
    • changedInput schema / properties / model / description
      Previous value: -"Model id override (default: server default; see gemini_list_models). gemini-3.1-flash-image (Nano Banana 2) is the versatile generalist workhorse — balances speed with state-of-the-art 4K generation, world knowledge, and reliable text rendering; excels at multi-reference-image processing and consistency. gemini-3-pro-image (Nano Banana Pro) is the premium choice for the most complex visual tasks — highest world knowledge, advanced localization, accurate brand consistency, precision creative control. gemini-3.1-flash-lite-image (Nano Banana 2 Lite) is the fastest/cheapest for simple tasks (1K only, no search grounding)."New value: +"Model id override (default: server default; see gemini_list_models). gemini-3.1-flash-image (Nano Banana 2, the default) is the generalist: fast, 4K, reliable text, strong multi-reference consistency. gemini-3-pro-image (Pro) is for the hardest work — best world knowledge, brand precision. gemini-3.1-flash-lite-image (Lite) is cheapest: 1K only, no search grounding."
    • changedInput schema / properties / orientation / description
      Previous value: -"Shape of the output, in plain terms: \"landscape\" (wide, 16:9), \"portrait\" (tall, 9:16) or \"square\" (1:1). Use this for a request phrased as landscape/portrait/vertical/horizontal. For any other proportion — 35mm photo (3:2), print (4:3), social (4:5), cinematic (21:9) — name it with aspect_ratio instead, which overrides this when both are given."New value: +"Output shape in plain terms: landscape (16:9), portrait (9:16) or square (1:1). For any other proportion — 3:2, 4:3, 4:5, 21:9 — use aspect_ratio, which wins if both are given."
    • changedInput schema / properties / style / description
      Previous value: -"Name of a saved style preset (see gemini_list_styles / gemini_save_style): its prompt fragment — and reference image, if it has one — is applied to the request automatically. Hosted connector only."New value: +"Name of a saved style preset (gemini_list_styles): its prompt fragment, and reference image if it has one, are applied automatically. Hosted connector only."
    • changedInput schema / properties / timeout_ms / description
      Previous value: -"Upstream request timeout in ms for this call (default: $GEMINI_TIMEOUT_MS, else 60000 — or 120000 when image_size is 4K, which routinely runs past 60s)"New value: +"Upstream timeout in ms for this call (default $GEMINI_TIMEOUT_MS, else 60000 — 120000 at 4K, which runs past 60s)"
  4. Changed2 schema fields changedv1.10.0
    • changedInput schema / properties / aspect_ratio / description
      Previous value: -"Output aspect ratio"New value: +"Exact output aspect ratio. For a plain landscape/portrait/square request, `orientation` is the shorthand; this wins if both are given."
    • addedInput schema / properties / orientation
      Added value: +{
      +  "description": "Shape of the output, in plain terms: \"landscape\" (wide, 16:9), \"portrait\" (tall, 9:16) or \"square\" (1:1). Use this for a request phrased as landscape/portrait/vertical/horizontal. For any other proportion — 35mm photo (3:2), print (4:3), social (4:5), cinematic (21:9) — name it with aspect_ratio instead, which overrides this when both are given.",
      +  "enum": [
      +    "landscape",
      +    "portrait",
      +    "square"
      +  ],
      +  "type": "string"
      +}
  5. Changed5 schema fields changedv1.7.0
    • changedInput schema / properties / async / description
      Previous value: -"Run in the background and return a job_id immediately instead of the image, so a long (Pro/4K) generation cannot hit the host tools/call timeout (-32001). Poll gemini_get_result with the job_id to fetch the result (jobs are per-process and expire ~10 min after completion)."New value: +"Run in the background and return a job_id immediately instead of the image, so a long (Pro/4K) generation cannot hit the host tools/call timeout (-32001). Poll gemini_get_result with the job_id to fetch the result. PREFER `max_wait_ms` on the hosted connector: it runs where the executor is only guaranteed to stay alive while the request is open, so this option is served there as a bounded wait rather than an immediate hand-off."
    • addedInput schema / properties / characters
      Added value: +{
      +  "description": "Names of saved characters (see gemini_list_characters / gemini_save_character): each one's reference image and description are attached to the request automatically, keeping recurring subjects consistent without re-sending anything. Hosted connector only.",
      +  "items": {
      +    "minLength": 1,
      +    "type": "string"
      +  },
      +  "maxItems": 8,
      +  "type": "array"
      +}
    • addedInput schema / properties / images_r2_keys
      Added value: +{
      +  "description": "Input images by r2_key from THIS connector's store: a signed upload (gemini_get_upload_url → curl PUT) or an earlier generation's media[].r2_key. The server reads its own bucket directly — no bytes in the conversation, no signed URL, no ~48h Files API expiry. Hosted connector only.",
      +  "items": {
      +    "minLength": 1,
      +    "type": "string"
      +  },
      +  "type": "array"
      +}
    • addedInput schema / properties / max_wait_ms
      Added value: +{
      +  "description": "Wait up to this many ms for the result; if generation is still running when the budget expires, return { job_id, status: \"running\" } immediately instead (poll gemini_get_result). Keeps fast results in-band while a slow batch can never trip the host tools/call timeout (-32001) — e.g. 20000 for multi-image sets. Ignored when async is set.",
      +  "exclusiveMinimum": 0,
      +  "maximum": 600000,
      +  "type": "integer"
      +}
    • addedInput schema / properties / style
      Added value: +{
      +  "description": "Name of a saved style preset (see gemini_list_styles / gemini_save_style): its prompt fragment — and reference image, if it has one — is applied to the request automatically. Hosted connector only.",
      +  "minLength": 1,
      +  "type": "string"
      +}
  6. First observedv1.2.0

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the sparse annotations (readOnlyHint: false, openWorldHint: true), the description discloses important behavioral traits: Gemini over-preserves the input, there is no edit-strength control, and local file inputs require a confirmation flow with confirmToken. This adds meaningful context the annotations do not provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but efficiently organized: purpose first, then alternative routing, then behavior caveats, then confirmation mechanics. Every sentence earns its place, and it stays readable despite the tool's 24-parameter complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex tool with no output schema, the description covers the key operational decisions an agent needs: when to use it versus gemini_interact, how to handle confirmation, what to expect regarding edit strength, and how to retry. The rich per-parameter schema fills in the remaining invocation details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds genuine semantics beyond the schema by explaining how to use seed and prompt wording to compensate for the lack of edit-strength control. It also clarifies the confirmation-token flow, which is central to correctly using confirmToken.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource: 'Edit or compose images' and clearly states the core invocation pattern (input images plus a text instruction). It also distinguishes itself from gemini_interact by naming the sibling and the condition that selects it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says when to prefer gemini_interact for a series of edits and when to use this tool for one-off edits or composing multiple distinct inputs. It also gives practical guidance for large structural changes (reroll with a different seed or more forceful wording), which helps an agent use the tool effectively.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.