gemini-mcp
This server exposes a Gemini-powered media generation MCP server with image, video, music, and file-management tools over stdio.
Image generation:
gemini_image_generatecreates images from text prompts with model, size, aspect ratio, orientation, seed, count, search grounding, and reference-image options.Image editing/composition:
gemini_image_editperforms one-off edits or multi-image composition from a text instruction.Consistent image sets:
gemini_image_setgenerates a master image plus consistent scene variations, with master/chain reference modes and style/character support.Multi-turn refinement:
gemini_interactchains iterative image edits via the Interactions API, withprevious_interaction_id,continue_last, and 404 recovery.Video generation:
gemini_video_generatesupports text-to-video, image-to-video, reference-to-video, interpolation, edit, and extend via Gemini omni, with resolution control and MP4 output.Music generation:
gemini_music_generatecreates MP3 tracks from text prompts via Lyria models, including clip, full-song, and pro variants.Async long-running jobs:
gemini_get_resultpolls jobs started withasync: trueor handed off bymax_wait_ms.Token/cost tracking:
gemini_token_usagereports session token usage and estimated USD cost.Files API management:
gemini_upload_file,gemini_list_files, andgemini_delete_fileupload, list, and delete reusable image/video/audio references.Model discovery:
gemini_list_modelslists available Gemini image models and the current default.Health check:
gemini_healthcheckverifies credentials and upstream reachability.Hosted-only extras: signed media URLs, inline media viewing, signed upload URLs, and persistent character/style libraries.
Provides tools for interacting with Google Gemini's media generation API, enabling image generation/editing, video generation, music generation, and file management through the Gemini platform.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@gemini-mcpgenerate an image of a cat wearing a hat"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
gemini-mcp
MCP server for Google Gemini media generation. Exposes thirteen tools to Claude over stdio: list available models, generate/edit/compose images, generate a consistent set of images from a master prompt, multi-turn image refinement (Interactions API), video generation (omni), music generation (Lyria), an async result poll for long generations, and Files API upload/list/delete for reusable image references. Output is written to disk by default (path returned) or returned inline as base64. Built on the Gemini v1beta API (generativelanguage.googleapis.com) using the Nano Banana / Nano Banana Pro (images), omni (video), and Lyria (music) model families.
Developed and maintained by AI (Claude Code).
Environment Variables
Variable | Required | Description |
| Yes | Your Google Gemini API key (aistudio.google.com/apikey) |
| No | Override the default image model (default: |
| No | Default directory for generated images (default: current working directory) |
| No | Directory to resolve bare input-image filenames against (so |
| No | Restrict local files streamed to the Gemini Files API ( |
| No | Upstream request timeout in ms (default: |
| No | Progress-notification cadence in ms while a generation runs (default: |
| No | How long to wait out interactions-store lag when a chained call 404s (default: |
Confirmations
Sending a local file to Google (the images / master_images / video_path inputs, or
gemini_upload_file's path) and every delete ask for confirmation first. A client that can
show a confirmation prompt (Claude Code) gets the real prompt. Elsewhere the first call does
nothing and returns a preview (the resolved local paths with their MIME types and sizes, or the
method/path being deleted) plus a confirmToken; only a repeat call with the same arguments and
that token proceeds. The token is single-use, expires, and is bound to the exact arguments — a
changed prompt or a swapped file between the two calls is refused. Pure text prompts, URLs,
files/ references and base64 inputs are not gated.
variable | default | |
|
| What a write does on a client that cannot show a confirmation prompt (claude.ai, Claude Desktop). |
|
| How long a token stays valid. |
| random per process | Signing key; set it only if tokens must survive a server restart. |
Long generations and client timeouts
4K / Pro-model generations can outrun an MCP host's own tools/call timeout (error -32001).
The server sends notifications/progress heartbeats so hosts that reset their timeout on
progress wait it out. If the host still gives up, the server-side generation usually completes
anyway: the image is written to the output dir, gemini_interact also writes an
<image>.json sidecar recording the interaction_id, and continue_last: true resumes the
interaction the lost response belonged to.
When a chained call 404s
A 404 on a request carrying previous_interaction_id is not proof the chain expired. The
only 404 body observed live is generic — "Requested entity was not found." — and never names
which entity. An unknown or renamed model id, and an expired Files API files/… uri
(~48h TTL), return exactly the same thing. So the server no longer asserts a cause it can't
establish: the upstream text is surfaced verbatim, and gemini_interact runs an experiment to
find out which it was.
Most often the id isn't missing at all — it just isn't visible yet. The interactions store is
eventually consistent, and a freshly created id can 404 while the same id resolves fine minutes
later; heavy turns (4K, Pro, thinking_level: high) are the likeliest to hit it, which is
exactly the turn you most want to chain from. So a chained 404 is retried with exponential
backoff for up to 120s (GEMINI_CHAIN_RETRY_MS) before anything is declared broken. The
404 generates nothing and isn't billed, so the wait costs only time.
After that budget is spent, the tool looks up that id's sidecar, re-attaches the image it produced, and re-issues the request without the chain:
The re-issue succeeds → the chain really was the problem, and you get your image anyway, reported as
chain_recovered: { expired_interaction_id, reanchored_on }. The 404'd attempt generates nothing, so this costs the one generation you'd have paid for re-anchoring manually.The re-issue 404s too → the interaction id was never the cause. You get told exactly that, with the upstream text, and pointed at the model id and any
files/…uri instead of being sent to chase an interaction that was fine all along.No sidecar matches the dead id → the original error, rather than a guess. Re-anchoring on the wrong picture would silently corrupt the edit.
Separately, continue_last no longer dies with the server process: with no in-memory id it
resumes from the newest <image>.json sidecar in the output dir and reports
continued_from_sidecar: true. That case was never an expired chain at all — the interaction
was alive upstream the whole time; only our memory of its id was gone.
For hosts whose timeout can't be tamed (e.g. Claude Desktop, a fixed ~30s cap that ignores progress), two guards make re-issuing safe and unnecessary:
async: truereturns ajob_idimmediately instead of the image, so the call can't time out at all; pollgemini_get_resultwith thejob_iduntil it'sdone.idempotency_keymakes a repeat call idempotent — a retry with the same key returns the recorded result (reused: true) instead of billing a second generation. (Even without a key, two identical in-flight calls are deduplicated automatically.)
Related MCP server: mmxomni
Tools
Tool | Description |
| List available Gemini image models and the current default |
| Generate image(s) from a text prompt |
| One-off edits or multi-image composition with a text instruction (for a series of edits, use |
| Generate a master image plus N consistent images referencing it |
| Preferred tool for iterative refinement: multi-turn generation/editing via the Interactions API — chain the returned |
| Generate a short video (text→video, image→video, two-still interpolation, |
| Generate music from a text prompt via a Lyria model — |
| Fetch an async generation started with |
| Token usage and an estimated USD cost for this session so far. Call it before and after a workflow and subtract to attribute that workflow's spend. Priced per call against each call's own model from a dated rate card ( |
| Upload an image (or video/audio) to the Gemini Files API once — from a |
| List the files currently uploaded under this API key, with MIME types and expiry times |
| Delete an uploaded file before its ~48h expiry (confirmed first — see Confirmations) |
| (hosted deployments only) Mint a fresh signed URL for generated media from its |
| (hosted deployments only) Return a generated image as an inline image block from its |
| (hosted deployments only) Mint a short-lived signed PUT URL so a shell can upload a reference image with no auth header; the PUT returns an |
| (hosted deployments only) Persistent per-account character library: save a reference image + description under a name, then pass |
| (hosted deployments only) Persistent per-account style presets: a reusable prompt fragment (optionally with a reference image), applied by passing |
Generation tools also share three throughput/latency controls: async: true (return a job_id
immediately), max_wait_ms (wait up to a budget, then hand back the job_id — fast results stay
in-band, slow batches never trip the host timeout), and idempotency_key (a retry returns the
recorded result instead of re-billing). On a hosted deployment, a gemini_image_set result with
more than one image also carries a bundle_url — one signed URL for a zip of every image in
the set — and set links are signed for ~7 days instead of the default ~48h. (A set too large to
zip safely in memory skips the bundle and says so via bundle_skipped; the per-image
links are unaffected.)
Seeing your images (hosted)
On a hosted deployment there is no filesystem, so a generated image has to come back as something you can open. It does: every result includes a URL, with no configuration.
{
"images": ["https://mcp.nullnet.app/b/<account>/gemini/gen/2026-07-29/ab12cd34-a-cat.png?exp=…&sig=…"],
"media": [{ "url": "https://…", "r2_key": "gen/2026-07-29/ab12cd34-a-cat.png",
"expires_at": "2026-07-31T12:00:00.000Z",
"curl_hint": "curl -sS -o a-cat.png \"https://…\"" }]
}Those links need no auth header — the signature is in the URL — so they work in a browser, in
curl, and in a chat message. They expire (48h by default) and the objects behind them are
swept on a retention schedule. The r2_key is the durable handle for that window:
gemini_sign_media(hosted deployments only) mints a fresh signed URL from anr2_key, so an expired link never forces you to re-generate — and re-pay for — the image.The interaction id is stored beside each image, so a hosted deployment recovers a lost turn the way the local one does:
continue_last: truesurvives a restart, a chained 404 is re-anchored on the image that interaction produced, andgemini_list_recent_mediashows the prompt, model and interaction id for every stored object rather than a wall of keys.gemini_view_media(hosted deployments only) returns the image itself, inline, from anr2_key. A link is not something a model can look at, so without this a generate → refine loop runs blind. It is a separate call rather than bytes on every result: you pay image tokens only on the turns you actually look.gemini_upload_filewithr2_keyturns media this server generated into a Files API reference (the server reads it back with its own key — no signature for you to mint), so a generated image can become the reference image for the next generation in one cheap call.Idempotent replays (
idempotency_key) re-mint the URLs inside the recorded result before returning it, so a reused result never carries a dead link.
Why this matters: MCP's inline image content blocks (inline: true) are visible to the
assistant but many chat clients never render them to the user, and the assistant cannot
extract bytes back out of its own context to save them elsewhere. A generation could bill
successfully and be invisible. A URL is the portable answer; inline remains available, but
it is no longer the only way to receive media.
How the bytes are served
The host stores each generated object and serves it back at a signed, expiring
URL — https://<host>/b/<account>/gemini/<key>?exp=&sig=. No setup, and no auth
header: the signature in the link is the authorization, so curl and a browser
both work.
Auth is a signed, expiring URL rather than an unlisted key. Random keys would
be simpler, but they never expire and never revoke: anything that ever logged or forwarded the
link keeps working forever. A signature scopes access to one object with a deadline, and
rotating MEDIA_URL_SECRET invalidates every outstanding link at once. The tradeoff is that
links are long and cannot be shortened by hand.
For assistants relaying a result
Show the user the URL. If your sandbox has network egress to the host, fetching it
and attaching the bytes as a file gives the nicest result; otherwise present the link itself.
Whether a given client renders  markdown inline varies by client — a bare URL is the
safe form, and a markdown link is a reasonable enhancement where you know it renders.
Retention
Variable | Default | Effect |
|
| Objects older than this are deleted by a daily cron — generated media ( |
| generated | HMAC key for |
Signed-URL lifetime is clamped to MEDIA_TTL_DAYS, so a link never outlives the object it
points at.
Prompts are stored too. Each generated object on a hosted deployment gets a small JSON
record beside it holding the prompt (first 500 characters), model and interaction id — that is
what gemini_list_recent_media shows and what lets a lost turn be recovered. The record lives
and is swept with its object, so it is kept for the same MEDIA_TTL_DAYS window, not longer.
Locally, gemini_interact writes an <image>.json sidecar (model, interaction id, options and
any text the model replied with — not the prompt itself) next to each image in the output dir;
those files are yours to keep or delete and never expire on their own. If your
prompts describe real people, bear in mind they sit in your bucket for the retention window.
Sending reference images without burning context
Every tool that takes a reference image — gemini_image_generate, gemini_image_edit,
gemini_image_set, gemini_interact, plus gemini_video_generate (reference stills) and
gemini_music_generate — accepts them four ways. Only one of them costs model context:
Parameter | Where the bytes travel | Context cost |
| the server downloads the https URL | none |
| a | none |
| the server reads its own store (hosted deployments only) | none |
| read off local disk (stdio builds only) | none |
| saved library entries, attached by name (hosted deployments only) | none |
| through the tool-call JSON | ~14k tokens per JPEG |
images_base64 is the fallback of last resort. It costs roughly 14k tokens per modest photo,
and it is silently corrupted whenever the file read that produced it was truncated — the
payload still looks like base64, so the failure surfaces as a bad generation rather than an
error. Prefer any of the other three.
images_url — the server fetches it
{ "prompt": "make it look like winter", "images_url": ["https://example.com/photo.jpg"] }Fetches are restricted to public https:// URLs — private, loopback and link-local hosts are
refused (IPv6 literals are parsed, so [::ffff:7f00:1] is caught as loopback), every redirect
hop is revalidated, and each hop is bounded by a timeout. The response must be
Content-Type: image/* and is capped at 15MB, enforced while streaming rather than trusted
from Content-Length. A failure names the offending URL. Anything over 6MB is uploaded to the
Files API and referenced by uri instead of inlined, since generateContent caps a whole
request near 20MB.
images_file_uris — upload once, reference many times
// 1. upload
{ "tool": "gemini_upload_file", "url": "https://example.com/photo.jpg" }
// → { "file_uri": "files/abc123", "mime_type": "image/jpeg", "expires": "..." }
// 2. reference it, as many times as you like
{ "prompt": "make it winter", "images_file_uris": ["files/abc123"] }
{ "prompt": "make it sunrise", "images_file_uris": ["files/abc123"] }Uploads are retained ~48h; after that the reference stops resolving (as a generic 404 —
see the chained-404 section above). gemini_image_set fetches or resolves such a reference
once and passes it to the master and every scene call.
On stdio builds, a local images path that gets referenced more than once in a session is
uploaded to the Files API automatically (keyed on path + mtime + size), so repeated edits of
the same photo stop re-sending the bytes. Editing the file invalidates the cached upload.
Signed upload URLs — no token at all (hosted)
This is the intended path for an agent with a shell: disk file → curl → r2_key → tool call,
with the image never entering the conversation. gemini_get_upload_url mints a short-lived
(~10 min) signed PUT URL, and the shell uploads with zero auth headers — the signature in
the URL is the authorization, mirroring how the download links work.
# 1. tool call: gemini_get_upload_url { filename: "photo.jpg", content_type: "image/jpeg" }
# → { upload_url, r2_key, expires_at, curl_hint }
# 2. shell:
curl -sS -X PUT -H "Content-Type: image/jpeg" --data-binary @photo.jpg "$UPLOAD_URL"
# → { "r2_key": "up/<tenant>/2026-07-31/ab12cd34-photo.jpg", "size_bytes": 812345, ... }The signature covers one tenant-scoped object key, the declared content type and the expiry;
uploads are capped at 15MB, enforced while reading the stream. Only raster image types are
accepted (jpeg/png/webp/gif/avif/heic/heif/bmp/tiff) — SVG is deliberately refused, because an
SVG is a scriptable document and /media serves from the server's own origin. The returned
r2_key is then usable three ways: directly as images_r2_keys on any generation tool (the
server reads its own bucket — no bytes in the conversation), permanently via
gemini_save_character, or as a ~48h Files API reference via gemini_upload_file({ r2_key }).
Uploads themselves follow the media retention schedule (up/ prefix, default 7 days).
Character & style library (hosted)
Recurring subjects and styles can be saved once, per account, with no expiry (the retention
cron deliberately skips the library's lib/ prefix):
// once:
{ "tool": "gemini_save_character", "name": "finn",
"description": "6-year-old boy, curly red hair", "image_r2_key": "up/…/photo.jpg" }
{ "tool": "gemini_save_style", "name": "bold-cartoon-sports",
"prompt_fragment": "bold cartoon style, thick outlines, saturated colors" }
// afterwards, on any generation:
{ "tool": "gemini_image_set",
"master_prompt": "Finn on a soccer field",
"scenes": ["kicking the ball", "celebrating a goal"],
"characters": ["finn"], "style": "bold-cartoon-sports" }Naming a character attaches its saved reference image and weaves its description into the
prompt; naming a style appends its fragment (and attaches its reference image, if it has one).
gemini_image_set passes character references to the master and every scene call, which is
what keeps the subject consistent across the set.
Quick Start
{
"mcpServers": {
"gemini": {
"command": "npx",
"args": ["-y", "@chrischall/gemini-mcp"],
"env": {
"GEMINI_API_KEY": "your-api-key-here"
}
}
}
}See SKILL.md for full usage documentation.
Available Tools
13 toolsgemini_delete_fileADestructive
Delete an uploaded file, image or photo (by file_uri) from the Gemini Files API before its ~48h expiry. Any tool call still referencing it will then fail with a generic 404, so delete only references you are finished with. Asks the user to confirm first: a confirmation prompt where the client supports one; otherwise the first call returns a preview and a confirmToken, and only a repeat call with that token proceeds (see MCP_CONFIRM_MODE).
| Name | Required | Description | Default |
|---|---|---|---|
| file_uri | Yes | The `files/<id>` reference (or full uri) to delete | |
| confirmToken | No | ONLY for the two-step confirmation fallback (a client without MCP elicitation). The confirmToken from this same tool's phase-1 "confirmation-required" response, passed back ONLY after the user has seen that preview and explicitly approved it in chat — never on the first call, never invented, never reused. Call again with the same arguments. Ignored when the client supports elicitation. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the destructiveHint annotation, the description discloses important behaviors: the 48h expiry context, the cascading 404 failure mode for dependent calls, and the confirmation/confirmToken mechanism. It is honest and detailed about the destructive consequences without contradicting the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, each earning its place: purpose + expiry, failure consequence + usage rule, and confirmation behavior. The most important information is front-loaded, and the destructive-safety caveats are organized clearly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive tool with no output schema, the description covers what an agent needs: what gets deleted, when deletion is safe, what happens to referencing calls, and how confirmation works. The reference to MCP_CONFIRM_MODE is a minor external hook, but the core call sequence is fully specified.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already covers both parameters with 100% coverage, so the baseline is 3. The description adds useful operational meaning around confirmToken (preview first, repeat call only after approval, MCP_CONFIRM_MODE), which is helpfully above baseline, though it does not significantly add syntax details for file_uri beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('Delete') and a precise resource ('uploaded file, image or photo (by file_uri) from the Gemini Files API'), which immediately disambiguates it from generation and listing siblings. It also adds meaningful context (the ~48h expiry) that the tool name alone cannot convey.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states both when to use the tool ('delete only references you are finished with') and when not to (any still-referenced file will fail with a generic 404). It also explains the two-step confirmation flow, giving the agent explicit operational guidance about how calls should proceed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gemini_get_resultA
Retrieve a generation started with async: true or handed off by max_wait_ms. Pass the returned job_id: while running it reports status "running"; on completion it returns the normal result (image URLs/paths + meta); on failure it raises the recorded error, including the case where the generation was killed before it finished. On the hosted connector job records are stored durably and survive a restart; on a local stdio server they live with the process and expire ~10 min after completion, where the output dir / .json sidecar is the fallback. A killed video/music job started with background: true is recovered from its upstream interaction when it finished there.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | The job_id returned by a generation tool called with async: true | |
| output_dir | No | Where to write media recovered from a killed job (default: $GEMINI_OUTPUT_DIR or cwd) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes far beyond the minimal annotations (readOnlyHint false, openWorldHint true). Discloses status polling, error raising, durability on hosted connector, expiration on local servers, and recovery of killed jobs. This adds substantial behavioral context not inferable from annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than average but every sentence contributes essential information. It is front-loaded with the core purpose, then expands on outcomes and storage behavior. No fluff or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a retrieval tool with no output schema, it fully specifies what is returned (image URLs/paths + meta) and what happens on failure (raises recorded error). It also covers lifecycle, persistence, and edge cases. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds context for job_id (returned by a generation tool with async: true) and output_dir (fallback for killed jobs), slightly enriching the schema. Does not fully explain all edge cases but adds meaningful value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (retrieve) and the specific resource (a generation started with async: true or max_wait_ms), and explains the three possible outcomes (running, success, failure). It is easily distinguished from sibling generation tools by naming the companion relationship.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use: to retrieve results of async generations or those handed off by max_wait_ms. It also describes fallback behavior for local stdio servers (output dir sidecar) and recovery for killed background jobs, effectively covering when not to rely solely on this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gemini_healthcheckVerify credentials and upstream reachabilityARead-onlyIdempotent
Resolves the credential the way real tools do, then makes one authenticated request to generativelanguage.googleapis.com. Reports which source supplied the credential, whether generativelanguage.googleapis.com accepted it, the round-trip time, and a plain-English hint distinguishing 'no credential' from 'credential rejected' from 'a generativelanguage.googleapis.com-side problem'. Read-only; never returns the credential itself. Call this when a real tool fails and you want to know which hop broke.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and openWorld hints. The description adds meaningful behavioral details: it never returns the credential, makes exactly one authenticated request, and reports a plain-English hint distinguishing 'no credential', 'credential rejected', and upstream problems. This goes well beyond what annotations convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences with no filler: the core behavior is front-loaded, followed by reported outputs, the read-only safety note, and the usage instruction. Every sentence earns its place and the description remains compact.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no parameters and no output schema, the description covers inputs, behavior, outputs, side effects, and invocation context. An agent has everything needed to decide when and how to call this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters and empty schema properties, so schema coverage is trivially 100%. There are no parameter semantics to describe, and the baseline of 4 for no-parameter tools applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific diagnostic purpose: resolve the credential the way real tools do, then make one authenticated request to generativelanguage.googleapis.com and report source, acceptance, RTT, and failure hint. This clearly distinguishes it from all sibling tools that generate images, video, or music, manage files, or interact with models.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly instructs to call this tool when a real tool fails and you want to know which hop broke, giving clear context for use. It doesn't name a specific alternative or say when not to use it, but the diagnostic positioning is unambiguous given the sibling list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gemini_image_editA
Edit or compose images: provide one or more input images (paths or base64), plus a text instruction. For a SERIES of successive edits to the same image, prefer gemini_interact (multi-turn) — it keeps edit context and avoids re-processing the full image each round; use gemini_image_edit for one-off edits or composing multiple distinct inputs. Gemini over-preserves the input; there is no edit-strength control — for large structural changes, reroll with a different seed or more forceful wording. Local file inputs are confirmed first: a confirmation prompt where the client supports one; otherwise the first call returns a preview and a confirmToken, and only a repeat call with that token proceeds (see MCP_CONFIRM_MODE).
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | Seed for reproducible generation; random if omitted | |
| async | No | Return a job_id immediately instead of the result, so a long generation cannot hit the host tools/call timeout (-32001); poll gemini_get_result. On the hosted connector prefer max_wait_ms — the executor only lives while a request is open, so async is served there as a bounded wait. | |
| model | No | Model id override (default: server default; see gemini_list_models). gemini-3.1-flash-image (Nano Banana 2, the default) is the generalist: fast, 4K, reliable text, strong multi-reference consistency. gemini-3-pro-image (Pro) is for the hardest work — best world knowledge, brand precision. gemini-3.1-flash-lite-image (Lite) is cheapest: 1K only, no search grounding. | |
| style | No | Name of a saved style preset (gemini_list_styles): its prompt fragment, and reference image if it has one, are applied automatically. Hosted connector only. | |
| images | No | Paths to input image file(s) (1 = edit, 2+ = compose) | |
| inline | No | Return the image as an inline image block you can SEE, instead of a path (stdio) or link (hosted). The default costs nothing to carry and hands back a reference you can reuse; use this when you need to check the result yourself. On the hosted connector gemini_view_media does the same for an image you already have | |
| prompt | Yes | Instruction describing the edit or composition | |
| filename | No | Base filename for the output image (extension stripped; default: slugified prompt) | |
| characters | No | Names of saved characters (gemini_list_characters): each one's reference image and description are attached automatically, keeping recurring subjects consistent without re-sending anything. Hosted connector only. | |
| image_size | No | Output resolution (512 = 0.5K, Flash-only) | |
| images_url | No | Input images as public https URLs — the server downloads them, so no image bytes cross the conversation. Preferred over images_base64, which costs ~14k tokens per photo. Max 15MB each, Content-Type image/*. | |
| output_dir | No | Directory to write images to (default: $GEMINI_OUTPUT_DIR or cwd) | |
| timeout_ms | No | Upstream timeout in ms for this call (default $GEMINI_TIMEOUT_MS, else 60000 — 120000 at 4K, which runs past 60s) | |
| max_wait_ms | No | Wait up to this many ms in-band, then hand back { job_id, status: "running" } to poll with gemini_get_result (e.g. 20000 for multi-image sets). Keeps fast results inline and slow ones off the host timeout (-32001). Ignored when async is set. | |
| orientation | No | Output shape in plain terms: landscape (16:9), portrait (9:16) or square (1:1). For any other proportion — 3:2, 4:3, 4:5, 21:9 — use aspect_ratio, which wins if both are given. | |
| aspect_ratio | No | Exact output aspect ratio. For a plain landscape/portrait/square request, `orientation` is the shorthand; this wins if both are given. | |
| confirmToken | No | ONLY for the two-step confirmation fallback (a client without MCP elicitation). The confirmToken from this same tool's phase-1 "confirmation-required" response, passed back ONLY after the user has seen that preview and explicitly approved it in chat — never on the first call, never invented, never reused. Call again with the same arguments. Ignored when the client supports elicitation. | |
| google_search | No | Ground the image in live Google Search results (current events, weather, data) | |
| images_base64 | No | Input images as base64 strings or data URIs. Last resort — about 14k tokens per photo; prefer images_url or images_file_uris. The server uploads each one and reports a file_uri under image_inputs: pass that to images_file_uris next time instead of re-sending the bytes. | |
| from_clipboard | No | Use the image currently on the macOS system clipboard as an input (downscaled to JPEG) | |
| images_r2_keys | No | Input images by r2_key from this connector's store: a signed upload (gemini_get_upload_url → curl PUT) or an earlier generation's media[].r2_key. The server reads its own bucket — no bytes in the conversation, no ~48h expiry. Hosted connector only. | |
| thinking_level | No | Reasoning depth (Gemini 3 models); higher can help complex/structural edits | |
| idempotency_key | No | Repeat calls with this key return the recorded result (reused: true) instead of billing a new generation. Set it when retrying after a host timeout (-32001). | |
| images_file_uris | No | Input images as Files API references ("files/<id>" or the full uri) from gemini_upload_file. Upload once and reuse across calls with no bytes in the conversation; retained ~48h, after which the reference stops resolving. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the sparse annotations (readOnlyHint: false, openWorldHint: true), the description discloses important behavioral traits: Gemini over-preserves the input, there is no edit-strength control, and local file inputs require a confirmation flow with confirmToken. This adds meaningful context the annotations do not provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but efficiently organized: purpose first, then alternative routing, then behavior caveats, then confirmation mechanics. Every sentence earns its place, and it stays readable despite the tool's 24-parameter complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with no output schema, the description covers the key operational decisions an agent needs: when to use it versus gemini_interact, how to handle confirmation, what to expect regarding edit strength, and how to retry. The rich per-parameter schema fills in the remaining invocation details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds genuine semantics beyond the schema by explaining how to use seed and prompt wording to compensate for the lack of edit-strength control. It also clarifies the confirmation-token flow, which is central to correctly using confirmToken.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource: 'Edit or compose images' and clearly states the core invocation pattern (input images plus a text instruction). It also distinguishes itself from gemini_interact by naming the sibling and the condition that selects it.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says when to prefer gemini_interact for a series of edits and when to use this tool for one-off edits or composing multiple distinct inputs. It also gives practical guidance for large structural changes (reroll with a different seed or more forceful wording), which helps an agent use the tool effectively.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gemini_image_generateA
Generate image(s) from a text prompt with a Gemini image model (Nano Banana / Nano Banana Pro). If the result will likely be refined iteratively, prefer gemini_interact (multi-turn) as the entry point. Local file inputs are confirmed first: a confirmation prompt where the client supports one; otherwise the first call returns a preview and a confirmToken, and only a repeat call with that token proceeds (see MCP_CONFIRM_MODE).
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | Seed for reproducible generation; random if omitted | |
| async | No | Return a job_id immediately instead of the result, so a long generation cannot hit the host tools/call timeout (-32001); poll gemini_get_result. On the hosted connector prefer max_wait_ms — the executor only lives while a request is open, so async is served there as a bounded wait. | |
| count | No | Number of independent images (default 1) | |
| model | No | Model id override (default: server default; see gemini_list_models). gemini-3.1-flash-image (Nano Banana 2, the default) is the generalist: fast, 4K, reliable text, strong multi-reference consistency. gemini-3-pro-image (Pro) is for the hardest work — best world knowledge, brand precision. gemini-3.1-flash-lite-image (Lite) is cheapest: 1K only, no search grounding. | |
| style | No | Name of a saved style preset (gemini_list_styles): its prompt fragment, and reference image if it has one, are applied automatically. Hosted connector only. | |
| images | No | Paths to reference input images (image-conditioned generation) | |
| inline | No | Return the image as an inline image block you can SEE, instead of a path (stdio) or link (hosted). The default costs nothing to carry and hands back a reference you can reuse; use this when you need to check the result yourself. On the hosted connector gemini_view_media does the same for an image you already have | |
| prompt | Yes | Text prompt describing the image | |
| filename | No | Base filename for the output image (extension stripped; default: slugified prompt) | |
| video_url | No | Public YouTube URL (or a previously uploaded Files API uri) as a video reference (video→image; use a Flash model e.g. gemini-3.1-flash-image) | |
| characters | No | Names of saved characters (gemini_list_characters): each one's reference image and description are attached automatically, keeping recurring subjects consistent without re-sending anything. Hosted connector only. | |
| image_size | No | Output resolution (512 = 0.5K, Flash-only) | |
| images_url | No | Reference images as public https URLs — the server downloads them, so no image bytes cross the conversation. Preferred over images_base64, which costs ~14k tokens per photo. Max 15MB each, Content-Type image/*. | |
| output_dir | No | Directory to write images to (default: $GEMINI_OUTPUT_DIR or cwd) | |
| timeout_ms | No | Upstream timeout in ms for this call (default $GEMINI_TIMEOUT_MS, else 60000 — 120000 at 4K, which runs past 60s) | |
| video_path | No | Path to a local video file — uploaded to the Gemini Files API (~48h retention, 2 GB max) and used as the video reference. Alternative to video_url. | |
| max_wait_ms | No | Wait up to this many ms in-band, then hand back { job_id, status: "running" } to poll with gemini_get_result (e.g. 20000 for multi-image sets). Keeps fast results inline and slow ones off the host timeout (-32001). Ignored when async is set. | |
| orientation | No | Output shape in plain terms: landscape (16:9), portrait (9:16) or square (1:1). For any other proportion — 3:2, 4:3, 4:5, 21:9 — use aspect_ratio, which wins if both are given. | |
| aspect_ratio | No | Exact output aspect ratio. For a plain landscape/portrait/square request, `orientation` is the shorthand; this wins if both are given. | |
| confirmToken | No | ONLY for the two-step confirmation fallback (a client without MCP elicitation). The confirmToken from this same tool's phase-1 "confirmation-required" response, passed back ONLY after the user has seen that preview and explicitly approved it in chat — never on the first call, never invented, never reused. Call again with the same arguments. Ignored when the client supports elicitation. | |
| google_search | No | Ground the image in live Google Search results (current events, weather, data) | |
| images_base64 | No | Reference images as base64 strings or data URIs. Last resort — about 14k tokens per photo; prefer images_url or images_file_uris. The server uploads each one and reports a file_uri under image_inputs: pass that to images_file_uris next time instead of re-sending the bytes. | |
| from_clipboard | No | Use the image currently on the macOS system clipboard as an input (downscaled to JPEG) | |
| images_r2_keys | No | Reference images by r2_key from this connector's store: a signed upload (gemini_get_upload_url → curl PUT) or an earlier generation's media[].r2_key. The server reads its own bucket — no bytes in the conversation, no ~48h expiry. Hosted connector only. | |
| thinking_level | No | Reasoning depth (Gemini 3 models); higher can help complex/structural edits | |
| idempotency_key | No | Repeat calls with this key return the recorded result (reused: true) instead of billing a new generation. Set it when retrying after a host timeout (-32001). | |
| images_file_uris | No | Reference images as Files API references ("files/<id>" or the full uri) from gemini_upload_file. Upload once and reuse across calls with no bytes in the conversation; retained ~48h, after which the reference stops resolving. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only convey readOnly=false and openWorld=true. The description adds a genuinely important behavioral trait: local file inputs are confirmed first, and on clients without MCP elicitation the first call returns a preview plus confirmToken, requiring a repeat call. This is valuable context beyond the annotations, though it does not mention output storage or latency, which are left to the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences with no filler: purpose, entry-point routing, and confirmation behavior. The decision-critical alternative (gemini_interact for iterative refinement) is front-loaded near the top, and every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 27-parameter tool with no output schema, the description is lean but sufficient for selection and invocation, aided by the very rich schema. It covers the two non-obvious call decisions: when to route to gemini_interact and what to expect with local file inputs. Return-value shape (path vs inline vs job_id) is not stated directly, but the schema's inline and async parameter descriptions cover that.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the individual parameter descriptions are extensive (e.g., confirmToken's usage rules, async/max_wait_ms interplay, images_r2_keys). The description paraphrases the confirmation flow but adds no per-parameter meaning beyond the schema, so the baseline score of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb-resource pair: 'Generate image(s) from a text prompt' with a Gemini image model. This clearly distinguishes the tool from siblings like gemini_video_generate and gemini_music_generate, and the explicit mention of 'image(s)' versus the edit-oriented sibling gemini_image_edit makes the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives an explicit alternative with a condition: 'If the result will likely be refined iteratively, prefer gemini_interact (multi-turn) as the entry point.' It also documents the local-file confirmation workflow, telling the agent when a first call returns only a preview/confirmToken versus when it proceeds. This is actionable routing guidance, not a restatement of the schema.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gemini_image_setA
Generate a consistent SET of images: a master image from master_prompt, then one image per scene that references the master so the subject/style stays consistent. Provide scenes (explicit per-image prompts) OR count (variations of the master). Scene generations run in parallel (reference_mode "master", the default). On the hosted connector: saved characters and a saved style can seed the whole set by name, multi-image results include a bundle_url zip of every image (one curl instead of N), and max_wait_ms returns a pollable job handle if the batch runs long. Local file inputs are confirmed first: a confirmation prompt where the client supports one; otherwise the first call returns a preview and a confirmToken, and only a repeat call with that token proceeds (see MCP_CONFIRM_MODE).
| Name | Required | Description | Default |
|---|---|---|---|
| seed | No | Seed for reproducible generation; random if omitted | |
| async | No | Return a job_id immediately instead of the result, so a long generation cannot hit the host tools/call timeout (-32001); poll gemini_get_result. On the hosted connector prefer max_wait_ms — the executor only lives while a request is open, so async is served there as a bounded wait. | |
| count | No | Number of variations of master_prompt (when scenes omitted) | |
| model | No | Model id override (default: server default; see gemini_list_models). gemini-3.1-flash-image (Nano Banana 2, the default) is the generalist: fast, 4K, reliable text, strong multi-reference consistency. gemini-3-pro-image (Pro) is for the hardest work — best world knowledge, brand precision. gemini-3.1-flash-lite-image (Lite) is cheapest: 1K only, no search grounding. | |
| style | No | Name of a saved style preset (gemini_list_styles): its prompt fragment, and reference image if it has one, are applied automatically. Hosted connector only. | |
| inline | No | Return the image as an inline image block you can SEE, instead of a path (stdio) or link (hosted). The default costs nothing to carry and hands back a reference you can reuse; use this when you need to check the result yourself. On the hosted connector gemini_view_media does the same for an image you already have | |
| scenes | No | Per-image prompts (1-8); each references the master | |
| basename | No | Base filename prefix for output images (default: slugified master_prompt) | |
| characters | No | Names of saved characters (gemini_list_characters): each one's reference image and description are attached automatically, keeping recurring subjects consistent without re-sending anything. Hosted connector only. | |
| image_size | No | Output resolution (512 = 0.5K, Flash-only) | |
| output_dir | No | Directory to write images to (default: $GEMINI_OUTPUT_DIR or cwd) | |
| timeout_ms | No | Upstream timeout in ms for this call (default $GEMINI_TIMEOUT_MS, else 60000 — 120000 at 4K, which runs past 60s) | |
| max_wait_ms | No | Wait up to this many ms in-band, then hand back { job_id, status: "running" } to poll with gemini_get_result (e.g. 20000 for multi-image sets). Keeps fast results inline and slow ones off the host timeout (-32001). Ignored when async is set. | |
| orientation | No | Output shape in plain terms: landscape (16:9), portrait (9:16) or square (1:1). For any other proportion — 3:2, 4:3, 4:5, 21:9 — use aspect_ratio, which wins if both are given. | |
| aspect_ratio | No | Exact output aspect ratio. For a plain landscape/portrait/square request, `orientation` is the shorthand; this wins if both are given. | |
| confirmToken | No | ONLY for the two-step confirmation fallback (a client without MCP elicitation). The confirmToken from this same tool's phase-1 "confirmation-required" response, passed back ONLY after the user has seen that preview and explicitly approved it in chat — never on the first call, never invented, never reused. Call again with the same arguments. Ignored when the client supports elicitation. | |
| google_search | No | Ground the image in live Google Search results (current events, weather, data) | |
| master_images | No | Reference image paths passed to the master generation call | |
| master_prompt | Yes | Prompt for the master/reference image | |
| from_clipboard | No | Use the image currently on the macOS system clipboard as an input (downscaled to JPEG) | |
| reference_mode | No | master: every image references the master (default). chain: each references the previous. | |
| thinking_level | No | Reasoning depth (Gemini 3 models); higher can help complex/structural edits | |
| idempotency_key | No | Repeat calls with this key return the recorded result (reused: true) instead of billing a new generation. Set it when retrying after a host timeout (-32001). | |
| master_images_url | No | Reference images passed to the master AND to every scene call (fetched once) as public https URLs — the server downloads them, so no image bytes cross the conversation. Preferred over images_base64, which costs ~14k tokens per photo. Max 15MB each, Content-Type image/*. | |
| master_images_base64 | No | Reference images for the master generation as base64 strings or data URIs. Last resort — about 14k tokens per photo; prefer master_images_url or master_images_file_uris. The server uploads each one and reports a file_uri under image_inputs: pass that to master_images_file_uris next time instead of re-sending the bytes. | |
| master_images_r2_keys | No | Reference images passed to the master AND to every scene call by r2_key from this connector's store: a signed upload (gemini_get_upload_url → curl PUT) or an earlier generation's media[].r2_key. The server reads its own bucket — no bytes in the conversation, no ~48h expiry. Hosted connector only. | |
| master_images_file_uris | No | Reference images passed to the master AND to every scene call as Files API references ("files/<id>" or the full uri) from gemini_upload_file. Upload once and reuse across calls with no bytes in the conversation; retained ~48h, after which the reference stops resolving. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Far exceeds what annotations provide (readOnlyHint=false, openWorldHint=true). Discloses parallel scene generation, the two-step confirmation flow with confirmToken, bundle_url aggregation, pollable job handles, model-specific behaviors, and timeout semantics. No contradiction with annotations — readOnlyHint=false aligns with generation being a mutating operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Long but densely informative — appropriate for a 27-parameter tool. The core purpose is front-loaded in the first sentence, and subsequent clauses deliver scoped specifics (parallel execution, hosted-connector features, confirmation flow). Minor redundancy with schema descriptions (model notes, orientation) means a few sentences could be trimmed, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with 27 parameters, no output schema, and platform-dependent behavior, this is remarkably complete. It covers the generation workflow, parallel execution, confirmation gating, async/polling strategy, hosted-vs-local differences, and model selection. Nothing an agent needs to correctly invoke or recover from timeouts is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with individually rich schema descriptions (e.g., token cost warnings on base64, retention windows on file_uris). The tool description adds meaningful value beyond the schema: bundle_url aggregation, confirmToken lifecycle, parallel execution of scenes, and reference_mode default. This is above the coverage-baseline of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a precise verb+resource: 'Generate a consistent SET of images: a master image from master_prompt, then one image per scene that references the master.' It immediately distinguishes this from the sibling gemini_image_generate (single image) and gemini_image_edit by describing the set-based workflow and consistency mechanism.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Extensive in-tool guidance: scenes vs count trade-off, reference_mode master vs chain, async vs max_wait_ms polling, hosted vs local connector differences, and confirmation flow. However, it never explicitly names alternatives or says 'use gemini_image_generate for a single image' — sibling routing is implied by the name and workflow rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gemini_interactA
Preferred tool for iterative refinement of a single image — multi-turn generation and editing via the Interactions API. To refine, pass the returned interaction id as previous_interaction_id on the next call; do NOT start a new interaction or re-send the image for each tweak. continue_last: true chains from this session's most recent interaction without threading the id. If a call times out client-side the generation usually still completes: the image and an <image>.json sidecar holding its interaction id land in the output dir, and continue_last: true still resumes it — look there before re-issuing, which bills a second generation. A chained 404 is re-anchored on the prior output and re-issued un-chained (chain_recovered); a second 404 means the interaction id was not the cause — check the model id / files uri. Output is JPEG. Local file inputs are confirmed first: a confirmation prompt where the client supports one; otherwise the first call returns a preview and a confirmToken, and only a repeat call with that token proceeds (see MCP_CONFIRM_MODE).
| Name | Required | Description | Default |
|---|---|---|---|
| async | No | Return a job_id immediately instead of the result, so a long generation cannot hit the host tools/call timeout (-32001); poll gemini_get_result. On the hosted connector prefer max_wait_ms — the executor only lives while a request is open, so async is served there as a bounded wait. | |
| input | Yes | Text prompt or editing instruction | |
| model | No | Model id override (default: server default; see gemini_list_models). gemini-3.1-flash-image (Nano Banana 2, the default) is the generalist: fast, 4K, reliable text, strong multi-reference consistency. gemini-3-pro-image (Pro) is for the hardest work — best world knowledge, brand precision. gemini-3.1-flash-lite-image (Lite) is cheapest: 1K only, no search grounding. | |
| images | No | Paths to reference input images. NEW reference images only (e.g. a style or target photo). When chaining with previous_interaction_id, do NOT re-attach the prior turn's output — the interaction already contains it, and re-sending it anchors the model against your edit. | |
| inline | No | Return base64 images inline instead of writing to disk | |
| filename | No | Base filename for the output image (extension stripped; default: slugified input) | |
| video_url | No | Public YouTube URL (or a previously uploaded Files API uri) as a video reference (video→image; use a Flash model e.g. gemini-3.1-flash-image) | |
| image_size | No | Output resolution (512 = 0.5K, Flash-only) | |
| images_url | No | Reference images. NEW reference images only (e.g. a style or target photo). When chaining with previous_interaction_id, do NOT re-attach the prior turn's output — the interaction already contains it, and re-sending it anchors the model against your edit. Given as public https URLs — the server downloads them, so no image bytes cross the conversation. Preferred over images_base64, which costs ~14k tokens per photo. Max 15MB each, Content-Type image/*. | |
| output_dir | No | Directory to write images to (default: $GEMINI_OUTPUT_DIR or cwd) | |
| timeout_ms | No | Upstream timeout in ms for this call (default $GEMINI_TIMEOUT_MS, else 60000 — 120000 at 4K, which runs past 60s) | |
| video_path | No | Path to a local video file — uploaded to the Gemini Files API (~48h retention, 2 GB max) and used as the video reference. Alternative to video_url. | |
| max_wait_ms | No | Wait up to this many ms in-band, then hand back { job_id, status: "running" } to poll with gemini_get_result (e.g. 20000 for multi-image sets). Keeps fast results inline and slow ones off the host timeout (-32001). Ignored when async is set. | |
| orientation | No | Output shape in plain terms: landscape (16:9), portrait (9:16) or square (1:1). For any other proportion — 3:2, 4:3, 4:5, 21:9 — use aspect_ratio, which wins if both are given. | |
| aspect_ratio | No | Exact output aspect ratio. For a plain landscape/portrait/square request, `orientation` is the shorthand; this wins if both are given. | |
| confirmToken | No | ONLY for the two-step confirmation fallback (a client without MCP elicitation). The confirmToken from this same tool's phase-1 "confirmation-required" response, passed back ONLY after the user has seen that preview and explicitly approved it in chat — never on the first call, never invented, never reused. Call again with the same arguments. Ignored when the client supports elicitation. | |
| search_types | No | Grounding search types (implies google_search). image_search (gemini-3.1-flash-image only) uses Google Image Search results as visual references; per Google ToS the returned grounding.search_suggestions HTML must then be displayed to the user. Cannot depict real people from web images. | |
| continue_last | No | Continue from the most recent interaction this server created (convenience for previous_interaction_id; an explicit id wins). Survives a server restart by falling back to the newest <image>.json sidecar in the output dir. | |
| google_search | No | Ground the image in live Google Search results (current events, weather, data) | |
| images_base64 | No | Reference images. NEW reference images only (e.g. a style or target photo). When chaining with previous_interaction_id, do NOT re-attach the prior turn's output — the interaction already contains it, and re-sending it anchors the model against your edit. Given as base64 strings or data URIs. Last resort — about 14k tokens per photo; prefer images_url or images_file_uris. The server uploads each one and reports a file_uri under image_inputs: pass that to images_file_uris next time instead of re-sending the bytes. | |
| from_clipboard | No | Use the image currently on the macOS system clipboard as an input (downscaled to JPEG) | |
| thinking_level | No | Reasoning depth; higher can help complex/structural edits | |
| idempotency_key | No | Repeat calls with this key return the recorded result (reused: true) instead of billing a new generation. Set it when retrying after a host timeout (-32001). | |
| images_file_uris | No | Reference images. NEW reference images only (e.g. a style or target photo). When chaining with previous_interaction_id, do NOT re-attach the prior turn's output — the interaction already contains it, and re-sending it anchors the model against your edit. Given as Files API references ("files/<id>" or the full uri) from gemini_upload_file. Upload once and reuse across calls with no bytes in the conversation; retained ~48h, after which the reference stops resolving. | |
| previous_interaction_id | No | ID from a prior gemini_interact call — continues that multi-turn conversation |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint:false and openWorldHint:true, and the description richly corroborates and extends them: it discloses that client-side timeouts still complete the generation (with a sidecar for recovery), that re-issuing bills a second generation, that chained 404s are re-anchored (chain_recovered), the two-step confirmation fallback with confirmToken, and JPEG output. No contradiction with annotations; the description carries substantial behavioral context beyond the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence earns its place — chaining, timeout recovery, 404 handling, confirmation, output format are all substantive and non-redundant. However, it is delivered as a single dense wall of text with no paragraph breaks or bullet structure, which hampers scanability for a 25-parameter tool; the density is justified by complexity but the formatting is not.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex mutation tool with no output schema, the description covers the critical failure modes: timeout recovery, chained 404 re-anchoring, billing implications, output sidecar location, and the confirmation flow. Given 25 parameters and no output schema, it cannot be exhaustive, and model-selection guidance lives in the schema rather than the description, so a 4 reflects thorough but not total coverage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3 per rubric. The tool description references previous_interaction_id, continue_last, output_dir, and confirmToken contextually, but does not add syntax or format details — the rich per-parameter explanations live in the schema itself, which already handles the semantic load.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb-resource-mechanism statement — 'iterative refinement of a single image — multi-turn generation and editing via the Interactions API' — which precisely scopes the tool and distinguishes it from siblings like gemini_image_generate, gemini_image_edit, and gemini_image_set. The chaining mechanics (previous_interaction_id, continue_last) further differentiate it as the multi-turn tool in a family of single-shot tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'Preferred tool for iterative refinement' opener gives clear usage context, and the 'do NOT start a new interaction or re-send the image for each tweak' instruction is explicit workflow guidance. However, it never names the alternatives it displaces (e.g., when to choose gemini_image_generate or gemini_image_edit instead), leaving exclusion criteria implicit rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gemini_list_filesARead-only
List files, images and photos currently uploaded to the Gemini Files API under this API key, with their reusable file_uri (files/<id>) references, MIME types and expiry times. Retention is ~48h, so an entry that has vanished has expired rather than failed.
| Name | Required | Description | Default |
|---|---|---|---|
| page_size | No | Maximum files to return (1-100, default 100) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, so the description doesn't need to restate safety. It adds valuable behavioral context: retention is ~48h, so vanished entries have expired rather than failed, and it discloses that results include file_uri, MIME types, and expiry times. This goes beyond the annotations and helps the agent interpret results correctly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core purpose and scope, followed by a useful retention caveat. No wasted words; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with one optional parameter and no output schema, the description is nearly complete. It explains what is returned (file_uri, MIME types, expiry times) and how to interpret missing entries. It doesn't describe pagination or the exact response format, but those are minor given the tool's simplicity and the annotations covering safety.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single optional page_size parameter, so the schema already documents it fully. The description doesn't add parameter-specific details, but with full coverage, baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists files, images, and photos uploaded to the Gemini Files API under the current API key, and distinguishes it from other operations by mentioning reusable file_uri references, MIME types, and expiry times. It uses a specific verb ('List') and resource ('files... under this API key'), making it easy to differentiate from siblings like gemini_upload_file and gemini_delete_file.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool: to check currently uploaded files and their references, and clarifies that missing entries mean expiry rather than failure. It doesn't explicitly name alternatives or exclusions, but the context is clear enough for an agent to select it over upload/delete/get_result tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gemini_list_modelsARead-only
List the Gemini image-generation models available to your API key (Nano Banana / Nano Banana Pro family), and the current default model.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already indicates no side effects. The description adds context that the tool lists models specific to the API key and includes the default model, but does not disclose additional behavioral traits beyond what annotations already provide. It confirms safe behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence that fully conveys the tool's purpose without extra words. It is appropriately sized and front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (no parameters, no output schema, no nested objects), the description is sufficiently complete. It explains what is returned (models list and default model). However, it does not detail the structure of each model entry, which is acceptable for a list operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has no parameters, so baseline is 4. The description adds meaning by specifying the model family (Nano Banana / Nano Banana Pro) and that the default model is included, which helps the agent understand what will be returned.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool lists Gemini image-generation models available to the API key, including a specific model family (Nano Banana / Nano Banana Pro) and the current default model. This provides specific verb and resource, distinguishing it from sibling tools like gemini_image_generate or gemini_list_files.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for checking available models and the default, but does not explicitly state when to use it versus alternatives, nor does it provide exclusions or when-not-to-use guidance. Since it's a simple listing tool, it's adequate but could be improved.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gemini_music_generateA
Generate music from a text prompt (mood, genre, instruments, structure, or lyrics inline) via a Lyria model: lyria-3-clip-preview (30s instrumental clip, default, cheapest), lyria-3.5 (full-length song with vocals) or lyria-3-pro-preview (longer-form). Output is MP3, written to disk (or returned inline). Single-turn: a track cannot be refined by a follow-up call, so put the whole brief in the prompt. Runs long — give it a max_wait_ms budget (or async: true + gemini_get_result on a local install), or raise timeout_ms. Needs a funded account. Local file inputs are confirmed first: a confirmation prompt where the client supports one; otherwise the first call returns a preview and a confirmToken, and only a repeat call with that token proceeds (see MCP_CONFIRM_MODE).
| Name | Required | Description | Default |
|---|---|---|---|
| async | No | Return a job_id immediately instead of the result, so a long generation cannot hit the host tools/call timeout (-32001); poll gemini_get_result. On the hosted connector prefer max_wait_ms — the executor only lives while a request is open, so async is served there as a bounded wait. | |
| model | No | Lyria model (default: lyria-3-clip-preview — 30s, $0.04). lyria-3.5 and lyria-3-pro-preview run minutes-long at $0.08. | |
| images | No | Optional reference image path(s) to condition the music | |
| inline | No | Return base64 audio inline instead of writing to disk. A 30s track is ~1.9MB of base64 — the default hands back a path instead | |
| prompt | Yes | Description of the music: mood, genre, instruments, tempo, structure, or lyrics | |
| filename | No | Base filename for the output audio (extension stripped; default: slugified prompt) | |
| background | No | Run the generation on Google's side and poll it, so a killed job can be recovered by gemini_get_result. Off by default — see gemini_video_generate | |
| images_url | No | Reference images as public https URLs — the server downloads them, so no image bytes cross the conversation. Preferred over images_base64, which costs ~14k tokens per photo. Max 15MB each, Content-Type image/*. | |
| output_dir | No | Directory to write audio to (default: $GEMINI_OUTPUT_DIR or cwd) | |
| timeout_ms | No | Upstream timeout in ms for this call (default $GEMINI_TIMEOUT_MS, else 60000 — 120000 at 4K, which runs past 60s) | |
| max_wait_ms | No | Wait up to this many ms in-band, then hand back { job_id, status: "running" } to poll with gemini_get_result (e.g. 20000 for multi-image sets). Keeps fast results inline and slow ones off the host timeout (-32001). Ignored when async is set. | |
| confirmToken | No | ONLY for the two-step confirmation fallback (a client without MCP elicitation). The confirmToken from this same tool's phase-1 "confirmation-required" response, passed back ONLY after the user has seen that preview and explicitly approved it in chat — never on the first call, never invented, never reused. Call again with the same arguments. Ignored when the client supports elicitation. | |
| images_base64 | No | Reference images as base64 strings or data URIs. Last resort — about 14k tokens per photo; prefer images_url or images_file_uris. The server uploads each one and reports a file_uri under image_inputs: pass that to images_file_uris next time instead of re-sending the bytes. | |
| from_clipboard | No | Use the image currently on the macOS clipboard as a reference | |
| idempotency_key | No | Repeat calls with this key return the recorded result (reused: true) instead of billing a new generation. Set it when retrying after a host timeout (-32001). | |
| images_file_uris | No | Reference images as Files API references ("files/<id>" or the full uri) from gemini_upload_file. Upload once and reuse across calls with no bytes in the conversation; retained ~48h, after which the reference stops resolving. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the minimal annotations, the description discloses that calls run long, may time out, write to disk or return inline, cannot be refined, can require a two-step confirmation fallback, and may need idempotency_key for retries. It also surfaces cost/speed differences between models. This goes well beyond what annotations provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded with the core action, model options, and output behavior before caveats. It is long, but the tool is complex; some model details are repeated from the schema and the long paragraph could be tightened, so it is not a perfect 5, but every sentence carries operational value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 16-parameter tool with no output schema, the description covers the critical operational context: runtime behavior, timeout/async handling, confirmation flow, retry/idempotency, file output vs inline, account prerequisite, and model selection. Nothing essential for an agent to call this tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3 and the schema carries the heavy lifting. The description adds meaningful extras beyond the schema: MP3 output format, 'full-length song with vocals' for lyria-3.5, 'longer-form' for pro, the single-turn prompt expectation, and the max_wait/async strategy. This justifies moving above baseline without overstating.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action ('Generate music from a text prompt'), a concrete resource (Lyria model), and the output format (MP3). It differentiates from siblings by explicitly naming music generation and model variants, so an agent can distinguish it from gemini_image_generate and gemini_video_generate without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives strong usage direction: it explains single-turn behavior, instructs putting the whole brief in the prompt, directs hosted vs local installs to max_wait_ms vs async + gemini_get_result, references gemini_video_generate for background polling, and warns about funded-account requirements. This is explicit when/how guidance with named alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gemini_token_usageARead-only
Token usage for this session so far — what every generation has cost in tokens, added up. Call it before and after a workflow and subtract to get that workflow's cost; call it after a single generation for that call's. Reports tokens AND an estimated USD cost, priced per call against each call's own model and stamped with the date its rates were read (override with GEMINI_RATE_CARD). Note there is no account-balance endpoint to query — Google Cloud is post-paid and its billing data lags by hours — so this is the accurate way to attribute spend to a call.
| Name | Required | Description | Default |
|---|---|---|---|
| reset | No | Zero the running total after reporting it, so the next call measures from here. Use it to bracket a workflow without arithmetic. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only indicate readOnlyHint=true, and the description adds substantial behavioral context: it reports tokens and estimated cost, prices per model, stamps rates with a date, supports GEMINI_RATE_CARD override, and explains the post-paid billing limitation. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place: core purpose, usage pattern, output type, cost calculation nuance, and why this tool is necessary. It is front-loaded with the main function and avoids fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tool with one optional parameter and no output schema, the description is complete: it explains scope, output content, usage patterns, the reset workflow indirectly, and the pricing/rate-card behavior. An agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already fully documents the reset parameter. The description adds useful context about cost calculation and rate-card override, but it does not add semantic detail about the parameter beyond what the schema provides. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it reports cumulative token usage and estimated USD cost for the session. It clearly distinguishes itself from the generation and file-management siblings by being a meter/attribution tool rather than a generative or file operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit when-to-use instructions: call before and after a workflow and subtract, or after a single generation. It also explains that no account-balance endpoint exists and billing data lags, positioning this tool as the reliable way to attribute spend.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gemini_upload_fileA
Upload a file — an image, reference photo, picture, screenshot, video or audio clip — to the Gemini Files API ONCE, and get back a reusable file_uri (files/<id>) to attach to later image, video or music generations. Keywords: upload, upload file, upload image, upload photo, attach, reference image, reference photo, file_uri, files api, image reference, reuse across calls. Use this instead of pasting base64 into a tool call: the reference is a short string, so no image bytes ever enter the conversation, and it can be reused across many generations until it expires (~48h). Provide exactly one of url (the server downloads it), data_base64, or path (a local file). A local path is confirmed first: a confirmation prompt where the client supports one; otherwise the first call returns a preview and a confirmToken, and only a repeat call with that token proceeds (see MCP_CONFIRM_MODE).
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | Public https URL the SERVER downloads and uploads (image/video/audio, up to 100MB). No bytes pass through the conversation. | |
| path | No | Path to a local file (absolute, or resolved against $GEMINI_INPUT_DIR). Confirm-gated like every other local-file input. | |
| r2_key | No | Unavailable on this server (media is written to local disk) — pass the file path via `path` instead. | |
| mime_type | No | Override the detected MIME type (sniffed from the bytes / taken from the server response otherwise) | |
| data_base64 | No | Raw base64 or a data: URI. Last resort — this is the one form that costs model context (~14k tokens for a modest JPEG). | |
| confirmToken | No | ONLY for the two-step confirmation fallback (a client without MCP elicitation). The confirmToken from this same tool's phase-1 "confirmation-required" response, passed back ONLY after the user has seen that preview and explicitly approved it in chat — never on the first call, never invented, never reused. Call again with the same arguments. Ignored when the client supports elicitation. | |
| display_name | No | Human-readable name recorded against the upload |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are minimal (readOnlyHint=false, openWorldHint=true), so the description carries the behavioral burden. It discloses the one-time upload semantics, ~48h expiry, that the file_uri is reusable across calls, context-cost impact of base64, server-side download behavior, and confirmation-token rules.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long, but dense and front-loaded with the core purpose. The keyword list and confirmation details are somewhat verbose, yet they carry practical value for an agent. Minor redundancy (e.g., image/reference photo/picture/screenshot) keeps it from a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 7 parameters, no output schema, and a nontrivial confirmation flow, the description is remarkably complete. It covers input modes, return value shape, expiration, context-cost tradeoffs, and the confirmToken lifecycle, leaving no critical call-time ambiguity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although schema coverage is 100%, the description adds significant meaning: exactly one source must be provided, url is server-downloaded, data_base64 is a costly last resort, r2_key is unusable, and confirmToken has strict two-phase usage rules. This goes well beyond the raw schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Upload') and a clear resource ('the Gemini Files API'), and explains the output is a reusable file_uri for later generations. This clearly differentiates it from sibling generation/list/delete tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly tells the agent to use this instead of pasting base64, explains when to use url vs path vs data_base64, and details the two-step confirmation behavior. This provides actionable routing guidance beyond just saying what the tool does.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gemini_video_generateA
Generate a short video (~10s) via the Gemini omni model: text→video, image→video / reference→video (supply reference image[s]), interpolate between two stills (pass first frame then last frame as images), or continue a prior video (task: "edit" or "extend" + previous_interaction_id / continue_last; extensions add ~3-10s each, to ~40s total). Cost scales with resolution — draft at 360p, keep at 1080p/4k. Output is written to disk as MP4 (video has no inline MCP block). Video runs long — give it a max_wait_ms budget (or async: true + gemini_get_result on a local install), or raise timeout_ms. Needs a funded account. Local file inputs are confirmed first: a confirmation prompt where the client supports one; otherwise the first call returns a preview and a confirmToken, and only a repeat call with that token proceeds (see MCP_CONFIRM_MODE).
| Name | Required | Description | Default |
|---|---|---|---|
| task | No | text_to_video (default), image_to_video / reference_to_video (need image input), or edit / extend (need previous_interaction_id) | |
| async | No | Return a job_id immediately instead of the result, so a long generation cannot hit the host tools/call timeout (-32001); poll gemini_get_result. On the hosted connector prefer max_wait_ms — the executor only lives while a request is open, so async is served there as a bounded wait. | |
| model | No | Model id override (default: gemini-omni-1.1-flash) | |
| images | No | Reference image path(s) for image_to_video / reference_to_video | |
| prompt | Yes | Description of the video to generate (or the edit instruction when task=edit) | |
| delivery | No | How the clip comes back: "uri" (default — a Files API link the server downloads, no size ceiling) or "inline" (base64, capped ~4MB) | |
| filename | No | Base filename for the output video (extension stripped; default: slugified prompt) | |
| background | No | Run the generation on Google's side and poll it, so a killed job can be recovered by gemini_get_result. Trade-off: retrieving a backgrounded interaction is unreliable today (some become permanently unreadable), so the default is off | |
| images_url | No | Reference stills as public https URLs — the server downloads them, so no image bytes cross the conversation. Preferred over images_base64, which costs ~14k tokens per photo. Max 15MB each, Content-Type image/*. | |
| output_dir | No | Directory to write the video to (default: $GEMINI_OUTPUT_DIR or cwd) | |
| resolution | No | Output resolution (default 720p). Video is billed per output token, so 360p costs roughly a third of 720p — use it for drafts | |
| timeout_ms | No | Upstream timeout in ms for this call (default $GEMINI_TIMEOUT_MS, else 60000 — 120000 at 4K, which runs past 60s) | |
| max_wait_ms | No | Wait up to this many ms in-band, then hand back { job_id, status: "running" } to poll with gemini_get_result (e.g. 20000 for multi-image sets). Keeps fast results inline and slow ones off the host timeout (-32001). Ignored when async is set. | |
| orientation | No | Output shape in plain terms: landscape (16:9), portrait (9:16) or square (1:1). For any other proportion — 3:2, 4:3, 4:5, 21:9 — use aspect_ratio, which wins if both are given. | |
| aspect_ratio | No | Exact output aspect ratio (omni: 16:9 or 9:16). `orientation` is the plain-language shorthand; this wins if both are given. | |
| confirmToken | No | ONLY for the two-step confirmation fallback (a client without MCP elicitation). The confirmToken from this same tool's phase-1 "confirmation-required" response, passed back ONLY after the user has seen that preview and explicitly approved it in chat — never on the first call, never invented, never reused. Call again with the same arguments. Ignored when the client supports elicitation. | |
| continue_last | No | Continue from the most recent video interaction this server created (explicit previous_interaction_id wins) | |
| images_base64 | No | Reference images as base64 strings or data URIs. Last resort — about 14k tokens per photo; prefer images_url or images_file_uris. The server uploads each one and reports a file_uri under image_inputs: pass that to images_file_uris next time instead of re-sending the bytes. | |
| from_clipboard | No | Use the image currently on the macOS clipboard as a reference | |
| idempotency_key | No | Repeat calls with this key return the recorded result (reused: true) instead of billing a new generation. Set it when retrying after a host timeout (-32001). | |
| images_file_uris | No | Reference stills as Files API references ("files/<id>" or the full uri) from gemini_upload_file. Upload once and reuse across calls with no bytes in the conversation; retained ~48h, after which the reference stops resolving. | |
| previous_interaction_id | No | Interaction id to edit/continue (with task: "edit") |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are minimal (readOnlyHint: false, openWorldHint: true), so the description carries the full behavioral burden — and it does impressively: cost scaling with resolution ('draft at 360p'), disk side effects ('Output is written to disk as MP4 (video has no inline MCP block)'), long-run timing behavior, the funded-account prerequisite, and the two-step confirmation/confirmToken fallback are all disclosed. The write-to-disk behavior is consistent with readOnlyHint: false, so there is no annotation contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
At roughly 200 words for a 22-parameter, 5-mode tool, the description is dense but every sentence carries load-bearing information, and the core function is front-loaded before any caveats. The progression (function → modes → cost → output → timing → access → confirmation) keeps the length navigable rather than rambling. Minor redundancy exists with the schema's per-parameter notes, but the prose is tighter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a high-complexity tool with no output schema, the description covers the critical runtime surfaces: each mode's input combinations, disk output, duration scaling, the async/job_id path, the max_wait_ms hand-back shape ({ job_id, status: "running" }), and the confirmToken flow. The remaining gap is the exact success-return shape of the normal synchronous path (what the MP4 reference looks like), which is only alluded to. Near-complete given the complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds genuinely non-schema semantics: interpolation requires passing first frame then last frame as images, extensions add ~3-10s each up to ~40s total, and each task value maps to a specific parameter set. The schema already covers most per-parameter semantics (transport trade-offs, defaults), so the added value is moderate rather than extensive.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource+scope ('Generate a short video (~10s) via the Gemini omni model') and immediately enumerates all five modes (text→video, image→video / reference→video, interpolation, edit/extend), making the tool's function unambiguous. This clearly differentiates it from siblings such as gemini_image_generate and gemini_music_generate, which produce other media types.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit mode routing with the parameters each mode requires ('supply reference image[s]', 'pass first frame then last frame as images', 'task: "edit" or "extend" + previous_interaction_id / continue_last'), and names the polling companion gemini_get_result plus the MCP_CONFIRM_MODE fallback. It also tells the agent how to handle long-running calls (max_wait_ms, async, timeout_ms). It lacks explicit when-not-to-use statements against sibling generators (e.g., when to prefer gemini_image_generate instead), but internal branching guidance is concrete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v2.3.0- Changed
gemini_delete_file2 fields changed- removed
Input schema / properties / confirmRemoved value: -{ - "description": "Must be true to proceed. Without this, the tool returns a preview.", - "type": "boolean" -} - added
Input schema / properties / confirmTokenAdded value: +{ + "description": "ONLY for the two-step confirmation fallback (a client without MCP elicitation). The confirmToken from this same tool's phase-1 \"confirmation-required\" response, passed back ONLY after the user has seen that preview and explicitly approved it in chat — never on the first call, never invented, never reused. Call again with the same arguments. Ignored when the client supports elicitation.", + "type": "string" +}
- Changed
gemini_image_edit2 fields changed- removed
Input schema / properties / confirmRemoved value: -{ - "description": "Must be true to proceed. Without this, the tool returns a preview.", - "type": "boolean" -} - added
Input schema / properties / confirmTokenAdded value: +{ + "description": "ONLY for the two-step confirmation fallback (a client without MCP elicitation). The confirmToken from this same tool's phase-1 \"confirmation-required\" response, passed back ONLY after the user has seen that preview and explicitly approved it in chat — never on the first call, never invented, never reused. Call again with the same arguments. Ignored when the client supports elicitation.", + "type": "string" +}
- Changed
gemini_image_generate2 fields changed- removed
Input schema / properties / confirmRemoved value: -{ - "description": "Must be true to proceed. Without this, the tool returns a preview.", - "type": "boolean" -} - added
Input schema / properties / confirmTokenAdded value: +{ + "description": "ONLY for the two-step confirmation fallback (a client without MCP elicitation). The confirmToken from this same tool's phase-1 \"confirmation-required\" response, passed back ONLY after the user has seen that preview and explicitly approved it in chat — never on the first call, never invented, never reused. Call again with the same arguments. Ignored when the client supports elicitation.", + "type": "string" +}
- Changed
gemini_image_set2 fields changed- removed
Input schema / properties / confirmRemoved value: -{ - "description": "Must be true to proceed. Without this, the tool returns a preview.", - "type": "boolean" -} - added
Input schema / properties / confirmTokenAdded value: +{ + "description": "ONLY for the two-step confirmation fallback (a client without MCP elicitation). The confirmToken from this same tool's phase-1 \"confirmation-required\" response, passed back ONLY after the user has seen that preview and explicitly approved it in chat — never on the first call, never invented, never reused. Call again with the same arguments. Ignored when the client supports elicitation.", + "type": "string" +}
- Changed
gemini_interact2 fields changed- removed
Input schema / properties / confirmRemoved value: -{ - "description": "Must be true to proceed. Without this, the tool returns a preview.", - "type": "boolean" -} - added
Input schema / properties / confirmTokenAdded value: +{ + "description": "ONLY for the two-step confirmation fallback (a client without MCP elicitation). The confirmToken from this same tool's phase-1 \"confirmation-required\" response, passed back ONLY after the user has seen that preview and explicitly approved it in chat — never on the first call, never invented, never reused. Call again with the same arguments. Ignored when the client supports elicitation.", + "type": "string" +}
- Changed
gemini_music_generate2 fields changed- removed
Input schema / properties / confirmRemoved value: -{ - "description": "Must be true to proceed. Without this, the tool returns a preview.", - "type": "boolean" -} - added
Input schema / properties / confirmTokenAdded value: +{ + "description": "ONLY for the two-step confirmation fallback (a client without MCP elicitation). The confirmToken from this same tool's phase-1 \"confirmation-required\" response, passed back ONLY after the user has seen that preview and explicitly approved it in chat — never on the first call, never invented, never reused. Call again with the same arguments. Ignored when the client supports elicitation.", + "type": "string" +}
- Changed
gemini_upload_file2 fields changed- removed
Input schema / properties / confirmRemoved value: -{ - "description": "Must be true to proceed. Without this, the tool returns a preview.", - "type": "boolean" -} - added
Input schema / properties / confirmTokenAdded value: +{ + "description": "ONLY for the two-step confirmation fallback (a client without MCP elicitation). The confirmToken from this same tool's phase-1 \"confirmation-required\" response, passed back ONLY after the user has seen that preview and explicitly approved it in chat — never on the first call, never invented, never reused. Call again with the same arguments. Ignored when the client supports elicitation.", + "type": "string" +}
- Changed
gemini_video_generate2 fields changed- removed
Input schema / properties / confirmRemoved value: -{ - "description": "Must be true to proceed. Without this, the tool returns a preview.", - "type": "boolean" -} - added
Input schema / properties / confirmTokenAdded value: +{ + "description": "ONLY for the two-step confirmation fallback (a client without MCP elicitation). The confirmToken from this same tool's phase-1 \"confirmation-required\" response, passed back ONLY after the user has seen that preview and explicitly approved it in chat — never on the first call, never invented, never reused. Call again with the same arguments. Ignored when the client supports elicitation.", + "type": "string" +}
12 tool updates
v2.0.0- Changed
gemini_delete_file1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
gemini_get_result1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
gemini_healthcheck1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
gemini_image_edit1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
gemini_image_generate1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
gemini_image_set1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
gemini_interact1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
gemini_list_files1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
gemini_music_generate1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
gemini_token_usage1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
gemini_upload_file1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
- Changed
gemini_video_generate1 field changed- changed
Input schema / $schemaPrevious value: -"http://json-schema.org/draft-07/schema#"New value: +"https://json-schema.org/draft/2020-12/schema"
6 tool updates
v1.14.0- Changed
gemini_image_edit13 fields changed- changed
Input schema / properties / async / descriptionPrevious value: -"Run in the background and return a job_id immediately instead of the image, so a long (Pro/4K) generation cannot hit the host tools/call timeout (-32001). Poll gemini_get_result with the job_id to fetch the result. PREFER `max_wait_ms` on the hosted connector: it runs where the executor is only guaranteed to stay alive while the request is open, so this option is served there as a bounded wait rather than an immediate hand-off."New value: +"Return a job_id immediately instead of the result, so a long generation cannot hit the host tools/call timeout (-32001); poll gemini_get_result. On the hosted connector prefer max_wait_ms — the executor only lives while a request is open, so async is served there as a bounded wait." - changed
Input schema / properties / characters / descriptionPrevious value: -"Names of saved characters (see gemini_list_characters / gemini_save_character): each one's reference image and description are attached to the request automatically, keeping recurring subjects consistent without re-sending anything. Hosted connector only."New value: +"Names of saved characters (gemini_list_characters): each one's reference image and description are attached automatically, keeping recurring subjects consistent without re-sending anything. Hosted connector only." - changed
Input schema / properties / idempotency_key / descriptionPrevious value: -"Opaque idempotency key: a repeat call with the same key returns the recorded result (reused: true) instead of billing a new generation. Set it when retrying after a host timeout (-32001) to avoid a duplicate charge."New value: +"Repeat calls with this key return the recorded result (reused: true) instead of billing a new generation. Set it when retrying after a host timeout (-32001)." - changed
Input schema / properties / images_base64 / descriptionPrevious value: -"Input images as base64 strings or data URIs. Last resort: prefer images_url or images_file_uris, which keep image bytes out of the conversation"New value: +"Input images as base64 strings or data URIs. Last resort — about 14k tokens per photo; prefer images_url or images_file_uris. The server uploads each one and reports a file_uri under image_inputs: pass that to images_file_uris next time instead of re-sending the bytes." - changed
Input schema / properties / images_file_uris / descriptionPrevious value: -"Input images by Gemini Files API reference (\"files/<id>\", or the full uri) from gemini_upload_file or POST /upload. Upload once, then reference it across as many calls as you like — no bytes are re-sent and none enter the conversation. Files are retained ~48h, after which the reference stops resolving."New value: +"Input images as Files API references (\"files/<id>\" or the full uri) from gemini_upload_file. Upload once and reuse across calls with no bytes in the conversation; retained ~48h, after which the reference stops resolving." - changed
Input schema / properties / images_r2_keys / descriptionPrevious value: -"Input images by r2_key from THIS connector's store: a signed upload (gemini_get_upload_url → curl PUT) or an earlier generation's media[].r2_key. The server reads its own bucket directly — no bytes in the conversation, no signed URL, no ~48h Files API expiry. Hosted connector only."New value: +"Input images by r2_key from this connector's store: a signed upload (gemini_get_upload_url → curl PUT) or an earlier generation's media[].r2_key. The server reads its own bucket — no bytes in the conversation, no ~48h expiry. Hosted connector only." - changed
Input schema / properties / images_url / descriptionPrevious value: -"Input images as public https URLs — the SERVER downloads them, so no image bytes travel through the conversation. Preferred over images_base64, which costs ~14k tokens per photo and breaks if a file read was truncated. Max 15MB each; must be a directly-linked image (Content-Type image/*)."New value: +"Input images as public https URLs — the server downloads them, so no image bytes cross the conversation. Preferred over images_base64, which costs ~14k tokens per photo. Max 15MB each, Content-Type image/*." - changed
Input schema / properties / inline / descriptionPrevious value: -"Return base64 images inline instead of writing to disk"New value: +"Return the image as an inline image block you can SEE, instead of a path (stdio) or link (hosted). The default costs nothing to carry and hands back a reference you can reuse; use this when you need to check the result yourself. On the hosted connector gemini_view_media does the same for an image you already have" - changed
Input schema / properties / max_wait_ms / descriptionPrevious value: -"Wait up to this many ms for the result; if generation is still running when the budget expires, return { job_id, status: \"running\" } immediately instead (poll gemini_get_result). Keeps fast results in-band while a slow batch can never trip the host tools/call timeout (-32001) — e.g. 20000 for multi-image sets. Ignored when async is set."New value: +"Wait up to this many ms in-band, then hand back { job_id, status: \"running\" } to poll with gemini_get_result (e.g. 20000 for multi-image sets). Keeps fast results inline and slow ones off the host timeout (-32001). Ignored when async is set." - changed
Input schema / properties / model / descriptionPrevious value: -"Model id override (default: server default; see gemini_list_models). gemini-3.1-flash-image (Nano Banana 2) is the versatile generalist workhorse — balances speed with state-of-the-art 4K generation, world knowledge, and reliable text rendering; excels at multi-reference-image processing and consistency. gemini-3-pro-image (Nano Banana Pro) is the premium choice for the most complex visual tasks — highest world knowledge, advanced localization, accurate brand consistency, precision creative control. gemini-3.1-flash-lite-image (Nano Banana 2 Lite) is the fastest/cheapest for simple tasks (1K only, no search grounding)."New value: +"Model id override (default: server default; see gemini_list_models). gemini-3.1-flash-image (Nano Banana 2, the default) is the generalist: fast, 4K, reliable text, strong multi-reference consistency. gemini-3-pro-image (Pro) is for the hardest work — best world knowledge, brand precision. gemini-3.1-flash-lite-image (Lite) is cheapest: 1K only, no search grounding." - changed
Input schema / properties / orientation / descriptionPrevious value: -"Shape of the output, in plain terms: \"landscape\" (wide, 16:9), \"portrait\" (tall, 9:16) or \"square\" (1:1). Use this for a request phrased as landscape/portrait/vertical/horizontal. For any other proportion — 35mm photo (3:2), print (4:3), social (4:5), cinematic (21:9) — name it with aspect_ratio instead, which overrides this when both are given."New value: +"Output shape in plain terms: landscape (16:9), portrait (9:16) or square (1:1). For any other proportion — 3:2, 4:3, 4:5, 21:9 — use aspect_ratio, which wins if both are given." - changed
Input schema / properties / style / descriptionPrevious value: -"Name of a saved style preset (see gemini_list_styles / gemini_save_style): its prompt fragment — and reference image, if it has one — is applied to the request automatically. Hosted connector only."New value: +"Name of a saved style preset (gemini_list_styles): its prompt fragment, and reference image if it has one, are applied automatically. Hosted connector only." - changed
Input schema / properties / timeout_ms / descriptionPrevious value: -"Upstream request timeout in ms for this call (default: $GEMINI_TIMEOUT_MS, else 60000 — or 120000 when image_size is 4K, which routinely runs past 60s)"New value: +"Upstream timeout in ms for this call (default $GEMINI_TIMEOUT_MS, else 60000 — 120000 at 4K, which runs past 60s)"
- Changed
gemini_image_generate13 fields changed- changed
Input schema / properties / async / descriptionPrevious value: -"Run in the background and return a job_id immediately instead of the image, so a long (Pro/4K) generation cannot hit the host tools/call timeout (-32001). Poll gemini_get_result with the job_id to fetch the result. PREFER `max_wait_ms` on the hosted connector: it runs where the executor is only guaranteed to stay alive while the request is open, so this option is served there as a bounded wait rather than an immediate hand-off."New value: +"Return a job_id immediately instead of the result, so a long generation cannot hit the host tools/call timeout (-32001); poll gemini_get_result. On the hosted connector prefer max_wait_ms — the executor only lives while a request is open, so async is served there as a bounded wait." - changed
Input schema / properties / characters / descriptionPrevious value: -"Names of saved characters (see gemini_list_characters / gemini_save_character): each one's reference image and description are attached to the request automatically, keeping recurring subjects consistent without re-sending anything. Hosted connector only."New value: +"Names of saved characters (gemini_list_characters): each one's reference image and description are attached automatically, keeping recurring subjects consistent without re-sending anything. Hosted connector only." - changed
Input schema / properties / idempotency_key / descriptionPrevious value: -"Opaque idempotency key: a repeat call with the same key returns the recorded result (reused: true) instead of billing a new generation. Set it when retrying after a host timeout (-32001) to avoid a duplicate charge."New value: +"Repeat calls with this key return the recorded result (reused: true) instead of billing a new generation. Set it when retrying after a host timeout (-32001)." - changed
Input schema / properties / images_base64 / descriptionPrevious value: -"Reference images as base64 strings or data URIs. Last resort: prefer images_url or images_file_uris, which keep image bytes out of the conversation"New value: +"Reference images as base64 strings or data URIs. Last resort — about 14k tokens per photo; prefer images_url or images_file_uris. The server uploads each one and reports a file_uri under image_inputs: pass that to images_file_uris next time instead of re-sending the bytes." - changed
Input schema / properties / images_file_uris / descriptionPrevious value: -"Reference images by Gemini Files API reference (\"files/<id>\", or the full uri) from gemini_upload_file or POST /upload. Upload once, then reference it across as many calls as you like — no bytes are re-sent and none enter the conversation. Files are retained ~48h, after which the reference stops resolving."New value: +"Reference images as Files API references (\"files/<id>\" or the full uri) from gemini_upload_file. Upload once and reuse across calls with no bytes in the conversation; retained ~48h, after which the reference stops resolving." - changed
Input schema / properties / images_r2_keys / descriptionPrevious value: -"Reference images by r2_key from THIS connector's store: a signed upload (gemini_get_upload_url → curl PUT) or an earlier generation's media[].r2_key. The server reads its own bucket directly — no bytes in the conversation, no signed URL, no ~48h Files API expiry. Hosted connector only."New value: +"Reference images by r2_key from this connector's store: a signed upload (gemini_get_upload_url → curl PUT) or an earlier generation's media[].r2_key. The server reads its own bucket — no bytes in the conversation, no ~48h expiry. Hosted connector only." - changed
Input schema / properties / images_url / descriptionPrevious value: -"Reference images as public https URLs — the SERVER downloads them, so no image bytes travel through the conversation. Preferred over images_base64, which costs ~14k tokens per photo and breaks if a file read was truncated. Max 15MB each; must be a directly-linked image (Content-Type image/*)."New value: +"Reference images as public https URLs — the server downloads them, so no image bytes cross the conversation. Preferred over images_base64, which costs ~14k tokens per photo. Max 15MB each, Content-Type image/*." - changed
Input schema / properties / inline / descriptionPrevious value: -"Return base64 images inline instead of writing to disk"New value: +"Return the image as an inline image block you can SEE, instead of a path (stdio) or link (hosted). The default costs nothing to carry and hands back a reference you can reuse; use this when you need to check the result yourself. On the hosted connector gemini_view_media does the same for an image you already have" - changed
Input schema / properties / max_wait_ms / descriptionPrevious value: -"Wait up to this many ms for the result; if generation is still running when the budget expires, return { job_id, status: \"running\" } immediately instead (poll gemini_get_result). Keeps fast results in-band while a slow batch can never trip the host tools/call timeout (-32001) — e.g. 20000 for multi-image sets. Ignored when async is set."New value: +"Wait up to this many ms in-band, then hand back { job_id, status: \"running\" } to poll with gemini_get_result (e.g. 20000 for multi-image sets). Keeps fast results inline and slow ones off the host timeout (-32001). Ignored when async is set." - changed
Input schema / properties / model / descriptionPrevious value: -"Model id override (default: server default; see gemini_list_models). gemini-3.1-flash-image (Nano Banana 2) is the versatile generalist workhorse — balances speed with state-of-the-art 4K generation, world knowledge, and reliable text rendering; excels at multi-reference-image processing and consistency. gemini-3-pro-image (Nano Banana Pro) is the premium choice for the most complex visual tasks — highest world knowledge, advanced localization, accurate brand consistency, precision creative control. gemini-3.1-flash-lite-image (Nano Banana 2 Lite) is the fastest/cheapest for simple tasks (1K only, no search grounding)."New value: +"Model id override (default: server default; see gemini_list_models). gemini-3.1-flash-image (Nano Banana 2, the default) is the generalist: fast, 4K, reliable text, strong multi-reference consistency. gemini-3-pro-image (Pro) is for the hardest work — best world knowledge, brand precision. gemini-3.1-flash-lite-image (Lite) is cheapest: 1K only, no search grounding." - changed
Input schema / properties / orientation / descriptionPrevious value: -"Shape of the output, in plain terms: \"landscape\" (wide, 16:9), \"portrait\" (tall, 9:16) or \"square\" (1:1). Use this for a request phrased as landscape/portrait/vertical/horizontal. For any other proportion — 35mm photo (3:2), print (4:3), social (4:5), cinematic (21:9) — name it with aspect_ratio instead, which overrides this when both are given."New value: +"Output shape in plain terms: landscape (16:9), portrait (9:16) or square (1:1). For any other proportion — 3:2, 4:3, 4:5, 21:9 — use aspect_ratio, which wins if both are given." - changed
Input schema / properties / style / descriptionPrevious value: -"Name of a saved style preset (see gemini_list_styles / gemini_save_style): its prompt fragment — and reference image, if it has one — is applied to the request automatically. Hosted connector only."New value: +"Name of a saved style preset (gemini_list_styles): its prompt fragment, and reference image if it has one, are applied automatically. Hosted connector only." - changed
Input schema / properties / timeout_ms / descriptionPrevious value: -"Upstream request timeout in ms for this call (default: $GEMINI_TIMEOUT_MS, else 60000 — or 120000 when image_size is 4K, which routinely runs past 60s)"New value: +"Upstream timeout in ms for this call (default $GEMINI_TIMEOUT_MS, else 60000 — 120000 at 4K, which runs past 60s)"
- Changed
gemini_image_set13 fields changed- changed
Input schema / properties / async / descriptionPrevious value: -"Run in the background and return a job_id immediately instead of the image, so a long (Pro/4K) generation cannot hit the host tools/call timeout (-32001). Poll gemini_get_result with the job_id to fetch the result. PREFER `max_wait_ms` on the hosted connector: it runs where the executor is only guaranteed to stay alive while the request is open, so this option is served there as a bounded wait rather than an immediate hand-off."New value: +"Return a job_id immediately instead of the result, so a long generation cannot hit the host tools/call timeout (-32001); poll gemini_get_result. On the hosted connector prefer max_wait_ms — the executor only lives while a request is open, so async is served there as a bounded wait." - changed
Input schema / properties / characters / descriptionPrevious value: -"Names of saved characters (see gemini_list_characters / gemini_save_character): each one's reference image and description are attached to the request automatically, keeping recurring subjects consistent without re-sending anything. Hosted connector only."New value: +"Names of saved characters (gemini_list_characters): each one's reference image and description are attached automatically, keeping recurring subjects consistent without re-sending anything. Hosted connector only." - changed
Input schema / properties / idempotency_key / descriptionPrevious value: -"Opaque idempotency key: a repeat call with the same key returns the recorded result (reused: true) instead of billing a new generation. Set it when retrying after a host timeout (-32001) to avoid a duplicate charge."New value: +"Repeat calls with this key return the recorded result (reused: true) instead of billing a new generation. Set it when retrying after a host timeout (-32001)." - changed
Input schema / properties / inline / descriptionPrevious value: -"Return base64 images inline instead of writing to disk"New value: +"Return the image as an inline image block you can SEE, instead of a path (stdio) or link (hosted). The default costs nothing to carry and hands back a reference you can reuse; use this when you need to check the result yourself. On the hosted connector gemini_view_media does the same for an image you already have" - changed
Input schema / properties / master_images_base64 / descriptionPrevious value: -"Reference images as base64 strings or data URIs for master generation. Last resort: prefer master_images_url or master_images_file_uris, which keep image bytes out of the conversation"New value: +"Reference images for the master generation as base64 strings or data URIs. Last resort — about 14k tokens per photo; prefer master_images_url or master_images_file_uris. The server uploads each one and reports a file_uri under image_inputs: pass that to master_images_file_uris next time instead of re-sending the bytes." - changed
Input schema / properties / master_images_file_uris / descriptionPrevious value: -"Reference images passed to the master AND to every scene call by Gemini Files API reference (\"files/<id>\", or the full uri) from gemini_upload_file or POST /upload. Upload once, then reference it across as many calls as you like — no bytes are re-sent and none enter the conversation. Files are retained ~48h, after which the reference stops resolving."New value: +"Reference images passed to the master AND to every scene call as Files API references (\"files/<id>\" or the full uri) from gemini_upload_file. Upload once and reuse across calls with no bytes in the conversation; retained ~48h, after which the reference stops resolving." - changed
Input schema / properties / master_images_r2_keys / descriptionPrevious value: -"Reference images passed to the master AND to every scene call by r2_key from THIS connector's store: a signed upload (gemini_get_upload_url → curl PUT) or an earlier generation's media[].r2_key. The server reads its own bucket directly — no bytes in the conversation, no signed URL, no ~48h Files API expiry. Hosted connector only."New value: +"Reference images passed to the master AND to every scene call by r2_key from this connector's store: a signed upload (gemini_get_upload_url → curl PUT) or an earlier generation's media[].r2_key. The server reads its own bucket — no bytes in the conversation, no ~48h expiry. Hosted connector only." - changed
Input schema / properties / master_images_url / descriptionPrevious value: -"Reference images passed to the master AND to every scene call (fetched once) as public https URLs — the SERVER downloads them, so no image bytes travel through the conversation. Preferred over images_base64, which costs ~14k tokens per photo and breaks if a file read was truncated. Max 15MB each; must be a directly-linked image (Content-Type image/*)."New value: +"Reference images passed to the master AND to every scene call (fetched once) as public https URLs — the server downloads them, so no image bytes cross the conversation. Preferred over images_base64, which costs ~14k tokens per photo. Max 15MB each, Content-Type image/*." - changed
Input schema / properties / max_wait_ms / descriptionPrevious value: -"Wait up to this many ms for the result; if generation is still running when the budget expires, return { job_id, status: \"running\" } immediately instead (poll gemini_get_result). Keeps fast results in-band while a slow batch can never trip the host tools/call timeout (-32001) — e.g. 20000 for multi-image sets. Ignored when async is set."New value: +"Wait up to this many ms in-band, then hand back { job_id, status: \"running\" } to poll with gemini_get_result (e.g. 20000 for multi-image sets). Keeps fast results inline and slow ones off the host timeout (-32001). Ignored when async is set." - changed
Input schema / properties / model / descriptionPrevious value: -"Model id override (default: server default; see gemini_list_models). gemini-3.1-flash-image (Nano Banana 2) is the versatile generalist workhorse — balances speed with state-of-the-art 4K generation, world knowledge, and reliable text rendering; excels at multi-reference-image processing and consistency. gemini-3-pro-image (Nano Banana Pro) is the premium choice for the most complex visual tasks — highest world knowledge, advanced localization, accurate brand consistency, precision creative control. gemini-3.1-flash-lite-image (Nano Banana 2 Lite) is the fastest/cheapest for simple tasks (1K only, no search grounding)."New value: +"Model id override (default: server default; see gemini_list_models). gemini-3.1-flash-image (Nano Banana 2, the default) is the generalist: fast, 4K, reliable text, strong multi-reference consistency. gemini-3-pro-image (Pro) is for the hardest work — best world knowledge, brand precision. gemini-3.1-flash-lite-image (Lite) is cheapest: 1K only, no search grounding." - changed
Input schema / properties / orientation / descriptionPrevious value: -"Shape of the output, in plain terms: \"landscape\" (wide, 16:9), \"portrait\" (tall, 9:16) or \"square\" (1:1). Use this for a request phrased as landscape/portrait/vertical/horizontal. For any other proportion — 35mm photo (3:2), print (4:3), social (4:5), cinematic (21:9) — name it with aspect_ratio instead, which overrides this when both are given."New value: +"Output shape in plain terms: landscape (16:9), portrait (9:16) or square (1:1). For any other proportion — 3:2, 4:3, 4:5, 21:9 — use aspect_ratio, which wins if both are given." - changed
Input schema / properties / style / descriptionPrevious value: -"Name of a saved style preset (see gemini_list_styles / gemini_save_style): its prompt fragment — and reference image, if it has one — is applied to the request automatically. Hosted connector only."New value: +"Name of a saved style preset (gemini_list_styles): its prompt fragment, and reference image if it has one, are applied automatically. Hosted connector only." - changed
Input schema / properties / timeout_ms / descriptionPrevious value: -"Upstream request timeout in ms for this call (default: $GEMINI_TIMEOUT_MS, else 60000 — or 120000 when image_size is 4K, which routinely runs past 60s)"New value: +"Upstream timeout in ms for this call (default $GEMINI_TIMEOUT_MS, else 60000 — 120000 at 4K, which runs past 60s)"
- Changed
gemini_interact9 fields changed- changed
Input schema / properties / async / descriptionPrevious value: -"Run in the background and return a job_id immediately instead of the image, so a long (Pro/4K) generation cannot hit the host tools/call timeout (-32001). Poll gemini_get_result with the job_id to fetch the result. PREFER `max_wait_ms` on the hosted connector: it runs where the executor is only guaranteed to stay alive while the request is open, so this option is served there as a bounded wait rather than an immediate hand-off."New value: +"Return a job_id immediately instead of the result, so a long generation cannot hit the host tools/call timeout (-32001); poll gemini_get_result. On the hosted connector prefer max_wait_ms — the executor only lives while a request is open, so async is served there as a bounded wait." - changed
Input schema / properties / idempotency_key / descriptionPrevious value: -"Opaque idempotency key: a repeat call with the same key returns the recorded result (reused: true) instead of billing a new generation. Set it when retrying after a host timeout (-32001) to avoid a duplicate charge."New value: +"Repeat calls with this key return the recorded result (reused: true) instead of billing a new generation. Set it when retrying after a host timeout (-32001)." - changed
Input schema / properties / images_base64 / descriptionPrevious value: -"Reference images as base64 strings or data URIs. Last resort: prefer images_url or images_file_uris, which keep image bytes out of the conversation. NEW reference images only (e.g. a style or target photo). When chaining with previous_interaction_id, do NOT re-attach the prior turn's output — the interaction already contains it, and re-sending it anchors the model against your edit."New value: +"Reference images. NEW reference images only (e.g. a style or target photo). When chaining with previous_interaction_id, do NOT re-attach the prior turn's output — the interaction already contains it, and re-sending it anchors the model against your edit. Given as base64 strings or data URIs. Last resort — about 14k tokens per photo; prefer images_url or images_file_uris. The server uploads each one and reports a file_uri under image_inputs: pass that to images_file_uris next time instead of re-sending the bytes." - changed
Input schema / properties / images_file_uris / descriptionPrevious value: -"Reference images. NEW reference images only (e.g. a style or target photo). When chaining with previous_interaction_id, do NOT re-attach the prior turn's output — the interaction already contains it, and re-sending it anchors the model against your edit. Given by Gemini Files API reference (\"files/<id>\", or the full uri) from gemini_upload_file or POST /upload. Upload once, then reference it across as many calls as you like — no bytes are re-sent and none enter the conversation. Files are retained ~48h, after which the reference stops resolving."New value: +"Reference images. NEW reference images only (e.g. a style or target photo). When chaining with previous_interaction_id, do NOT re-attach the prior turn's output — the interaction already contains it, and re-sending it anchors the model against your edit. Given as Files API references (\"files/<id>\" or the full uri) from gemini_upload_file. Upload once and reuse across calls with no bytes in the conversation; retained ~48h, after which the reference stops resolving." - changed
Input schema / properties / images_url / descriptionPrevious value: -"Reference images. NEW reference images only (e.g. a style or target photo). When chaining with previous_interaction_id, do NOT re-attach the prior turn's output — the interaction already contains it, and re-sending it anchors the model against your edit. Given as public https URLs — the SERVER downloads them, so no image bytes travel through the conversation. Preferred over images_base64, which costs ~14k tokens per photo and breaks if a file read was truncated. Max 15MB each; must be a directly-linked image (Content-Type image/*)."New value: +"Reference images. NEW reference images only (e.g. a style or target photo). When chaining with previous_interaction_id, do NOT re-attach the prior turn's output — the interaction already contains it, and re-sending it anchors the model against your edit. Given as public https URLs — the server downloads them, so no image bytes cross the conversation. Preferred over images_base64, which costs ~14k tokens per photo. Max 15MB each, Content-Type image/*." - changed
Input schema / properties / max_wait_ms / descriptionPrevious value: -"Wait up to this many ms for the result; if generation is still running when the budget expires, return { job_id, status: \"running\" } immediately instead (poll gemini_get_result). Keeps fast results in-band while a slow batch can never trip the host tools/call timeout (-32001) — e.g. 20000 for multi-image sets. Ignored when async is set."New value: +"Wait up to this many ms in-band, then hand back { job_id, status: \"running\" } to poll with gemini_get_result (e.g. 20000 for multi-image sets). Keeps fast results inline and slow ones off the host timeout (-32001). Ignored when async is set." - changed
Input schema / properties / model / descriptionPrevious value: -"Model id override (default: server default; see gemini_list_models). gemini-3.1-flash-image (Nano Banana 2) is the versatile generalist workhorse — balances speed with state-of-the-art 4K generation, world knowledge, and reliable text rendering; excels at multi-reference-image processing and consistency. gemini-3-pro-image (Nano Banana Pro) is the premium choice for the most complex visual tasks — highest world knowledge, advanced localization, accurate brand consistency, precision creative control. gemini-3.1-flash-lite-image (Nano Banana 2 Lite) is the fastest/cheapest for simple tasks (1K only, no search grounding)."New value: +"Model id override (default: server default; see gemini_list_models). gemini-3.1-flash-image (Nano Banana 2, the default) is the generalist: fast, 4K, reliable text, strong multi-reference consistency. gemini-3-pro-image (Pro) is for the hardest work — best world knowledge, brand precision. gemini-3.1-flash-lite-image (Lite) is cheapest: 1K only, no search grounding." - changed
Input schema / properties / orientation / descriptionPrevious value: -"Shape of the output, in plain terms: \"landscape\" (wide, 16:9), \"portrait\" (tall, 9:16) or \"square\" (1:1). Use this for a request phrased as landscape/portrait/vertical/horizontal. For any other proportion — 35mm photo (3:2), print (4:3), social (4:5), cinematic (21:9) — name it with aspect_ratio instead, which overrides this when both are given."New value: +"Output shape in plain terms: landscape (16:9), portrait (9:16) or square (1:1). For any other proportion — 3:2, 4:3, 4:5, 21:9 — use aspect_ratio, which wins if both are given." - changed
Input schema / properties / timeout_ms / descriptionPrevious value: -"Upstream request timeout in ms for this call (default: $GEMINI_TIMEOUT_MS, else 60000 — or 120000 when image_size is 4K, which routinely runs past 60s)"New value: +"Upstream timeout in ms for this call (default $GEMINI_TIMEOUT_MS, else 60000 — 120000 at 4K, which runs past 60s)"
- Changed
gemini_music_generate13 fields changed- changed
Input schema / properties / async / descriptionPrevious value: -"Run in the background and return a job_id immediately instead of the image, so a long (Pro/4K) generation cannot hit the host tools/call timeout (-32001). Poll gemini_get_result with the job_id to fetch the result. PREFER `max_wait_ms` on the hosted connector: it runs where the executor is only guaranteed to stay alive while the request is open, so this option is served there as a bounded wait rather than an immediate hand-off."New value: +"Return a job_id immediately instead of the result, so a long generation cannot hit the host tools/call timeout (-32001); poll gemini_get_result. On the hosted connector prefer max_wait_ms — the executor only lives while a request is open, so async is served there as a bounded wait." - removed
Input schema / properties / audio_formatRemoved value: -{ - "description": "Output format (default mp3). wav is lyria-3-pro-preview-only.", - "enum": [ - "mp3", - "wav" - ], - "type": "string" -} - removed
Input schema / properties / continue_lastRemoved value: -{ - "description": "Continue from the most recent music interaction this server created (explicit previous_interaction_id wins)", - "type": "boolean" -} - changed
Input schema / properties / idempotency_key / descriptionPrevious value: -"Opaque idempotency key: a repeat call with the same key returns the recorded result (reused: true) instead of billing a new generation. Set it when retrying after a host timeout (-32001) to avoid a duplicate charge."New value: +"Repeat calls with this key return the recorded result (reused: true) instead of billing a new generation. Set it when retrying after a host timeout (-32001)." - changed
Input schema / properties / images_base64 / descriptionPrevious value: -"Reference images as base64 strings or data URIs. Last resort: prefer images_url or images_file_uris, which keep image bytes out of the conversation"New value: +"Reference images as base64 strings or data URIs. Last resort — about 14k tokens per photo; prefer images_url or images_file_uris. The server uploads each one and reports a file_uri under image_inputs: pass that to images_file_uris next time instead of re-sending the bytes." - changed
Input schema / properties / images_file_uris / descriptionPrevious value: -"Reference images by Gemini Files API reference (\"files/<id>\", or the full uri) from gemini_upload_file or POST /upload. Upload once, then reference it across as many calls as you like — no bytes are re-sent and none enter the conversation. Files are retained ~48h, after which the reference stops resolving."New value: +"Reference images as Files API references (\"files/<id>\" or the full uri) from gemini_upload_file. Upload once and reuse across calls with no bytes in the conversation; retained ~48h, after which the reference stops resolving." - changed
Input schema / properties / images_url / descriptionPrevious value: -"Reference images as public https URLs — the SERVER downloads them, so no image bytes travel through the conversation. Preferred over images_base64, which costs ~14k tokens per photo and breaks if a file read was truncated. Max 15MB each; must be a directly-linked image (Content-Type image/*)."New value: +"Reference images as public https URLs — the server downloads them, so no image bytes cross the conversation. Preferred over images_base64, which costs ~14k tokens per photo. Max 15MB each, Content-Type image/*." - changed
Input schema / properties / inline / descriptionPrevious value: -"Return base64 audio inline instead of writing to disk"New value: +"Return base64 audio inline instead of writing to disk. A 30s track is ~1.9MB of base64 — the default hands back a path instead" - changed
Input schema / properties / max_wait_ms / descriptionPrevious value: -"Wait up to this many ms for the result; if generation is still running when the budget expires, return { job_id, status: \"running\" } immediately instead (poll gemini_get_result). Keeps fast results in-band while a slow batch can never trip the host tools/call timeout (-32001) — e.g. 20000 for multi-image sets. Ignored when async is set."New value: +"Wait up to this many ms in-band, then hand back { job_id, status: \"running\" } to poll with gemini_get_result (e.g. 20000 for multi-image sets). Keeps fast results inline and slow ones off the host timeout (-32001). Ignored when async is set." - changed
Input schema / properties / model / descriptionPrevious value: -"Lyria model (default: lyria-3-clip-preview). Pro is longer-form and supports WAV."New value: +"Lyria model (default: lyria-3-clip-preview — 30s, $0.04). lyria-3.5 and lyria-3-pro-preview run minutes-long at $0.08." - changed
Input schema / properties / model / enumPrevious value: -[ - "lyria-3-clip-preview", - "lyria-3-pro-preview" -]New value: +[ + "lyria-3-clip-preview", + "lyria-3.5", + "lyria-3-pro-preview" +] - removed
Input schema / properties / previous_interaction_idRemoved value: -{ - "description": "Interaction id to continue from", - "type": "string" -} - changed
Input schema / properties / timeout_ms / descriptionPrevious value: -"Upstream request timeout in ms for this call (default: $GEMINI_TIMEOUT_MS, else 60000 — or 120000 when image_size is 4K, which routinely runs past 60s)"New value: +"Upstream timeout in ms for this call (default $GEMINI_TIMEOUT_MS, else 60000 — 120000 at 4K, which runs past 60s)"
- Changed
gemini_video_generate12 fields changed- changed
Input schema / properties / async / descriptionPrevious value: -"Run in the background and return a job_id immediately instead of the image, so a long (Pro/4K) generation cannot hit the host tools/call timeout (-32001). Poll gemini_get_result with the job_id to fetch the result. PREFER `max_wait_ms` on the hosted connector: it runs where the executor is only guaranteed to stay alive while the request is open, so this option is served there as a bounded wait rather than an immediate hand-off."New value: +"Return a job_id immediately instead of the result, so a long generation cannot hit the host tools/call timeout (-32001); poll gemini_get_result. On the hosted connector prefer max_wait_ms — the executor only lives while a request is open, so async is served there as a bounded wait." - changed
Input schema / properties / idempotency_key / descriptionPrevious value: -"Opaque idempotency key: a repeat call with the same key returns the recorded result (reused: true) instead of billing a new generation. Set it when retrying after a host timeout (-32001) to avoid a duplicate charge."New value: +"Repeat calls with this key return the recorded result (reused: true) instead of billing a new generation. Set it when retrying after a host timeout (-32001)." - changed
Input schema / properties / images_base64 / descriptionPrevious value: -"Reference images as base64 strings or data URIs. Last resort: prefer images_url or images_file_uris, which keep image bytes out of the conversation"New value: +"Reference images as base64 strings or data URIs. Last resort — about 14k tokens per photo; prefer images_url or images_file_uris. The server uploads each one and reports a file_uri under image_inputs: pass that to images_file_uris next time instead of re-sending the bytes." - changed
Input schema / properties / images_file_uris / descriptionPrevious value: -"Reference stills by Gemini Files API reference (\"files/<id>\", or the full uri) from gemini_upload_file or POST /upload. Upload once, then reference it across as many calls as you like — no bytes are re-sent and none enter the conversation. Files are retained ~48h, after which the reference stops resolving."New value: +"Reference stills as Files API references (\"files/<id>\" or the full uri) from gemini_upload_file. Upload once and reuse across calls with no bytes in the conversation; retained ~48h, after which the reference stops resolving." - changed
Input schema / properties / images_url / descriptionPrevious value: -"Reference stills as public https URLs — the SERVER downloads them, so no image bytes travel through the conversation. Preferred over images_base64, which costs ~14k tokens per photo and breaks if a file read was truncated. Max 15MB each; must be a directly-linked image (Content-Type image/*)."New value: +"Reference stills as public https URLs — the server downloads them, so no image bytes cross the conversation. Preferred over images_base64, which costs ~14k tokens per photo. Max 15MB each, Content-Type image/*." - changed
Input schema / properties / max_wait_ms / descriptionPrevious value: -"Wait up to this many ms for the result; if generation is still running when the budget expires, return { job_id, status: \"running\" } immediately instead (poll gemini_get_result). Keeps fast results in-band while a slow batch can never trip the host tools/call timeout (-32001) — e.g. 20000 for multi-image sets. Ignored when async is set."New value: +"Wait up to this many ms in-band, then hand back { job_id, status: \"running\" } to poll with gemini_get_result (e.g. 20000 for multi-image sets). Keeps fast results inline and slow ones off the host timeout (-32001). Ignored when async is set." - changed
Input schema / properties / model / descriptionPrevious value: -"Model id override (default: gemini-omni-flash-preview)"New value: +"Model id override (default: gemini-omni-1.1-flash)" - changed
Input schema / properties / orientation / descriptionPrevious value: -"Shape of the output, in plain terms: \"landscape\" (wide, 16:9), \"portrait\" (tall, 9:16) or \"square\" (1:1). Use this for a request phrased as landscape/portrait/vertical/horizontal. For any other proportion — 35mm photo (3:2), print (4:3), social (4:5), cinematic (21:9) — name it with aspect_ratio instead, which overrides this when both are given."New value: +"Output shape in plain terms: landscape (16:9), portrait (9:16) or square (1:1). For any other proportion — 3:2, 4:3, 4:5, 21:9 — use aspect_ratio, which wins if both are given." - added
Input schema / properties / resolutionAdded value: +{ + "description": "Output resolution (default 720p). Video is billed per output token, so 360p costs roughly a third of 720p — use it for drafts", + "enum": [ + "360p", + "720p", + "1080p", + "4k" + ], + "type": "string" +} - changed
Input schema / properties / task / descriptionPrevious value: -"text_to_video (default), image_to_video / reference_to_video (need image input), or edit (needs previous_interaction_id)"New value: +"text_to_video (default), image_to_video / reference_to_video (need image input), or edit / extend (need previous_interaction_id)" - changed
Input schema / properties / task / enumPrevious value: -[ - "text_to_video", - "image_to_video", - "reference_to_video", - "edit" -]New value: +[ + "text_to_video", + "image_to_video", + "reference_to_video", + "edit", + "extend" +] - changed
Input schema / properties / timeout_ms / descriptionPrevious value: -"Upstream request timeout in ms for this call (default: $GEMINI_TIMEOUT_MS, else 60000 — or 120000 when image_size is 4K, which routinely runs past 60s)"New value: +"Upstream timeout in ms for this call (default $GEMINI_TIMEOUT_MS, else 60000 — 120000 at 4K, which runs past 60s)"
1 tool update
v1.12.0- Added
gemini_healthcheck
1 tool update
v1.11.1- Added
gemini_token_usage
7 tool updates
v1.10.0- Changed
gemini_get_result1 field changed- added
Input schema / properties / output_dirAdded value: +{ + "description": "Where to write media recovered from a killed job (default: $GEMINI_OUTPUT_DIR or cwd)", + "type": "string" +}
- Changed
gemini_image_edit2 fields changed- changed
Input schema / properties / aspect_ratio / descriptionPrevious value: -"Output aspect ratio"New value: +"Exact output aspect ratio. For a plain landscape/portrait/square request, `orientation` is the shorthand; this wins if both are given." - added
Input schema / properties / orientationAdded value: +{ + "description": "Shape of the output, in plain terms: \"landscape\" (wide, 16:9), \"portrait\" (tall, 9:16) or \"square\" (1:1). Use this for a request phrased as landscape/portrait/vertical/horizontal. For any other proportion — 35mm photo (3:2), print (4:3), social (4:5), cinematic (21:9) — name it with aspect_ratio instead, which overrides this when both are given.", + "enum": [ + "landscape", + "portrait", + "square" + ], + "type": "string" +}
- Changed
gemini_image_generate2 fields changed- changed
Input schema / properties / aspect_ratio / descriptionPrevious value: -"Output aspect ratio"New value: +"Exact output aspect ratio. For a plain landscape/portrait/square request, `orientation` is the shorthand; this wins if both are given." - added
Input schema / properties / orientationAdded value: +{ + "description": "Shape of the output, in plain terms: \"landscape\" (wide, 16:9), \"portrait\" (tall, 9:16) or \"square\" (1:1). Use this for a request phrased as landscape/portrait/vertical/horizontal. For any other proportion — 35mm photo (3:2), print (4:3), social (4:5), cinematic (21:9) — name it with aspect_ratio instead, which overrides this when both are given.", + "enum": [ + "landscape", + "portrait", + "square" + ], + "type": "string" +}
- Changed
gemini_image_set2 fields changed- changed
Input schema / properties / aspect_ratio / descriptionPrevious value: -"Output aspect ratio"New value: +"Exact output aspect ratio. For a plain landscape/portrait/square request, `orientation` is the shorthand; this wins if both are given." - added
Input schema / properties / orientationAdded value: +{ + "description": "Shape of the output, in plain terms: \"landscape\" (wide, 16:9), \"portrait\" (tall, 9:16) or \"square\" (1:1). Use this for a request phrased as landscape/portrait/vertical/horizontal. For any other proportion — 35mm photo (3:2), print (4:3), social (4:5), cinematic (21:9) — name it with aspect_ratio instead, which overrides this when both are given.", + "enum": [ + "landscape", + "portrait", + "square" + ], + "type": "string" +}
- Changed
gemini_interact2 fields changed- changed
Input schema / properties / aspect_ratio / descriptionPrevious value: -"Output aspect ratio"New value: +"Exact output aspect ratio. For a plain landscape/portrait/square request, `orientation` is the shorthand; this wins if both are given." - added
Input schema / properties / orientationAdded value: +{ + "description": "Shape of the output, in plain terms: \"landscape\" (wide, 16:9), \"portrait\" (tall, 9:16) or \"square\" (1:1). Use this for a request phrased as landscape/portrait/vertical/horizontal. For any other proportion — 35mm photo (3:2), print (4:3), social (4:5), cinematic (21:9) — name it with aspect_ratio instead, which overrides this when both are given.", + "enum": [ + "landscape", + "portrait", + "square" + ], + "type": "string" +}
- Changed
gemini_music_generate1 field changed- added
Input schema / properties / backgroundAdded value: +{ + "description": "Run the generation on Google's side and poll it, so a killed job can be recovered by gemini_get_result. Off by default — see gemini_video_generate", + "type": "boolean" +}
- Changed
gemini_video_generate4 fields changed- changed
Input schema / properties / aspect_ratio / descriptionPrevious value: -"Output aspect ratio (omni: 16:9 or 9:16)"New value: +"Exact output aspect ratio (omni: 16:9 or 9:16). `orientation` is the plain-language shorthand; this wins if both are given." - added
Input schema / properties / backgroundAdded value: +{ + "description": "Run the generation on Google's side and poll it, so a killed job can be recovered by gemini_get_result. Trade-off: retrieving a backgrounded interaction is unreliable today (some become permanently unreadable), so the default is off", + "type": "boolean" +} - added
Input schema / properties / deliveryAdded value: +{ + "description": "How the clip comes back: \"uri\" (default — a Files API link the server downloads, no size ceiling) or \"inline\" (base64, capped ~4MB)", + "enum": [ + "inline", + "uri" + ], + "type": "string" +} - added
Input schema / properties / orientationAdded value: +{ + "description": "Shape of the output, in plain terms: \"landscape\" (wide, 16:9), \"portrait\" (tall, 9:16) or \"square\" (1:1). Use this for a request phrased as landscape/portrait/vertical/horizontal. For any other proportion — 35mm photo (3:2), print (4:3), social (4:5), cinematic (21:9) — name it with aspect_ratio instead, which overrides this when both are given.", + "enum": [ + "landscape", + "portrait", + "square" + ], + "type": "string" +}
6 tool updates
v1.7.0- Changed
gemini_image_edit5 fields changed- changed
Input schema / properties / async / descriptionPrevious value: -"Run in the background and return a job_id immediately instead of the image, so a long (Pro/4K) generation cannot hit the host tools/call timeout (-32001). Poll gemini_get_result with the job_id to fetch the result (jobs are per-process and expire ~10 min after completion)."New value: +"Run in the background and return a job_id immediately instead of the image, so a long (Pro/4K) generation cannot hit the host tools/call timeout (-32001). Poll gemini_get_result with the job_id to fetch the result. PREFER `max_wait_ms` on the hosted connector: it runs where the executor is only guaranteed to stay alive while the request is open, so this option is served there as a bounded wait rather than an immediate hand-off." - added
Input schema / properties / charactersAdded value: +{ + "description": "Names of saved characters (see gemini_list_characters / gemini_save_character): each one's reference image and description are attached to the request automatically, keeping recurring subjects consistent without re-sending anything. Hosted connector only.", + "items": { + "minLength": 1, + "type": "string" + }, + "maxItems": 8, + "type": "array" +} - added
Input schema / properties / images_r2_keysAdded value: +{ + "description": "Input images by r2_key from THIS connector's store: a signed upload (gemini_get_upload_url → curl PUT) or an earlier generation's media[].r2_key. The server reads its own bucket directly — no bytes in the conversation, no signed URL, no ~48h Files API expiry. Hosted connector only.", + "items": { + "minLength": 1, + "type": "string" + }, + "type": "array" +} - added
Input schema / properties / max_wait_msAdded value: +{ + "description": "Wait up to this many ms for the result; if generation is still running when the budget expires, return { job_id, status: \"running\" } immediately instead (poll gemini_get_result). Keeps fast results in-band while a slow batch can never trip the host tools/call timeout (-32001) — e.g. 20000 for multi-image sets. Ignored when async is set.", + "exclusiveMinimum": 0, + "maximum": 600000, + "type": "integer" +} - added
Input schema / properties / styleAdded value: +{ + "description": "Name of a saved style preset (see gemini_list_styles / gemini_save_style): its prompt fragment — and reference image, if it has one — is applied to the request automatically. Hosted connector only.", + "minLength": 1, + "type": "string" +}
- Changed
gemini_image_generate5 fields changed- changed
Input schema / properties / async / descriptionPrevious value: -"Run in the background and return a job_id immediately instead of the image, so a long (Pro/4K) generation cannot hit the host tools/call timeout (-32001). Poll gemini_get_result with the job_id to fetch the result (jobs are per-process and expire ~10 min after completion)."New value: +"Run in the background and return a job_id immediately instead of the image, so a long (Pro/4K) generation cannot hit the host tools/call timeout (-32001). Poll gemini_get_result with the job_id to fetch the result. PREFER `max_wait_ms` on the hosted connector: it runs where the executor is only guaranteed to stay alive while the request is open, so this option is served there as a bounded wait rather than an immediate hand-off." - added
Input schema / properties / charactersAdded value: +{ + "description": "Names of saved characters (see gemini_list_characters / gemini_save_character): each one's reference image and description are attached to the request automatically, keeping recurring subjects consistent without re-sending anything. Hosted connector only.", + "items": { + "minLength": 1, + "type": "string" + }, + "maxItems": 8, + "type": "array" +} - added
Input schema / properties / images_r2_keysAdded value: +{ + "description": "Reference images by r2_key from THIS connector's store: a signed upload (gemini_get_upload_url → curl PUT) or an earlier generation's media[].r2_key. The server reads its own bucket directly — no bytes in the conversation, no signed URL, no ~48h Files API expiry. Hosted connector only.", + "items": { + "minLength": 1, + "type": "string" + }, + "type": "array" +} - added
Input schema / properties / max_wait_msAdded value: +{ + "description": "Wait up to this many ms for the result; if generation is still running when the budget expires, return { job_id, status: \"running\" } immediately instead (poll gemini_get_result). Keeps fast results in-band while a slow batch can never trip the host tools/call timeout (-32001) — e.g. 20000 for multi-image sets. Ignored when async is set.", + "exclusiveMinimum": 0, + "maximum": 600000, + "type": "integer" +} - added
Input schema / properties / styleAdded value: +{ + "description": "Name of a saved style preset (see gemini_list_styles / gemini_save_style): its prompt fragment — and reference image, if it has one — is applied to the request automatically. Hosted connector only.", + "minLength": 1, + "type": "string" +}
- Changed
gemini_image_set5 fields changed- changed
Input schema / properties / async / descriptionPrevious value: -"Run in the background and return a job_id immediately instead of the image, so a long (Pro/4K) generation cannot hit the host tools/call timeout (-32001). Poll gemini_get_result with the job_id to fetch the result (jobs are per-process and expire ~10 min after completion)."New value: +"Run in the background and return a job_id immediately instead of the image, so a long (Pro/4K) generation cannot hit the host tools/call timeout (-32001). Poll gemini_get_result with the job_id to fetch the result. PREFER `max_wait_ms` on the hosted connector: it runs where the executor is only guaranteed to stay alive while the request is open, so this option is served there as a bounded wait rather than an immediate hand-off." - added
Input schema / properties / charactersAdded value: +{ + "description": "Names of saved characters (see gemini_list_characters / gemini_save_character): each one's reference image and description are attached to the request automatically, keeping recurring subjects consistent without re-sending anything. Hosted connector only.", + "items": { + "minLength": 1, + "type": "string" + }, + "maxItems": 8, + "type": "array" +} - added
Input schema / properties / master_images_r2_keysAdded value: +{ + "description": "Reference images passed to the master AND to every scene call by r2_key from THIS connector's store: a signed upload (gemini_get_upload_url → curl PUT) or an earlier generation's media[].r2_key. The server reads its own bucket directly — no bytes in the conversation, no signed URL, no ~48h Files API expiry. Hosted connector only.", + "items": { + "minLength": 1, + "type": "string" + }, + "type": "array" +} - added
Input schema / properties / max_wait_msAdded value: +{ + "description": "Wait up to this many ms for the result; if generation is still running when the budget expires, return { job_id, status: \"running\" } immediately instead (poll gemini_get_result). Keeps fast results in-band while a slow batch can never trip the host tools/call timeout (-32001) — e.g. 20000 for multi-image sets. Ignored when async is set.", + "exclusiveMinimum": 0, + "maximum": 600000, + "type": "integer" +} - added
Input schema / properties / styleAdded value: +{ + "description": "Name of a saved style preset (see gemini_list_styles / gemini_save_style): its prompt fragment — and reference image, if it has one — is applied to the request automatically. Hosted connector only.", + "minLength": 1, + "type": "string" +}
- Changed
gemini_interact2 fields changed- changed
Input schema / properties / async / descriptionPrevious value: -"Run in the background and return a job_id immediately instead of the image, so a long (Pro/4K) generation cannot hit the host tools/call timeout (-32001). Poll gemini_get_result with the job_id to fetch the result (jobs are per-process and expire ~10 min after completion)."New value: +"Run in the background and return a job_id immediately instead of the image, so a long (Pro/4K) generation cannot hit the host tools/call timeout (-32001). Poll gemini_get_result with the job_id to fetch the result. PREFER `max_wait_ms` on the hosted connector: it runs where the executor is only guaranteed to stay alive while the request is open, so this option is served there as a bounded wait rather than an immediate hand-off." - added
Input schema / properties / max_wait_msAdded value: +{ + "description": "Wait up to this many ms for the result; if generation is still running when the budget expires, return { job_id, status: \"running\" } immediately instead (poll gemini_get_result). Keeps fast results in-band while a slow batch can never trip the host tools/call timeout (-32001) — e.g. 20000 for multi-image sets. Ignored when async is set.", + "exclusiveMinimum": 0, + "maximum": 600000, + "type": "integer" +}
- Changed
gemini_music_generate2 fields changed- changed
Input schema / properties / async / descriptionPrevious value: -"Run in the background and return a job_id immediately instead of the image, so a long (Pro/4K) generation cannot hit the host tools/call timeout (-32001). Poll gemini_get_result with the job_id to fetch the result (jobs are per-process and expire ~10 min after completion)."New value: +"Run in the background and return a job_id immediately instead of the image, so a long (Pro/4K) generation cannot hit the host tools/call timeout (-32001). Poll gemini_get_result with the job_id to fetch the result. PREFER `max_wait_ms` on the hosted connector: it runs where the executor is only guaranteed to stay alive while the request is open, so this option is served there as a bounded wait rather than an immediate hand-off." - added
Input schema / properties / max_wait_msAdded value: +{ + "description": "Wait up to this many ms for the result; if generation is still running when the budget expires, return { job_id, status: \"running\" } immediately instead (poll gemini_get_result). Keeps fast results in-band while a slow batch can never trip the host tools/call timeout (-32001) — e.g. 20000 for multi-image sets. Ignored when async is set.", + "exclusiveMinimum": 0, + "maximum": 600000, + "type": "integer" +}
- Changed
gemini_video_generate2 fields changed- changed
Input schema / properties / async / descriptionPrevious value: -"Run in the background and return a job_id immediately instead of the image, so a long (Pro/4K) generation cannot hit the host tools/call timeout (-32001). Poll gemini_get_result with the job_id to fetch the result (jobs are per-process and expire ~10 min after completion)."New value: +"Run in the background and return a job_id immediately instead of the image, so a long (Pro/4K) generation cannot hit the host tools/call timeout (-32001). Poll gemini_get_result with the job_id to fetch the result. PREFER `max_wait_ms` on the hosted connector: it runs where the executor is only guaranteed to stay alive while the request is open, so this option is served there as a bounded wait rather than an immediate hand-off." - added
Input schema / properties / max_wait_msAdded value: +{ + "description": "Wait up to this many ms for the result; if generation is still running when the budget expires, return { job_id, status: \"running\" } immediately instead (poll gemini_get_result). Keeps fast results in-band while a slow batch can never trip the host tools/call timeout (-32001) — e.g. 20000 for multi-image sets. Ignored when async is set.", + "exclusiveMinimum": 0, + "maximum": 600000, + "type": "integer" +}
1 tool update
v1.4.0- Changed
gemini_upload_file1 field changed- added
Input schema / properties / r2_keyAdded value: +{ + "description": "Unavailable on this server (media is written to local disk) — pass the file path via `path` instead.", + "minLength": 1, + "type": "string" +}
11 tool updates
v1.2.0- First observed
gemini_delete_file - First observed
gemini_get_result - First observed
gemini_image_edit - First observed
gemini_image_generate - First observed
gemini_image_set - First observed
gemini_interact - First observed
gemini_list_files - First observed
gemini_list_models - First observed
gemini_music_generate - First observed
gemini_upload_file - First observed
gemini_video_generate
TDQS
Scored across 13 tools
The four image-related tools (gemini_image_generate, gemini_image_edit, gemini_image_set, gemini_interact) have overlapping purposes, though the detailed descriptions with cross-references help guide selection. The other tools are clearly distinct, so the set is not highly ambiguous, but agents may still struggle to choose the right image tool at first.
All tools share a consistent 'gemini_' prefix and snake_case, but the word order is mixed: some use verb_noun (list_models, upload_file), others noun_verb (image_generate, music_generate), and a few are compound nouns or bare verbs (healthcheck, interact). This mixed convention is readable but not uniform.
13 tools is well within the ideal 3-15 range for a media generation server. Each tool covers a distinct part of the workflow: generation, editing, async retrieval, file management, model listing, health check, and usage tracking. No tool feels redundant or missing.
The tool surface covers the full lifecycle for media generation: create (image/video/music), edit, refine, set generation, async results, file upload/list/delete, model listing, and cost tracking. There are no obvious dead ends or critical missing operations for the stated purpose.
Maintenance
Related MCP Connectors
MCP server for Google Veo AI video generation
Focused MCP server for OpenAI image/audio generation (v2.0.0). Wraps endpoints via HAPI CLI.
MCP server for Clipkit — gives AI agents a video toolbox via the Clipkit schema.
Generate AI images and videos from any compatible MCP client.
Related MCP Servers
- AlicenseAqualityFmaintenanceProduction-grade MCP server for image and video understanding and generation across Gemini, OpenAI, and Grok.5140 PyPI4Apache 2.0
- AlicenseNot gradedqualityDmaintenanceModel Context Protocol server exposing MiniMax's image, speech, music, and video generation APIs as MCP tools for use with any MCP-aware host.4,211 npmMIT
- FlicenseAqualityDmaintenanceWraps Google Gemini's image generation API as an MCP server, enabling text-to-image, image editing, and grounded search workflows from any MCP client.2-
- AlicenseNot gradedqualityDmaintenanceRemote MCP server that exposes Google Gemini's text, image, video (Veo), and audio transcription capabilities as tools any MCP client can call directly.95 npmMIT