gemini-image-mcp
Drives the gemini.google.com consumer web UI to generate images from text prompts, saving the full-resolution results to disk and returning the file paths along with Gemini's reply text and chat URL. Supports continuing the remembered conversation for iterative edits (e.g. "same character, night scene") or starting a fresh chat, plus a status tool to check whether the signed-in Google session is still valid.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@gemini-image-mcpgenerate an image of a cozy cabin in a snowy forest at dusk"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
gemini-image-mcp
MCP server that drives gemini.google.com in headless Chrome to generate images, saves them to disk and returns the paths. Supports continuing the same conversation (for iterative edits) or starting a new one.
It signs in with your own Google account through a normal browser session, so it needs no API key -- and equally, it gives no API guarantees. Read the Caveats before relying on it.
Requirements
Node.js 18+
Google Chrome installed (the real browser -- Chromium builds are not used)
A Google account with access to Gemini
A desktop session you can open a browser window in, for the one-time sign-in (a headless server with no display cannot complete it locally)
Linux and macOS should work -- Chrome is located automatically in the usual places, and
GEMINI_MCP_CHROME overrides the path -- but it has only been exercised on Windows.
Related MCP server: nano-banana-mcp
How the Google session is connected
There is no API key and no password handling. The server uses a dedicated
persistent Chrome profile at ~/.gemini-image-mcp/profile:
npm run loginopens your real Chrome (not bundled Chromium) with that profile, visible.You sign in to Google by hand in that window — 2FA, passkey, whatever your account uses.
Close the window. Chrome has written the cookies into the profile directory.
The server then launches the same profile in Chrome's headless mode, so it is already signed in. Nothing in this code ever sees or stores a credential.
Chrome is spawned directly with --remote-debugging-port and attached via
connectOverCDP, rather than through Playwright's launchPersistentContext.
Playwright drives Chrome over --remote-debugging-pipe, which Chrome 154 refuses
for this profile (it exits immediately with code 21); the port transport works.
The session lasts as long as Google keeps it alive (typically weeks/months). When it
expires, every tool returns NOT_SIGNED_IN — re-run npm run login.
Using a separate profile (rather than your everyday Chrome profile) matters: Chrome locks a profile directory while it is running, so automation against your main profile would fail whenever Chrome is open.
Setup
git clone https://github.com/sedvis/gemini-image-mcp.git
cd gemini-image-mcp
npm install
npm run login # opens Chrome; sign in by hand, onceCheck it worked, then try a generation straight from the CLI:
node src/smoke.mjs status
node src/smoke.mjs gen "a red fox in a misty forest at sunrise" newRegistering with an MCP client
Any MCP client that speaks stdio works. The server is node <abs-path>/src/server.mjs:
{
"mcpServers": {
"gemini-image": {
"command": "node",
"args": ["/abs/path/to/gemini-image-mcp/src/server.mjs"]
}
}
}For Claude Code specifically: if you open this repo as your project, the bundled
.mcp.json is picked up automatically -- it resolves its own path:
"args": ["${CLAUDE_PROJECT_DIR:-.}/src/server.mjs"]CLAUDE_PROJECT_DIR is set in the spawned server's environment rather than Claude
Code's own, so the :-. default is what makes the expansion work from .mcp.json;
a bare relative path is not reliably resolved against the project root.
To use it from other projects, register it once globally with an absolute path:
claude mcp add gemini-image --scope user -- node /abs/path/to/gemini-image-mcp/src/server.mjsTools
Tool | Purpose |
|
|
| Forget the remembered thread, open a fresh chat |
| Is the profile still signed in? |
conversation: "continue" reuses the remembered chat URL, which is persisted to
~/.gemini-image-mcp/state.json — so follow-ups like "same character, night scene"
keep working across server restarts.
Env vars
Var | Default | Meaning |
|
| Set to |
|
| Profile + state location |
|
| Default save directory |
| auto-detected | Full path to the Chrome executable, if it is not in a standard location |
Troubleshooting
Symptom | Cause / fix |
| Session expired, or an interstitial is blocking. Re-run |
| Profile locked by another Chrome using the same |
Image returned but | Both capture paths failed. See How images are saved below. |
File is a | The download button was missed, so the downscaled preview was read off the DOM instead. Re-run; check the |
Wrong/duplicate image returned | Only the newest |
How images are saved
Images are saved via Gemini's own "Download full-sized image" button, which yields the real asset -- a ~2816x1536 JPEG at 300 DPI.
This matters: the <img> rendered in the page is only a downscaled ~1024px preview, so
reading pixels off the DOM element silently loses most of the resolution. That read is
kept purely as a fallback, and files it produces are marked -preview- in the filename.
The fallback also cannot use fetch: image URLs are CORS-blocked from the page context,
and a fresh turn serves the image as a blob: URL, so it goes through a canvas read.
Caveats
This drives the consumer web UI, so it depends on Gemini's DOM. If Google reshuffles the markup, generation returns a
warningplus a debug screenshot instead of files; the selectors live in oneSELobject at the top ofsrc/gemini.mjs.Automating a consumer web account is outside Gemini's intended use; for anything production-facing, the Gemini API with an API key is the supported route.
Google may occasionally show an interstitial (consent, "verify it's you") that headless cannot clear. Run with
GEMINI_MCP_HEADLESS=0once to clear it by hand.The profile directory holds a live Google session. Treat it like a credential: it is kept outside the repo (in
~/.gemini-image-mcp/) so it is never committed, and it is per-machine -- do not copy it between machines or users.Rate limits, model choice and output size are whatever your Google account's Gemini gives you; this server does not control them.
License
MIT -- see LICENSE.
Available Tools
3 toolsgemini_generate_imageGenerate an image with GeminiA
Sends a prompt to gemini.google.com in a headless Chrome signed in with the local Google profile, waits for the generated image(s), saves them to disk and returns the file paths. conversation="continue" keeps the existing chat thread (so you can iterate: "make it darker"); conversation="new" starts a fresh chat.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | The image prompt, or a follow-up edit when continuing. | |
| save_dir | No | Directory for the saved images. Default: /root/.gemini-image-mcp/images | |
| conversation | No | "new" starts a fresh Gemini chat; "continue" stays in the current thread. | continue |
| timeout_seconds | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the execution environment (headless Chrome signed in with the local Google profile), the auth/profile requirement, the blocking wait for generated images, disk persistence, and the return shape (file paths). It does not mention failure modes, timeouts, or rate limiting, which would be needed for a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no padding. The first front-loads the mechanism and output behavior; the second explains the conversation toggle with a concrete iteration example. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter tool with no output schema, the description covers what the tool does, the auth/environment dependency, and the return value (file paths). It is nearly complete, with only timeout behavior and error handling left unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75%, so the baseline is 3, but the description adds real meaning: it clarifies the conversation enum with concrete iteration semantics and the "make it darker" example, and confirms output goes to disk as file paths. The prompt's follow-up-edit behavior is also reinforced. timeout_seconds and save_dir get no added context, keeping it from a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource (generate images via gemini.google.com) and explains the mechanism: headless Chrome with the local Google profile, saving images to disk and returning file paths. It does not name the sibling tools (gemini_new_conversation, gemini_session_status) or explicitly contrast with them, so it falls short of full sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied through the conversation parameter semantics: "continue" for iteration ("make it darker") and "new" for a fresh chat. However, there is no explicit guidance on when to choose this tool versus gemini_new_conversation, nor any prerequisites or exclusions stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gemini_new_conversationStart a new Gemini conversationA
Drops the remembered chat thread and opens a fresh one.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It usefully discloses that the tool 'drops the remembered chat thread,' indicating a context reset, which is a key behavioral trait. However, it does not cover other potentially relevant aspects like reversibility, effect on other sessions, or any prerequisites, leaving room for more transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words. It efficiently conveys the core behavior in a compact form.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, parameterless tool with no output schema, the description provides the essential behavior (resetting the conversation). While it lacks usage guidelines and deeper behavioral details, it is sufficient for an agent to understand the primary effect. The absence of annotations means a bit more context could help, but overall it is nearly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline score of 4 applies. The description does not need to explain parameter semantics, and the empty schema is self-explanatory.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action (drops the remembered chat thread and opens a fresh one) and clearly identifies the resource (conversation). It distinguishes itself from siblings like gemini_session_status and gemini_generate_image by focusing on resetting the conversation context. The only minor gap is that it doesn't explicitly name Gemini, but the title and name do.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers no guidance on when to use this tool versus alternatives, no prerequisites, and no exclusions. It simply states what the tool does, leaving the agent to infer appropriate usage contexts.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gemini_session_statusCheck the Gemini browser sessionA
Reports whether the local Chrome profile still has a valid Google/Gemini session. If signedIn is false, the user must run "npm run login" in the gemini-image-mcp project and sign in by hand (profile: /root/.gemini-image-mcp/profile).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it does well: it names the concrete state field (signedIn), the profile path, and the manual recovery step. It stops short of describing the return shape beyond that one field or what happens on repeated checks, but for a zero-arg status probe this is substantial behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core purpose, then the actionable fallback. The remediation detail is longer than strictly necessary but earns its place by being directly actionable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description names the field the agent should inspect (signedIn) and the failure remediation (npm run login with the profile path), which is everything an agent needs to interpret and act on the result for this simple probe.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes no parameters, so the baseline is 4. The description correctly implies a pure no-argument status check and points the agent to the relevant field/state rather than any input.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: it reports whether the local Chrome profile still holds a valid Google/Gemini session, anchored to the concrete field signedIn. Sibling tools (gemini_new_conversation, gemini_generate_image) are generation actions, so this diagnostic/status tool is clearly distinguishable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explains the condition that selects this tool's output (signedIn false) and gives the exact remediation (
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
v1.0.0- First observed
gemini_generate_image - First observed
gemini_new_conversation - First observed
gemini_session_status
TDQS
Scored across 3 tools
Each tool targets a distinct action: resetting the chat thread, checking session status, and generating images. There is mild overlap since gemini_new_conversation and gemini_generate_image's conversation="new" both start a fresh chat, but descriptions make the distinction clear.
All three tools follow a consistent gemini_<verb>_<noun> snake_case pattern (new_conversation, session_status, generate_image). No deviations in style or casing.
Three tools is slightly lean for the surface, but the server's scope is narrow (image generation via a browser session) and each tool earns its place. No obvious padding or missing operations-driven bloat.
The core lifecycle—verify session, start/continue a conversation, generate and save images—is covered. Login is handled externally via npm run login, so the only minor gap is no tooling around listing or managing previously saved images.
Maintenance
Related MCP Connectors
LLM chat, text tools, image generation, editing, batch image jobs, and asynchronous video generation
Generate images, videos, voiceovers, and captions from a chat prompt.
Create images and videos from prompts, with options for image mixing, reference images, and start/…
Generate and edit images, create videos, quote credit costs, and retrieve private results.
Related MCP Servers
- FlicenseAqualityDmaintenanceGenerates images from text prompts using ElevenLabs' image generation API via browser automation, with support for multiple models and persistent authentication.3-
- FlicenseNot gradedqualityDmaintenanceGenerates images using Google's Gemini model, with options to save to file or return as base64 data URL.-
- AlicenseAqualityDmaintenanceEnables image generation and transformation using Gemini API, returning file paths instead of base64 data to avoid token limit issues in Claude Code.21MIT
- FlicenseNot gradedqualityDmaintenanceEnables image generation and editing via third-party relay services. Returns local file paths and Markdown display hints.-