Skip to main content
Glama
README.md
# gemini-image-mcp

MCP server that drives **gemini.google.com** in headless Chrome to generate images,
saves them to disk and returns the paths. Supports continuing the same conversation
(for iterative edits) or starting a new one.

It signs in with *your own* Google account through a normal browser session, so it needs
no API key -- and equally, it gives no API guarantees. Read the [Caveats](#caveats)
before relying on it.

## Requirements

- **Node.js 18+**
- **Google Chrome** installed (the real browser -- Chromium builds are not used)
- A **Google account with access to Gemini**
- A desktop session you can open a browser window in, for the one-time sign-in
  (a headless server with no display cannot complete it locally)

Linux and macOS should work -- Chrome is located automatically in the usual places, and
`GEMINI_MCP_CHROME` overrides the path -- but it has only been exercised on Windows.

## How the Google session is connected

There is no API key and no password handling. The server uses a **dedicated
persistent Chrome profile** at `~/.gemini-image-mcp/profile`:

1. `npm run login` opens your *real* Chrome (not bundled Chromium) with that profile, visible.
2. **You** sign in to Google by hand in that window — 2FA, passkey, whatever your account uses.
3. Close the window. Chrome has written the cookies into the profile directory.
4. The server then launches the same profile in Chrome's headless mode, so it is
   already signed in. Nothing in this code ever sees or stores a credential.

Chrome is spawned directly with `--remote-debugging-port` and attached via
`connectOverCDP`, rather than through Playwright's `launchPersistentContext`.
Playwright drives Chrome over `--remote-debugging-pipe`, which Chrome 154 refuses
for this profile (it exits immediately with code 21); the port transport works.

The session lasts as long as Google keeps it alive (typically weeks/months). When it
expires, every tool returns `NOT_SIGNED_IN` — re-run `npm run login`.

Using a separate profile (rather than your everyday Chrome profile) matters: Chrome
locks a profile directory while it is running, so automation against your main profile
would fail whenever Chrome is open.

## Setup

```bash
git clone https://github.com/sedvis/gemini-image-mcp.git
cd gemini-image-mcp
npm install
npm run login      # opens Chrome; sign in by hand, once
```

Check it worked, then try a generation straight from the CLI:

```bash
node src/smoke.mjs status
node src/smoke.mjs gen "a red fox in a misty forest at sunrise" new
```

### Registering with an MCP client

Any MCP client that speaks stdio works. The server is `node <abs-path>/src/server.mjs`:

```json
{
  "mcpServers": {
    "gemini-image": {
      "command": "node",
      "args": ["/abs/path/to/gemini-image-mcp/src/server.mjs"]
    }
  }
}
```

For **Claude Code** specifically: if you open **this repo** as your project, the bundled
`.mcp.json` is picked up automatically -- it resolves its own path:

```json
"args": ["${CLAUDE_PROJECT_DIR:-.}/src/server.mjs"]
```

`CLAUDE_PROJECT_DIR` is set in the spawned server's environment rather than Claude
Code's own, so the `:-.` default is what makes the expansion work from `.mcp.json`;
a bare relative path is not reliably resolved against the project root.

To use it from *other* projects, register it once globally with an absolute path:

```bash
claude mcp add gemini-image --scope user -- node /abs/path/to/gemini-image-mcp/src/server.mjs
```

## Tools

| Tool | Purpose |
|---|---|
| `gemini_generate_image` | `prompt`, `conversation` (`new` \| `continue`), `save_dir?`, `timeout_seconds?` → saved file paths + Gemini's reply text + chat URL |
| `gemini_new_conversation` | Forget the remembered thread, open a fresh chat |
| `gemini_session_status` | Is the profile still signed in? |

`conversation: "continue"` reuses the remembered chat URL, which is persisted to
`~/.gemini-image-mcp/state.json` — so follow-ups like *"same character, night scene"*
keep working across server restarts.

## Env vars

| Var | Default | Meaning |
|---|---|---|
| `GEMINI_MCP_HEADLESS` | `1` | Set to `0` to watch the browser — useful when selectors break |
| `GEMINI_MCP_HOME` | `~/.gemini-image-mcp` | Profile + state location |
| `GEMINI_MCP_OUT` | `~/.gemini-image-mcp/images` | Default save directory |
| `GEMINI_MCP_CHROME` | auto-detected | Full path to the Chrome executable, if it is not in a standard location |

## Troubleshooting

| Symptom | Cause / fix |
|---|---|
| `NOT_SIGNED_IN` | Session expired, or an interstitial is blocking. Re-run `npm run login`. |
| `Chrome exited with code 21` | Profile locked by another Chrome using the same `--user-data-dir`, or a stale `SingletonLock` in the profile. Close it / delete the lock. |
| Image returned but `files: []` | Both capture paths failed. See *How images are saved* below. |
| File is a `-preview-` PNG | The download button was missed, so the downscaled preview was read off the DOM instead. Re-run; check the `SEL.download` selector. |
| Wrong/duplicate image returned | Only the newest `model-response` bubble is harvested — earlier turns lazy-load and swap their `src`, so global image counting is not reliable. |

## How images are saved

Images are saved via Gemini's own **"Download full-sized image"** button, which yields
the real asset -- a ~2816x1536 JPEG at 300 DPI.

This matters: the `<img>` rendered in the page is only a downscaled ~1024px preview, so
reading pixels off the DOM element silently loses most of the resolution. That read is
kept purely as a fallback, and files it produces are marked `-preview-` in the filename.
The fallback also cannot use `fetch`: image URLs are CORS-blocked from the page context,
and a fresh turn serves the image as a `blob:` URL, so it goes through a canvas read.

## Caveats

- This drives the consumer web UI, so it depends on Gemini's DOM. If Google reshuffles
  the markup, generation returns a `warning` plus a debug screenshot instead of files;
  the selectors live in one `SEL` object at the top of `src/gemini.mjs`.
- Automating a consumer web account is outside Gemini's intended use; for anything
  production-facing, the Gemini API with an API key is the supported route.
- Google may occasionally show an interstitial (consent, "verify it's you") that
  headless cannot clear. Run with `GEMINI_MCP_HEADLESS=0` once to clear it by hand.
- The profile directory holds a live Google session. Treat it like a credential: it is
  kept outside the repo (in `~/.gemini-image-mcp/`) so it is never committed, and it is
  per-machine -- do not copy it between machines or users.
- Rate limits, model choice and output size are whatever your Google account's Gemini
  gives you; this server does not control them.

## License

MIT -- see [LICENSE](LICENSE).

TDQS

A3.9/5.0

Scored across 3 tools

Disambiguation4/5

Each tool targets a distinct action: resetting the chat thread, checking session status, and generating images. There is mild overlap since gemini_new_conversation and gemini_generate_image's conversation="new" both start a fresh chat, but descriptions make the distinction clear.

Naming Consistency5/5

All three tools follow a consistent gemini_<verb>_<noun> snake_case pattern (new_conversation, session_status, generate_image). No deviations in style or casing.

Tool Count4/5

Three tools is slightly lean for the surface, but the server's scope is narrow (image generation via a browser session) and each tool earns its place. No obvious padding or missing operations-driven bloat.

Completeness4/5

The core lifecycle—verify session, start/continue a conversation, generate and save images—is covered. Login is handled externally via npm run login, so the only minor gap is no tooling around listing or managing previously saved images.

Maintenance

ActivityMaintained
ResponsivenessNo issues