Skip to main content
Glama
README.md
# gemini-image-mcp

A tiny [MCP](https://modelcontextprotocol.io) server that analyzes images with Google's
Gemini vision models. It exposes one tool, `analyze_image`, that takes a local image path
(or URL) plus an optional prompt and returns Gemini's text answer.

**Why:** it lets an agent (e.g. Claude Code) read screenshots, diagrams, charts, or UI
states *by reference* — the raw image bytes go to Gemini, and only the text answer comes
back, so they never bloat the calling agent's context window.

It talks straight to the Gemini REST API (`generativelanguage.googleapis.com`) with `fetch` —
no Google SDK, no `gemini-cli`, nothing tied to the deprecated consumer CLI.

## Setup

```sh
npm install
npm run build
cp .env.example .env   # then put your key in .env
```

Get an API key at <https://aistudio.google.com/apikey>.

### Configuration

Set via `.env` (loaded automatically from the repo root) or ambient environment:

| Variable         | Required | Default               | Notes                                              |
| ---------------- | -------- | --------------------- | -------------------------------------------------- |
| `GEMINI_API_KEY` | yes      | —                     | Your AI Studio key. Never commit it.               |
| `GEMINI_MODEL`   | no       | `gemini-flash-latest` | Use `gemini-pro-latest` for harder visual reasoning. |

The key is sent as an `x-goog-api-key` header (kept out of URLs/logs) and is never written
to a tracked file — `.env` is gitignored.

## Use with Claude Code

```sh
claude mcp add gemini-image -- node /absolute/path/to/gemini-image-mcp/dist/index.js
```

The server loads its own `.env`, so no key needs to live in Claude's config. Restart Claude
Code, then it can call `analyze_image` with an image path and an optional prompt.

## Tool: `analyze_image`

| Argument | Type              | Required | Description                                                                                   |
| -------- | ----------------- | -------- | --------------------------------------------------------------------------------------------- |
| `image`  | string \| string[] | yes      | A single local file path or `http(s)` URL, or an array of them to compare/reason about together. |
| `prompt` | string            | no       | What to ask. Defaults to a detailed description (one image) or a comparison (several).         |
| `model`  | string            | no       | Per-call model override.                                                                       |

Pass several images to compare them (before/after, spot-the-difference, "do these match"). Each is
labelled `Image 1`, `Image 2`, … in order, so the prompt can refer to them. All images ride in a single
Gemini request.

Supported inputs: PNG, JPEG, WebP, GIF, BMP, HEIC/HEIF, and PDF.

## Smoke test

Verify the key + API + image path end to end, without the MCP layer. The prompt comes
first (pass `""` for the default), then one or more images:

```sh
npm run smoke -- "What does this image say?" ./test/sample.png
npm run smoke -- "What changed between these?" ./before.png ./after.png
```

## License

MIT

TDQS

A3.9/5.0

Scored across 1 tool

Disambiguation5/5

Only one tool exists, so there is no possibility of confusion or overlap with other tools.

Naming Consistency5/5

The single tool name 'analyze_image' follows a clear verb_noun pattern and is self-consistent.

Tool Count2/5

A single tool is too few for the apparent scope of image analysis; users would likely expect additional tools for model selection or listing capabilities.

Completeness2/5

The tool surface is severely limited, lacking any supporting tools such as model listing or status checks, leaving obvious gaps in functionality.

Maintenance

ActivityInactive
ResponsivenessNo issues