Explain Image MCP Server
# Explain Image MCP Server
A **zero-dependency** [MCP](https://modelcontextprotocol.io) (Model Context Protocol) server that lets any AI agent — including text-only models — **"see" images**.
The agent passes an image (local path, URL, or data URL) plus a prompt to the `describe_image` tool. The server forwards both to a Gemini vision model through the **OpenAI-compatible** REST endpoint and returns the model's text interpretation. Because the agent supplies the prompt, it controls exactly what the model returns: a description, OCR of visible text, an object list, structured JSON, and so on.
No external libraries are used — the MCP JSON-RPC protocol is implemented by hand over stdio, and the Gemini request is a plain `fetch()`.
```
AI agent ── describe_image(image, prompt) ──► MCP server (stdio, JSON-RPC 2.0)
│
▼
POST /chat/completions (OpenAI-compatible)
│
▼
Gemini vision model
```
## Requirements
- Node.js >= 18.17 (global `fetch` required)
- A [Gemini API key](https://aistudio.google.com/apikey)
## Configuration
| Env var | Default | Description |
|---|---|---|
| `GEMINI_API_KEY` | *(required)* | Google AI Studio API key |
| `GEMINI_MODEL` | `gemini-2.5-flash` | Default model id; override per call with the `model` argument |
| `GEMINI_BASE_URL` | `https://generativelanguage.googleapis.com/v1beta/openai` | OpenAI-compatible base URL (swap to add another provider later) |
## Install with an MCP client
**san:**
```sh
san mcp add -e GEMINI_API_KEY=your-key explain-image -- node /path/to/explain-image-mcp-server/src/index.js
```
**Claude Desktop** (`claude_desktop_config.json`):
```json
{
"mcpServers": {
"explain-image": {
"command": "node",
"args": ["/path/to/explain-image-mcp-server/src/index.js"],
"env": { "GEMINI_API_KEY": "your-key" }
}
}
}
```
## Tools
### `describe_image`
Analyze one or more images with a Gemini vision model. The agent supplies the prompt.
| Argument | Type | Required | Description |
|---|---|---|---|
| `image` | string \| string[] | yes | Local file path, `http(s)` URL, `data:` URL, or an array of these |
| `prompt` | string | no | What the model should return. Defaults to a detailed description |
| `model` | string | no | Gemini model id, overrides `GEMINI_MODEL` |
| `max_tokens` | integer | no | Maximum response length |
**Example agent call:**
```
describe_image(
image: "/screenshots/bug.png",
prompt: "Describe this UI bug precisely: what is shown, and what looks wrong?"
)
```
### `list_models`
List the model ids available on the configured OpenAI-compatible endpoint.
## Default model & cost
The default is `gemini-2.5-flash` — a cost-efficient GA Flash model with image
input (Google pricing, July 2026): **$0.30 / 1M input tokens, $2.50 / 1M output
tokens** (vs $1.50 / $7.50 for `gemini-3.6-flash`). For 2.5 Flash/Flash-Lite
models the server sends `reasoning_effort: "none"`, disabling thinking so no
output tokens are spent on reasoning. Set `GEMINI_MODEL` (or pass `model` per
call) to use a different model, e.g. `gemini-3.6-flash` for higher quality at a
higher price.
## Development
```sh
npm test # end-to-end smoke test (protocol handshake + request formatting, no real key needed)
```
The smoke test runs the full MCP handshake against a fake OpenAI-compatible upstream and validates the request shape (auth header, message body, base64 data URL) plus error paths.
## Security note
The API key is stored in plaintext wherever the MCP client saves env vars. Keep it out of version control — `.san/` is git-ignored in this repo for that reason.
TDQS
Scored across 2 tools
The two tools have completely distinct purposes: describe_image processes images, while list_models enumerates available models. There is no overlap or ambiguity in their roles.
Both tools follow the identical verb_noun naming pattern (describe_image and list_models), creating a clear and predictable convention. This consistency makes the tool set easy to navigate.
With only 2 tools, the server is lean but appropriately scoped for its narrow purpose of explaining images. The core describe_image tool is supported by list_models, giving just enough functionality without bloat.
For the stated purpose of image explanation, the server fully covers the domain. describe_image is flexible via custom prompts, supporting descriptions, OCR, object listing, and more, with no obvious gaps for a single-purpose MCP server.