Explain Image MCP Server
Explain Image MCP Server
A zero-dependency MCP (Model Context Protocol) server that lets any AI agent — including text-only models — "see" images.
The agent passes an image (local path, URL, or data URL) plus a prompt to the describe_image tool. The server forwards both to a Gemini vision model through the OpenAI-compatible REST endpoint and returns the model's text interpretation. Because the agent supplies the prompt, it controls exactly what the model returns: a description, OCR of visible text, an object list, structured JSON, and so on.
No external libraries are used — the MCP JSON-RPC protocol is implemented by hand over stdio, and the Gemini request is a plain fetch().
AI agent ── describe_image(image, prompt) ──► MCP server (stdio, JSON-RPC 2.0)
│
▼
POST /chat/completions (OpenAI-compatible)
│
▼
Gemini vision modelRequirements
Node.js >= 18.17 (global
fetchrequired)
Configuration
Env var | Default | Description |
| (required) | Google AI Studio API key |
|
| Default model id; override per call with the |
|
| OpenAI-compatible base URL (swap to add another provider later) |
Install with an MCP client
san:
san mcp add -e GEMINI_API_KEY=your-key explain-image -- node /path/to/explain-image-mcp-server/src/index.jsClaude Desktop (claude_desktop_config.json):
{
"mcpServers": {
"explain-image": {
"command": "node",
"args": ["/path/to/explain-image-mcp-server/src/index.js"],
"env": { "GEMINI_API_KEY": "your-key" }
}
}
}Tools
describe_image
Analyze one or more images with a Gemini vision model. The agent supplies the prompt.
Argument | Type | Required | Description |
| string | string[] | yes | Local file path, |
| string | no | What the model should return. Defaults to a detailed description |
| string | no | Gemini model id, overrides |
| integer | no | Maximum response length |
Example agent call:
describe_image(
image: "/screenshots/bug.png",
prompt: "Describe this UI bug precisely: what is shown, and what looks wrong?"
)list_models
List the model ids available on the configured OpenAI-compatible endpoint.
Default model & cost
The default is gemini-2.5-flash — a cost-efficient GA Flash model with image
input (Google pricing, July 2026): $0.30 / 1M input tokens, $2.50 / 1M output
tokens (vs $1.50 / $7.50 for gemini-3.6-flash). For 2.5 Flash/Flash-Lite
models the server sends reasoning_effort: "none", disabling thinking so no
output tokens are spent on reasoning. Set GEMINI_MODEL (or pass model per
call) to use a different model, e.g. gemini-3.6-flash for higher quality at a
higher price.
Development
npm test # end-to-end smoke test (protocol handshake + request formatting, no real key needed)The smoke test runs the full MCP handshake against a fake OpenAI-compatible upstream and validates the request shape (auth header, message body, base64 data URL) plus error paths.
Security note
The API key is stored in plaintext wherever the MCP client saves env vars. Keep it out of version control — .san/ is git-ignored in this repo for that reason.