Skip to main content
Glama
hsain9357

Explain Image MCP Server

by hsain9357
README.md
# Explain Image MCP Server

A **zero-dependency** [MCP](https://modelcontextprotocol.io) (Model Context Protocol) server that lets any AI agent — including text-only models — **"see" images**.

The agent passes an image (local path, URL, or data URL) plus a prompt to the `describe_image` tool. The server forwards both to a Gemini vision model through the **OpenAI-compatible** REST endpoint and returns the model's text interpretation. Because the agent supplies the prompt, it controls exactly what the model returns: a description, OCR of visible text, an object list, structured JSON, and so on.

No external libraries are used — the MCP JSON-RPC protocol is implemented by hand over stdio, and the Gemini request is a plain `fetch()`.

```
AI agent ── describe_image(image, prompt) ──► MCP server (stdio, JSON-RPC 2.0)
                                                  │
                                                  ▼
                              POST /chat/completions  (OpenAI-compatible)
                                                  │
                                                  ▼
                                            Gemini vision model
```

## Requirements

- Node.js >= 18.17 (global `fetch` required)
- A [Gemini API key](https://aistudio.google.com/apikey)

## Configuration

| Env var | Default | Description |
|---|---|---|
| `GEMINI_API_KEY` | *(required)* | Google AI Studio API key |
| `GEMINI_MODEL` | `gemini-2.5-flash` | Default model id; override per call with the `model` argument |
| `GEMINI_BASE_URL` | `https://generativelanguage.googleapis.com/v1beta/openai` | OpenAI-compatible base URL (swap to add another provider later) |

## Install with an MCP client

**san:**

```sh
san mcp add -e GEMINI_API_KEY=your-key explain-image -- node /path/to/explain-image-mcp-server/src/index.js
```

**Claude Desktop** (`claude_desktop_config.json`):

```json
{
  "mcpServers": {
    "explain-image": {
      "command": "node",
      "args": ["/path/to/explain-image-mcp-server/src/index.js"],
      "env": { "GEMINI_API_KEY": "your-key" }
    }
  }
}
```

## Tools

### `describe_image`

Analyze one or more images with a Gemini vision model. The agent supplies the prompt.

| Argument | Type | Required | Description |
|---|---|---|---|
| `image` | string \| string[] | yes | Local file path, `http(s)` URL, `data:` URL, or an array of these |
| `prompt` | string | no | What the model should return. Defaults to a detailed description |
| `model` | string | no | Gemini model id, overrides `GEMINI_MODEL` |
| `max_tokens` | integer | no | Maximum response length |

**Example agent call:**

```
describe_image(
  image: "/screenshots/bug.png",
  prompt: "Describe this UI bug precisely: what is shown, and what looks wrong?"
)
```

### `list_models`

List the model ids available on the configured OpenAI-compatible endpoint.

## Default model & cost

The default is `gemini-2.5-flash` — a cost-efficient GA Flash model with image
input (Google pricing, July 2026): **$0.30 / 1M input tokens, $2.50 / 1M output
tokens** (vs $1.50 / $7.50 for `gemini-3.6-flash`). For 2.5 Flash/Flash-Lite
models the server sends `reasoning_effort: "none"`, disabling thinking so no
output tokens are spent on reasoning. Set `GEMINI_MODEL` (or pass `model` per
call) to use a different model, e.g. `gemini-3.6-flash` for higher quality at a
higher price.

## Development

```sh
npm test   # end-to-end smoke test (protocol handshake + request formatting, no real key needed)
```

The smoke test runs the full MCP handshake against a fake OpenAI-compatible upstream and validates the request shape (auth header, message body, base64 data URL) plus error paths.

## Security note

The API key is stored in plaintext wherever the MCP client saves env vars. Keep it out of version control — `.san/` is git-ignored in this repo for that reason.

TDQS

A4.4/5.0

Scored across 2 tools

Disambiguation5/5

The two tools have completely distinct purposes: describe_image processes images, while list_models enumerates available models. There is no overlap or ambiguity in their roles.

Naming Consistency5/5

Both tools follow the identical verb_noun naming pattern (describe_image and list_models), creating a clear and predictable convention. This consistency makes the tool set easy to navigate.

Tool Count4/5

With only 2 tools, the server is lean but appropriately scoped for its narrow purpose of explaining images. The core describe_image tool is supported by list_models, giving just enough functionality without bloat.

Completeness5/5

For the stated purpose of image explanation, the server fully covers the domain. describe_image is flexible via custom prompts, supporting descriptions, OCR, object listing, and more, with no obvious gaps for a single-purpose MCP server.