Skip to main content
Glama
README.md
# mcp-vision-server

A lightweight [MCP](https://modelcontextprotocol.io) server exposing a single tool — `describe_image` — that forwards a local image file to **any** OpenAI-compatible vision endpoint and returns the text description.

Most LLMs can't see an image you drop into a chat. This server bridges that gap: the host AI calls `describe_image` with a file path, your vision model does the actual seeing, and the description flows back into the conversation.

```
┌──────────┐     MCP (stdio)      ┌─────────────────┐      HTTP POST     ┌──────────────────┐
│  Client  │ ─── describe_image ─▶│  mcp-vision-srv │ ── image+prompt ──▶│  Vision Endpoint │
│  (AI)    │ ◀── text description │   (this repo)   │ ◀──── JSON resp ───│  (vLLM/Ollama/…) │
└──────────┘                      └─────────────────┘                    └──────────────────┘
```

The server speaks MCP over **stdio** — it doesn't serve HTTP itself. It reads the file from disk, base64-encodes it, POSTs to your endpoint, and hands the response back as tool output.

## Requirements

- Python ≥ 3.10
- [`uv`](https://github.com/astral-sh/uv) installed
- **A running vision endpoint** that speaks the OpenAI chat-completions schema

That last one is not optional. This server has no model of its own — without something listening at `VISION_ENDPOINT`, every call fails. See [Supported Endpoints](#supported-endpoints).

## Install

No clone needed — `uvx` can run it straight from the repo:

```bash
uvx --from git+https://github.com/joshsssn/mcp-vision-server mcp-vision-server
```

Your MCP client will run this for you once configured; the command above is mainly useful to check the server starts.

**Claude Code**, one line:

```bash
claude mcp add vision \
  --env VISION_ENDPOINT=http://localhost:11434/v1/chat/completions \
  --env VISION_MODEL=llama3.2-vision \
  -- uvx --from git+https://github.com/joshsssn/mcp-vision-server mcp-vision-server
```

**VS Code (Copilot Chat)** — add to `.vscode/mcp.json` in your workspace, or to your user settings:

```jsonc
{
  "servers": {
    "mcp-vision-server": {
      "type": "stdio",
      "command": "uvx",
      "args": [
        "--from",
        "git+https://github.com/joshsssn/mcp-vision-server",
        "mcp-vision-server"
      ],
      "env": {
        "VISION_ENDPOINT": "https://api.openai.com/v1/chat/completions",
        "VISION_API_KEY": "sk-your-key-here",
        "VISION_MODEL": "gpt-4o"
      }
    }
  }
}
```

**Claude Desktop** — add to `claude_desktop_config.json`:

```jsonc
{
  "mcpServers": {
    "mcp-vision-server": {
      "command": "uvx",
      "args": [
        "--from",
        "git+https://github.com/joshsssn/mcp-vision-server",
        "mcp-vision-server"
      ],
      "env": {
        "VISION_ENDPOINT": "https://api.openai.com/v1/chat/completions",
        "VISION_API_KEY": "sk-your-key-here",
        "VISION_MODEL": "gpt-4o"
      }
    }
  }
}
```

> Working from a local clone instead? Replace the `git+https://…` argument with the path to the repo root.

### Local development

```bash
git clone https://github.com/joshsssn/mcp-vision-server.git
cd mcp-vision-server
cp .env.example .env    # edit with your endpoint URL, API key, and model name
uvx --from . mcp-vision-server
```

## Configuration

All settings come from environment variables, or a `.env` file in the project root:

| Variable | Default | Description |
|---|---|---|
| `VISION_ENDPOINT` | `http://localhost:11434/v1/chat/completions` | OpenAI-compatible vision API URL |
| `VISION_API_KEY` | `ollama` | Bearer token for the endpoint |
| `VISION_MODEL` | `llama3.2-vision` | Model name the endpoint expects |
| `VISION_TIMEOUT` | `60` | Request timeout in seconds |
| `VISION_MAX_IMAGE_BYTES` | `20971520` (20 MB) | Reject images larger than this |

The defaults target a local Ollama install, so `ollama run llama3.2-vision` plus a zero-config client entry is enough to get going.

The `env` block in your client config overrides `.env`. Use `.env` for local dev, the client block for anything you share — no keys end up in the repo either way.

See [`.env.example`](.env.example) for a ready-to-copy template.

## Tool: `describe_image`

| Parameter | Type | Required | Default |
|---|---|---|---|
| `image` | `string` | yes | — |
| `prompt` | `string` | no | `Describe this image in detail.` |

**`image`** must be an absolute path to a file on disk — not a URL, not a pasted image. Supported formats: PNG, JPG, JPEG, GIF, WEBP.

```
C:\Users\photos\cat.jpg
/home/user/images/screenshot.png
```

**`prompt`** is any natural-language instruction for the vision model:

- `"Describe this image in detail."`
- `"Extract all text from this image."`
- `"What colors dominate this image?"`
- `"Is there a person in this image? If so, describe them."`

Example call:

```json
{
  "tool": "describe_image",
  "arguments": {
    "image": "/home/user/photos/sunset.jpg",
    "prompt": "What colors dominate this image? Is there a person in it?"
  }
}
```

## Supported Endpoints

Anything implementing the OpenAI chat-completions schema with `image_url` content parts:

| Endpoint | Example URL | Notes |
|---|---|---|
| [Ollama](https://ollama.com) | `http://localhost:11434/v1/chat/completions` | Easiest local setup |
| [LM Studio](https://lmstudio.ai) | `http://localhost:1234/v1/chat/completions` | GUI-friendly |
| [vLLM](https://github.com/vllm-project/vllm) | `http://localhost:8000/v1/chat/completions` | Great for self-hosted |
| [OpenAI](https://platform.openai.com) | `https://api.openai.com/v1/chat/completions` | Cloud, requires API key |
| Custom / self-hosted | `https://your-box/v1/chat/completions` | Anything OpenAI-compatible |

## Troubleshooting

**`invalid peer certificate: UnknownIssuer` on startup.** `uv` uses its own certificate store and doesn't know about TLS-inspecting antivirus software or corporate proxies. Add `"UV_NATIVE_TLS": "1"` to the `env` block to make it use the system store instead.

**Connection refused / timeout on every call.** Nothing is listening at `VISION_ENDPOINT`. Confirm the endpoint independently before blaming the server:

```bash
curl $VISION_ENDPOINT -H "Content-Type: application/json" \
  -d '{"model":"llama3.2-vision","messages":[{"role":"user","content":"hi"}]}'
```

**Tool doesn't appear in the client.** Check the MCP logs — in VS Code, the Output panel has a channel per server. A failure to install the package shows up as a non-zero exit before `initialize` ever completes.

**Image rejected for size.** Raise `VISION_MAX_IMAGE_BYTES`, or downscale first. Base64 encoding inflates the payload by roughly a third, so the limit is deliberately conservative.

## Project Structure

```
mcp-vision-server/
├── pyproject.toml
├── .env.example
├── .gitignore
├── README.md
├── LICENCE
└── src/
    └── mcp_vision_server/
        ├── __init__.py
        └── server.py
```

## License

MIT — do whatever.

TDQS

A4.1/5.0

Scored across 1 tool

Disambiguation5/5

With only one tool, there is no possibility of confusion or overlap. The tool's purpose is uniquely defined.

Naming Consistency5/5

The single tool follows a clear verb_noun pattern (describe_image), which is consistent and descriptive.

Tool Count3/5

A single tool is on the thin side, but reasonable for a narrowly scoped vision-description server. It feels minimal but not trivial.

Completeness5/5

The tool fully covers the server's stated purpose of describing images from local files, with no obvious missing operations for that domain.

Maintenance

ActivitySlowing
ResponsivenessNo issues