Skip to main content
Glama
knownothing20

vision-mcp

README.md
# vision-mcp

Vision MCP Server for AI Agents - give your text-only model eyes.

Use any OpenAI-compatible vision model API through MCP tools for image
analysis, OCR, error diagnosis, diagram understanding, and chart analysis.

[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
[![Version](https://img.shields.io/badge/version-1.1.0-blue)](https://github.com/knownothing20/vision-mcp)

[中文文档](./README_CN.md)

---

## One-line Summary

Give AI Agents visual understanding through MCP tools. Works with any
OpenAI-compatible vision API (GLM-4V, Qwen-VL, MiMo-V2.5, GPT-4o, etc.).

**Analogy**: Vision MCP is the eyes for text-only models like GLM-5.1.

---

## Architecture

```text
User sends image path or URL
  -> AI Agent (text model, e.g. GLM-5.1)
  -> vision-mcp tools (MCP)
  -> Vision Model API (e.g. MiMo-V2.5-Free, Qwen3-VL, GLM-4V)
  -> Returns text description
  -> AI Agent continues with understanding
```

---

## Features

| Tool | Description |
|------|-------------|
| `image_analysis` | General image understanding |
| `image_analysis_url` | Analyze image from URL |
| `extract_text` | OCR - extract text from screenshots |
| `diagnose_error` | Analyze error screenshots, suggest fixes |
| `understand_diagram` | Architecture/flow/UML diagram analysis |
| `analyze_chart` | Chart and data visualization analysis |

---

## Supported Vision APIs

| Provider | Model | API Base | Cost |
|----------|-------|----------|------|
| OpenCode Zen (MiMo-V2.5) | `mimo-v2.5-free` | `https://opencode.ai/zen/v1` | Free |
| Zhipu (GLM-4V) | `glm-4v` | `https://open.bigmodel.cn/api/paas/v4` | Paid |
| Qwen-VL (via proxy) | `Qwen3-VL-ms` | Your proxy URL | Varies |
| OpenAI | `gpt-4o` | `https://api.openai.com/v1` | Paid |
| Any OpenAI-compatible | Custom | Custom | Varies |

> MiMo-V2.5 is a native omnimodal model by Xiaomi that supports text, image,
> video, and audio. The free tier on OpenCode Zen is a great starting point.

---

## Quick Start

### 1. Clone

```bash
git clone https://github.com/knownothing20/vision-mcp.git
cd vision-mcp
npm install
```

### 2. Configure

```bash
cp local/.env.example local/.env
# Edit local/.env, fill in your API key and model
```

Minimal config for MiMo-V2.5 Free (OpenCode Zen):

```env
VISION_API_KEY=your-opencode-zen-api-key
VISION_API_BASE=https://opencode.ai/zen/v1
VISION_MODEL=mimo-v2.5-free
```

### 3. Register in opencode

Add to `opencode.json` (replace path with your actual skill dir):

```json
{
  "mcp": {
    "vision-mcp": {
      "command": ["node", "/path/to/vision-mcp/index.js"],
      "enabled": true,
      "type": "local"
    }
  }
}
```

Or use `node sync.cjs` to auto-register (recommended).

### 4. Restart opencode

Restart your AI tool to activate MCP registration.

---

## Configuration

Edit `local/.env`:

| Variable | Description | Default |
|----------|-------------|---------|
| `VISION_API_KEY` | API key (required) | - |
| `VISION_API_BASE` | API endpoint | `https://open.bigmodel.cn/api/paas/v4` |
| `VISION_MODEL` | Model name | `glm-4v` |

Priority: `local/.env` > `.env` (root) > environment variables.

### Provider Examples

<details>
<summary>OpenCode Zen (MiMo-V2.5 Free) - Recommended free option</summary>

```env
VISION_API_KEY=sk-your-zen-api-key
VISION_API_BASE=https://opencode.ai/zen/v1
VISION_MODEL=mimo-v2.5-free
```

Get your API key at https://opencode.ai/auth

</details>

<details>
<summary>Zhipu GLM-4V</summary>

```env
VISION_API_KEY=your-zhipu-api-key
VISION_API_BASE=https://open.bigmodel.cn/api/paas/v4
VISION_MODEL=glm-4v
```

</details>

<details>
<summary>Qwen-VL via local proxy</summary>

```env
VISION_API_KEY=your-proxy-key
VISION_API_BASE=http://192.168.x.x:8317/v1
VISION_MODEL=Qwen3-VL-ms
```

</details>

<details>
<summary>OpenAI GPT-4o</summary>

```env
VISION_API_KEY=your-openai-api-key
VISION_API_BASE=https://api.openai.com/v1
VISION_MODEL=gpt-4o
```

</details>

---

## Usage

After setup, provide image file paths to your AI agent:

```
"Analyze this image: C:\Users\you\Desktop\screenshot.png"
"Extract text from /path/to/document.png"
"Diagnose this error: C:\screenshots\error.png"
"What does this architecture diagram mean? /path/to/diagram.png"
"Analyze chart trends: /path/to/chart.png"
```

> **Note**: Your main AI model (e.g. GLM-5.1) may not support image input
> directly. Provide the **file path** instead of embedding the image in chat,
> and the agent will use vision-mcp tools to analyze it.

---

## Directory Structure

```
vision-mcp/
├── index.js              # MCP server main program
├── sync.cjs              # Repo <-> skill dir sync tool
├── package.json          # Dependencies
├── SKILL.md              # Agent documentation
├── README.md             # This file
├── README_CN.md          # Chinese documentation
├── AGENT_GUIDE.md        # Agent install guide
├── INSTALL_OTHER_MACHINE.md  # Install on another machine
├── CHANGELOG.md          # Version changelog
├── LICENSE               # MIT License
├── .gitignore
└── local/                # Local private config (gitignored)
    └── .env.example      # Config template
```

---

## License

MIT License - See [LICENSE](./LICENSE)