vision-mcp
README.md
# vision-mcp
Vision MCP Server for AI Agents - give your text-only model eyes.
Use any OpenAI-compatible vision model API through MCP tools for image
analysis, OCR, error diagnosis, diagram understanding, and chart analysis.
[](https://opensource.org/licenses/MIT)
[](https://github.com/knownothing20/vision-mcp)
[中文文档](./README_CN.md)
---
## One-line Summary
Give AI Agents visual understanding through MCP tools. Works with any
OpenAI-compatible vision API (GLM-4V, Qwen-VL, MiMo-V2.5, GPT-4o, etc.).
**Analogy**: Vision MCP is the eyes for text-only models like GLM-5.1.
---
## Architecture
```text
User sends image path or URL
-> AI Agent (text model, e.g. GLM-5.1)
-> vision-mcp tools (MCP)
-> Vision Model API (e.g. MiMo-V2.5-Free, Qwen3-VL, GLM-4V)
-> Returns text description
-> AI Agent continues with understanding
```
---
## Features
| Tool | Description |
|------|-------------|
| `image_analysis` | General image understanding |
| `image_analysis_url` | Analyze image from URL |
| `extract_text` | OCR - extract text from screenshots |
| `diagnose_error` | Analyze error screenshots, suggest fixes |
| `understand_diagram` | Architecture/flow/UML diagram analysis |
| `analyze_chart` | Chart and data visualization analysis |
---
## Supported Vision APIs
| Provider | Model | API Base | Cost |
|----------|-------|----------|------|
| OpenCode Zen (MiMo-V2.5) | `mimo-v2.5-free` | `https://opencode.ai/zen/v1` | Free |
| Zhipu (GLM-4V) | `glm-4v` | `https://open.bigmodel.cn/api/paas/v4` | Paid |
| Qwen-VL (via proxy) | `Qwen3-VL-ms` | Your proxy URL | Varies |
| OpenAI | `gpt-4o` | `https://api.openai.com/v1` | Paid |
| Any OpenAI-compatible | Custom | Custom | Varies |
> MiMo-V2.5 is a native omnimodal model by Xiaomi that supports text, image,
> video, and audio. The free tier on OpenCode Zen is a great starting point.
---
## Quick Start
### 1. Clone
```bash
git clone https://github.com/knownothing20/vision-mcp.git
cd vision-mcp
npm install
```
### 2. Configure
```bash
cp local/.env.example local/.env
# Edit local/.env, fill in your API key and model
```
Minimal config for MiMo-V2.5 Free (OpenCode Zen):
```env
VISION_API_KEY=your-opencode-zen-api-key
VISION_API_BASE=https://opencode.ai/zen/v1
VISION_MODEL=mimo-v2.5-free
```
### 3. Register in opencode
Add to `opencode.json` (replace path with your actual skill dir):
```json
{
"mcp": {
"vision-mcp": {
"command": ["node", "/path/to/vision-mcp/index.js"],
"enabled": true,
"type": "local"
}
}
}
```
Or use `node sync.cjs` to auto-register (recommended).
### 4. Restart opencode
Restart your AI tool to activate MCP registration.
---
## Configuration
Edit `local/.env`:
| Variable | Description | Default |
|----------|-------------|---------|
| `VISION_API_KEY` | API key (required) | - |
| `VISION_API_BASE` | API endpoint | `https://open.bigmodel.cn/api/paas/v4` |
| `VISION_MODEL` | Model name | `glm-4v` |
Priority: `local/.env` > `.env` (root) > environment variables.
### Provider Examples
<details>
<summary>OpenCode Zen (MiMo-V2.5 Free) - Recommended free option</summary>
```env
VISION_API_KEY=sk-your-zen-api-key
VISION_API_BASE=https://opencode.ai/zen/v1
VISION_MODEL=mimo-v2.5-free
```
Get your API key at https://opencode.ai/auth
</details>
<details>
<summary>Zhipu GLM-4V</summary>
```env
VISION_API_KEY=your-zhipu-api-key
VISION_API_BASE=https://open.bigmodel.cn/api/paas/v4
VISION_MODEL=glm-4v
```
</details>
<details>
<summary>Qwen-VL via local proxy</summary>
```env
VISION_API_KEY=your-proxy-key
VISION_API_BASE=http://192.168.x.x:8317/v1
VISION_MODEL=Qwen3-VL-ms
```
</details>
<details>
<summary>OpenAI GPT-4o</summary>
```env
VISION_API_KEY=your-openai-api-key
VISION_API_BASE=https://api.openai.com/v1
VISION_MODEL=gpt-4o
```
</details>
---
## Usage
After setup, provide image file paths to your AI agent:
```
"Analyze this image: C:\Users\you\Desktop\screenshot.png"
"Extract text from /path/to/document.png"
"Diagnose this error: C:\screenshots\error.png"
"What does this architecture diagram mean? /path/to/diagram.png"
"Analyze chart trends: /path/to/chart.png"
```
> **Note**: Your main AI model (e.g. GLM-5.1) may not support image input
> directly. Provide the **file path** instead of embedding the image in chat,
> and the agent will use vision-mcp tools to analyze it.
---
## Directory Structure
```
vision-mcp/
├── index.js # MCP server main program
├── sync.cjs # Repo <-> skill dir sync tool
├── package.json # Dependencies
├── SKILL.md # Agent documentation
├── README.md # This file
├── README_CN.md # Chinese documentation
├── AGENT_GUIDE.md # Agent install guide
├── INSTALL_OTHER_MACHINE.md # Install on another machine
├── CHANGELOG.md # Version changelog
├── LICENSE # MIT License
├── .gitignore
└── local/ # Local private config (gitignored)
└── .env.example # Config template
```
---
## License
MIT License - See [LICENSE](./LICENSE)
This server cannot be deployed
Maintenance
ActivityStale
ResponsivenessNo issues