Vison-MCP
# Vison-MCP
> MCP server for vision AI — screenshots to code, OCR, error diagnosis, and image analysis via OpenAI-compatible APIs.
## Supported Tools
| Tool | Description |
|------|-------------|
| `image_analysis` | General visual understanding — describe any image in detail |
| `extract_text_from_screenshot` | OCR optimized for terminals, code, documents, and general content |
| `ui_to_artifact` | Convert UI screenshots into code, prompts, specs, or descriptions |
| `diagnose_error_screenshot` | Analyze error screenshots and propose actionable fixes |
| `understand_technical_diagram` | Interpret architecture diagrams, flowcharts, UML, ER, and system diagrams |
| `analyze_data_visualization` | Read charts and dashboards to surface insights, trends, and anomalies |
| `ui_diff_check` | Compare two UI screenshots to flag visual differences and implementation drift |
| `video_analysis` | Inspect videos (MP4/MOV/M4V) — scene detection, event analysis, content summarization |
## Installation
```bash
git clone https://github.com/Lin-zhibo/Vison-MCP.git
cd Vison-MCP
npm install
npm run build
```
## Configuration
Set the following environment variables:
| Variable | Required | Default | Description |
|----------|----------|---------|-------------|
| `VISIONAI_API_KEY` | Yes | — | API authentication key |
| `VISIONAI_BASE_URL` | Yes | — | OpenAI-compatible API endpoint |
| `VISIONAI_MODEL_NAME` | No | `gpt-4o` | Vision model to use |
## Usage
### With Claude Code
Add to your `.claude/settings.json` or `claude_desktop_config.json`:
```json
{
"mcpServers": {
"vison-mcp": {
"command": "node",
"args": ["/path/to/Vison-MCP/dist/index.js"],
"env": {
"VISIONAI_API_KEY": "your-api-key",
"VISIONAI_BASE_URL": "https://api.openai.com/v1",
"VISIONAI_MODEL_NAME": "gpt-4o"
}
}
}
}
```
### Local Development
```bash
# Copy environment template
cp .env.example .env
# Edit .env with your API credentials
# Build and run
npm run build
npm start
```
## Requirements
- Node.js >= 18.0.0
- An OpenAI-compatible vision API endpoint (GPT-4o, Claude Vision, or compatible)
## License
MIT
TDQS
Scored across 8 tools
Each tool has a distinct purpose: analyzing data visualizations, diagnosing error screenshots, extracting text, general image analysis, UI diff checking, UI-to-artifact conversion, technical diagram interpretation, and video analysis. No two tools overlap in function.
Names use mixed conventions: some follow verb_noun (e.g., analyze_data_visualization, diagnose_error_screenshot), others are noun_noun or use abbreviations (e.g., image_analysis, ui_diff_check, ui_to_artifact). This inconsistency may cause confusion.
With 8 tools, the server covers a broad range of vision tasks without being overwhelming. The count is slightly on the lower side but still well-scoped for a vision-focused MCP.
The tool set covers major vision domains: data visualization, error diagnosis, text extraction, general analysis, UI comparison, code generation, technical diagrams, and video. Minor gaps like object detection or facial recognition are acceptable for a general-purpose server.