Skip to main content
Glama
README.md
# Vison-MCP

> MCP server for vision AI — screenshots to code, OCR, error diagnosis, and image analysis via OpenAI-compatible APIs.

## Supported Tools

| Tool | Description |
|------|-------------|
| `image_analysis` | General visual understanding — describe any image in detail |
| `extract_text_from_screenshot` | OCR optimized for terminals, code, documents, and general content |
| `ui_to_artifact` | Convert UI screenshots into code, prompts, specs, or descriptions |
| `diagnose_error_screenshot` | Analyze error screenshots and propose actionable fixes |
| `understand_technical_diagram` | Interpret architecture diagrams, flowcharts, UML, ER, and system diagrams |
| `analyze_data_visualization` | Read charts and dashboards to surface insights, trends, and anomalies |
| `ui_diff_check` | Compare two UI screenshots to flag visual differences and implementation drift |
| `video_analysis` | Inspect videos (MP4/MOV/M4V) — scene detection, event analysis, content summarization |

## Installation

```bash
git clone https://github.com/Lin-zhibo/Vison-MCP.git
cd Vison-MCP
npm install
npm run build
```

## Configuration

Set the following environment variables:

| Variable | Required | Default | Description |
|----------|----------|---------|-------------|
| `VISIONAI_API_KEY` | Yes | — | API authentication key |
| `VISIONAI_BASE_URL` | Yes | — | OpenAI-compatible API endpoint |
| `VISIONAI_MODEL_NAME` | No | `gpt-4o` | Vision model to use |

## Usage

### With Claude Code

Add to your `.claude/settings.json` or `claude_desktop_config.json`:

```json
{
  "mcpServers": {
    "vison-mcp": {
      "command": "node",
      "args": ["/path/to/Vison-MCP/dist/index.js"],
      "env": {
        "VISIONAI_API_KEY": "your-api-key",
        "VISIONAI_BASE_URL": "https://api.openai.com/v1",
        "VISIONAI_MODEL_NAME": "gpt-4o"
      }
    }
  }
}
```

### Local Development

```bash
# Copy environment template
cp .env.example .env
# Edit .env with your API credentials

# Build and run
npm run build
npm start
```

## Requirements

- Node.js >= 18.0.0
- An OpenAI-compatible vision API endpoint (GPT-4o, Claude Vision, or compatible)

## License

MIT

TDQS

A3.7/5.0

Scored across 8 tools

Disambiguation5/5

Each tool has a distinct purpose: analyzing data visualizations, diagnosing error screenshots, extracting text, general image analysis, UI diff checking, UI-to-artifact conversion, technical diagram interpretation, and video analysis. No two tools overlap in function.

Naming Consistency3/5

Names use mixed conventions: some follow verb_noun (e.g., analyze_data_visualization, diagnose_error_screenshot), others are noun_noun or use abbreviations (e.g., image_analysis, ui_diff_check, ui_to_artifact). This inconsistency may cause confusion.

Tool Count4/5

With 8 tools, the server covers a broad range of vision tasks without being overwhelming. The count is slightly on the lower side but still well-scoped for a vision-focused MCP.

Completeness4/5

The tool set covers major vision domains: data visualization, error diagnosis, text extraction, general analysis, UI comparison, code generation, technical diagrams, and video. Minor gaps like object detection or facial recognition are acceptable for a general-purpose server.

Maintenance

ActivityInactive
ResponsivenessNo issues