Skip to main content
Glama
knownothing20

vision-mcp

vision-mcp

Vision MCP Server for AI Agents - give your text-only model eyes.

Use any OpenAI-compatible vision model API through MCP tools for image analysis, OCR, error diagnosis, diagram understanding, and chart analysis.

License: MIT Version

中文文档


One-line Summary

Give AI Agents visual understanding through MCP tools. Works with any OpenAI-compatible vision API (GLM-4V, Qwen-VL, MiMo-V2.5, GPT-4o, etc.).

Analogy: Vision MCP is the eyes for text-only models like GLM-5.1.


Related MCP server: Vision MCP Server

Architecture

User sends image path or URL
  -> AI Agent (text model, e.g. GLM-5.1)
  -> vision-mcp tools (MCP)
  -> Vision Model API (e.g. MiMo-V2.5-Free, Qwen3-VL, GLM-4V)
  -> Returns text description
  -> AI Agent continues with understanding

Features

Tool

Description

image_analysis

General image understanding

image_analysis_url

Analyze image from URL

extract_text

OCR - extract text from screenshots

diagnose_error

Analyze error screenshots, suggest fixes

understand_diagram

Architecture/flow/UML diagram analysis

analyze_chart

Chart and data visualization analysis


Supported Vision APIs

Provider

Model

API Base

Cost

OpenCode Zen (MiMo-V2.5)

mimo-v2.5-free

https://opencode.ai/zen/v1

Free

Zhipu (GLM-4V)

glm-4v

https://open.bigmodel.cn/api/paas/v4

Paid

Qwen-VL (via proxy)

Qwen3-VL-ms

Your proxy URL

Varies

OpenAI

gpt-4o

https://api.openai.com/v1

Paid

Any OpenAI-compatible

Custom

Custom

Varies

MiMo-V2.5 is a native omnimodal model by Xiaomi that supports text, image, video, and audio. The free tier on OpenCode Zen is a great starting point.


Quick Start

1. Clone

git clone https://github.com/knownothing20/vision-mcp.git
cd vision-mcp
npm install

2. Configure

cp local/.env.example local/.env
# Edit local/.env, fill in your API key and model

Minimal config for MiMo-V2.5 Free (OpenCode Zen):

VISION_API_KEY=your-opencode-zen-api-key
VISION_API_BASE=https://opencode.ai/zen/v1
VISION_MODEL=mimo-v2.5-free

3. Register in opencode

Add to opencode.json (replace path with your actual skill dir):

{
  "mcp": {
    "vision-mcp": {
      "command": ["node", "/path/to/vision-mcp/index.js"],
      "enabled": true,
      "type": "local"
    }
  }
}

Or use node sync.cjs to auto-register (recommended).

4. Restart opencode

Restart your AI tool to activate MCP registration.


Configuration

Edit local/.env:

Variable

Description

Default

VISION_API_KEY

API key (required)

-

VISION_API_BASE

API endpoint

https://open.bigmodel.cn/api/paas/v4

VISION_MODEL

Model name

glm-4v

Priority: local/.env > .env (root) > environment variables.

Provider Examples

VISION_API_KEY=sk-your-zen-api-key
VISION_API_BASE=https://opencode.ai/zen/v1
VISION_MODEL=mimo-v2.5-free

Get your API key at https://opencode.ai/auth

VISION_API_KEY=your-zhipu-api-key
VISION_API_BASE=https://open.bigmodel.cn/api/paas/v4
VISION_MODEL=glm-4v
VISION_API_KEY=your-proxy-key
VISION_API_BASE=http://192.168.x.x:8317/v1
VISION_MODEL=Qwen3-VL-ms
VISION_API_KEY=your-openai-api-key
VISION_API_BASE=https://api.openai.com/v1
VISION_MODEL=gpt-4o

Usage

After setup, provide image file paths to your AI agent:

"Analyze this image: C:\Users\you\Desktop\screenshot.png"
"Extract text from /path/to/document.png"
"Diagnose this error: C:\screenshots\error.png"
"What does this architecture diagram mean? /path/to/diagram.png"
"Analyze chart trends: /path/to/chart.png"

Note: Your main AI model (e.g. GLM-5.1) may not support image input directly. Provide the file path instead of embedding the image in chat, and the agent will use vision-mcp tools to analyze it.


Directory Structure

vision-mcp/
├── index.js              # MCP server main program
├── sync.cjs              # Repo <-> skill dir sync tool
├── package.json          # Dependencies
├── SKILL.md              # Agent documentation
├── README.md             # This file
├── README_CN.md          # Chinese documentation
├── AGENT_GUIDE.md        # Agent install guide
├── INSTALL_OTHER_MACHINE.md  # Install on another machine
├── CHANGELOG.md          # Version changelog
├── LICENSE               # MIT License
├── .gitignore
└── local/                # Local private config (gitignored)
    └── .env.example      # Config template

License

MIT License - See LICENSE

Related MCP Connectors

Related MCP Servers

  • F
    license
    A
    quality
    Not graded
    maintenance
    Enables AI agents to analyze images through vision AI providers (Gemini, OpenAI, Claude), performing tasks like image description, object detection with bounding boxes, region-specific analysis, and precise color extraction without consuming context window with raw pixels.
    4
    -