my-own-vision-mcp
Enables AI agents to use OpenAI-compatible vision models such as GPT-4o for image analysis, OCR, structured extraction, image comparison, and screenshot-to-accessibility-tree conversion.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@my-own-vision-mcpAnalyze this error screenshot and explain what's wrong"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
my-own-vision-mcp
A standalone Model Context Protocol (MCP) server that gives AI agents vision capabilities — image analysis, OCR, structured extraction, and image comparison.
It calls any OpenAI-compatible vision API directly (Qwen-VL, GPT-4o, Claude, GLM-4V, etc.) — no Python, no extra services, just Node.js.
Works with: Claude Code · Cursor · Windsurf · Cline · opencode · openclaw · any MCP-compatible client
The idea
Most capable coding agents run on text-only models — fast and cheap, but blind to images. You could switch to a multimodal model for everything, but that's expensive: vision tokens cost 5-20x more than text tokens, and most coding tasks don't need vision at all.
This project takes a different approach:
┌─────────────────────────┐
user request ───▶ │ text-only agent model │ ← cheap, fast, handles 95% of work
│ (opencode / openclaw / │
│ Claude Code / Cursor) │
└──────────┬──────────────┘
│ "I need to see this image"
│ calls MCP tool
▼
┌─────────────────────────┐
│ dedicated vision model │ ← only invoked when needed
│ (Qwen-VL / GPT-4o / │
│ GLM-4V / local vLLM) │
└─────────────────────────┘Extend capabilities — a text-only agent gains on-demand vision: OCR, image description, screenshot-to-UI-tree, structured extraction
Save cost — the expensive vision model is called only when an image is involved, not on every turn
Decouple models — swap the agent model and the vision model independently; use a cheap local model for coding and a powerful cloud model for vision, or vice versa
Related MCP server: Vision MCP Server
Why use this?
AI coding agents (Claude Code, Cursor, Windsurf, Cline, opencode, openclaw, etc.) can't see images. This MCP server bridges that gap by exposing vision tools that the agent can call autonomously:
Scenario | Tool | Example |
Describe a photo / screenshot |
| "What's in this error screenshot?" |
Extract text from images (OCR) |
| Read a scanned document, receipt, or meme |
Extract structured data from an image |
| Pull |
Compare two images |
| "Did the UI change between these two screenshots?" |
Key features
4 tools + 3 prompts covering general vision tasks
Any OpenAI-compatible API — configure your endpoint and key, done
Multiple input formats — file path, base64, or URL
Automatic image preprocessing — resize/compress via
sharp(optional), resolution follows provider capabilityProvider fallback chain — auto-failover across providers with health tracking
Structured logging — stderr + auto-rotating log files, one per host client
Retry with backoff — configurable retry on 429/5xx, empty-response retry, JSON-mode fallback
Zero Python dependency — pure TypeScript/Node.js
Quick start
git clone https://github.com/vectorequa/my-own-vision-mcp.git
cd my-own-vision-mcp
npm install
npm run build1. Configure your API key
Create ~/.config/my-own-vision-mcp/my-own-vision-mcp.json (on Windows: %USERPROFILE%\.config\my-own-vision-mcp\my-own-vision-mcp.json):
{
"llm": {
"providers": {
"qwen": {
"url": "https://your-api-endpoint/v1",
"api_key": "your-actual-api-key"
}
}
}
}This file is deep-merged over the project config.json. Only url and api_key need to be set here; model/max_tokens/timeout come from the project config.
Alternatively, set env var MY_OWN_VISION_MCP_API_KEY.
2. Register with your MCP host
opencode (opencode.json or ~/.config/opencode/opencode.json):
{
"mcp": {
"my-own-vision-mcp": {
"type": "local",
"command": ["node", "dist/index.js"],
"cwd": "/path/to/my-own-vision-mcp",
"environment": {
"MY_OWN_VISION_MCP_CLIENT": "opencode"
}
}
}
}openclaw (~/.openclaw/openclaw.json):
{
"mcp": {
"servers": {
"my-own-vision-mcp": {
"command": "node",
"args": ["dist/index.js"],
"cwd": "/path/to/my-own-vision-mcp",
"transport": "stdio",
"enabled": true,
"env": {
"MY_OWN_VISION_MCP_CLIENT": "openclaw"
}
}
}
}
}Any other MCP-compatible client — use stdio transport, command node dist/index.js, working directory set to the project root.
3. Verify
Ask your agent to call the ping tool. You should get:
{
"status": "ok",
"provider": "qwen",
"model": "your-model-name",
"max_tokens": 16384,
"timeout": 120,
"all_providers": { "qwen": { "model": "your-model-name" } },
"vision": {
"max_image_dim": 1280,
"jpeg_quality": 85
}
}(Full response also includes extra_notes, capabilities, and all providers — call ping to see all fields.)
Tools
Tool | Description | Key params |
| Analyze or compare image(s). Pass single image or array for multi-image. |
|
| OCR: extract all text, preserving layout. Auto-detects language. |
|
| Extract structured JSON guided by a schema. |
|
| Check server health and config. | — |
Prompts (user-invoked workflows)
Prompt | What it does |
| Thin redirect → calls |
| Thin redirect → calls |
| Thin redirect → calls |
All image inputs accept: file path, base64 string, or URL (http/https).
Common parameters
All tools (except ping) accept:
Parameter | Description |
| Max output tokens. 2048=brief, 8192=detailed, 16384=large. |
| Image resolution: |
| Preferred LLM provider name (use |
Configuration
Three-layer merge (low → high priority)
config.json (project) → ~/.config/my-own-vision-mcp/my-own-vision-mcp.json (user) → env varsProject config.json (in repo, non-sensitive)
{
"llm": {
"default_provider": "qwen",
"providers": {
"qwen": {
"enable": true,
"url": "https://your-api-endpoint/v1",
"api_key": "YOUR_API_KEY",
"model": "YOUR_MODEL",
"max_tokens": 16384,
"timeout": 120,
"retry": {
"max_retries": 3,
"max_504_retries": 1,
"base_delay": 1.0,
"max_delay": 30.0,
"jitter": 0.5,
"retry_on_status": [429, 500, 502, 503, 504],
"retry_504_delay": 10.0,
"empty_retries": 3,
"empty_retry_delay": 1.5
}
}
}
},
"vision": {
"max_image_dim": 1280,
"jpeg_quality": 85,
"max_image_size": 20971520,
"url_timeout": 30
},
"logging": {
"max_file_size": 1048576,
"max_files": 10
}
}User config ~/.config/my-own-vision-mcp/my-own-vision-mcp.json (sensitive, not in repo)
{
"llm": {
"providers": {
"qwen": {
"url": "https://your-real-endpoint/v1",
"api_key": "sk-your-real-api-key"
}
}
}
}Key config fields
Field | Default | Description |
| true | Set to false to disable a provider (won't be listed or callable) |
| 4096 | Max output tokens per request |
| 60 | Request timeout in seconds |
| — | Retry config: |
| — | Per-provider vision capabilities: |
| 1280 | Default max image dimension (px) if provider doesn't specify |
| 85 | Default JPEG compression quality if provider doesn't specify |
| 20MB | Max input image file size |
| 30 | Timeout (s) for fetching images from URLs |
| 1MB | Log file rotation threshold |
| 10 | Max rotated log files to keep |
Environment variable overrides
Variable | Purpose |
| Override project config file path |
| Override default provider's API key |
| Client name for log file naming (e.g., |
Logging
Logs go to both stderr and file logs/<client>.log with auto-rotation.
Client name: set via
MY_OWN_VISION_MCP_CLIENTenv var in the host's MCP configOptional: defaults to
default→logs/default.logRotation: file exceeds
logging.max_file_size(default 1MB) → rotates, keeping at mostlogging.max_files(default 10) filesstdout is reserved for MCP protocol — all logs go to stderr only
Image preprocessing (optional)
Install sharp for resize/compress before sending to LLM:
npm install sharpWithout sharp, images are sent as-is (raw base64). With sharp, images are resized to provider.capabilities.max_image_dim (or vision.max_image_dim fallback) and compressed to JPEG.
Development
npm run dev # run via tsx (no build needed)
npm run build # compile to dist/
npm start # run compiled outputTests (hand-written, no framework):
npx tsx test/image-loader-test.ts
npx tsx test/retry-test.ts
npx tsx test/json-utils-test.tsVersioning
This project uses dual versioning:
System | Where | Format | Example | Purpose |
SemVer |
|
|
| Dependency compatibility |
CalVer | Git tag + GitHub release |
|
| Release timeline |
package.jsonversion follows Semantic Versioning — breaking changes bump MAJOR, new features bump MINOR, fixes bump PATCHGit release tags follow calendar versioning —
v2026.09.0is the first release in Sep 2026,v2026.09.1is the second, etc.Each GitHub release title shows both:
v2026.09.0 (SemVer 0.1.2)
Compatible LLM providers
Any endpoint that implements the OpenAI POST /v1/chat/completions format with vision support:
Qwen-VL (Qwen-VL-Max, Qwen2-VL, Qwen3-VL, etc.) via DashScope or self-hosted
OpenAI GPT-4o / GPT-4o-mini
Google Gemini (Gemini 2.0 Flash, Gemini 1.5 Pro) via OpenAI-compatible proxy
GLM-4V (Zhipu AI)
Llama Vision (Llama 3.2 Vision) via Ollama / vLLM
Pixtral (Mistral)
InternVL (OpenVLM)
OpenRouter — any vision model on OpenRouter (free tier supported)
Local models via vLLM, Ollama, LM Studio, etc.
Configure multiple providers in config.json and select per-tool-call via the provider parameter.
Contributing
Issues and PRs welcome! If this project saves you time or tokens, please ⭐ star the repo — it helps others find it.
License
This server cannot be deployed
Maintenance
Related MCP Connectors
Generate images, GIFs, videos, and PDFs from HTML, URLs, or templates — from your AI agent.
Image & PDF tools for AI agents: compress, convert, resize, PDF, AI vision, pipeline.
Give agents eyes on any web page: structured context, and changes explained in plain language.
OCR, transcription, file extraction, and image generation for AI agents via MCP.
Related MCP Servers
- FlicenseAqualityNot gradedmaintenanceEnables AI agents to analyze images through vision AI providers (Gemini, OpenAI, Claude), performing tasks like image description, object detection with bounding boxes, region-specific analysis, and precise color extraction without consuming context window with raw pixels.4-
- AlicenseAqualityDmaintenanceEnables AI agents to analyze images, extract text, compare images, and analyze video through any OpenAI-compatible vision model.4305 npm20MIT
- AlicenseAqualityCmaintenanceGives text-only coding agents the ability to 'see' images, videos, and screenshots by routing them to a vision model and returning structured text.834 npm1MIT
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to analyze images using any OpenAI-compatible vision API, providing tools for image analysis, OCR, error diagnosis, diagram understanding, and chart analysis.MIT