mcp-vision-server
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| IMAGE_MODEL | No | 生图模型,回退 VISION_MODEL | |
| VISION_MODEL | Yes | 默认视觉模型 | |
| VISION_API_KEY | Yes | 视觉 Key | |
| VISION_API_BASE_URL | Yes | 视觉 API 根地址 | |
| VISION_BACKUP_MODEL | No | ocr 备用视觉模型,回退 VISION_MODEL | |
| VISION_ANALYZE_MODEL | No | analyze 专用模型,回退 VISION_MODEL |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": true
} |
| logging | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| vision_analyzeA | Understand, explain, or describe an image using a vision model. Use for interpreting screenshots, UI, diagrams, charts, or error messages that need reasoning. Not for verbatim text extraction (use vision_ocr). |
| vision_ocrA | Extract exact text from an image (OCR). Use when the user wants literal text copied verbatim from a screenshot, code image, terminal output, document, or receipt. Not for explaining or understanding images (use vision_analyze). |
| image_generateA | Generate one or more images from a text prompt through an OpenAI-compatible images/generations endpoint. Images are returned inline as base64 when available, otherwise as download URLs. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 3 tools
Each tool has a clearly distinct purpose: vision_analyze for understanding/explaining images, vision_ocr for verbatim text extraction, and image_generate for creating images. The descriptions explicitly cross-reference each other to prevent confusion, leaving no ambiguity.
The naming is mostly consistent with a domain prefix and action (vision_analyze, vision_ocr), but image_generate deviates by using 'image_' instead of 'vision_'. This is a minor inconsistency that does not hinder readability, but it is a noticeable break from the established pattern.
Three tools is a well-scoped count for a vision server, covering analysis, OCR, and generation. Each tool addresses a distinct core capability, and there are no redundant or missing tools that would suggest over- or under-engineering.
For the stated domain of vision tasks, the tool surface is complete: understanding (analyze), text extraction (OCR), and creation (generate). There are no obvious missing operations that would force an agent into a dead end; the tools cover the primary workflows one would expect from a vision server.