vision_mcp
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| VISION_MODEL | No | Vision model name. | qwen-vl-plus |
| VISION_API_KEY | No | API key for the vision model. Not needed for local deployments. | not-needed |
| VISION_API_BASE | No | API base URL for the vision model, without /chat/completions suffix. | http://localhost:8000/v1 |
| VISION_MAX_TOKENS | No | Maximum output tokens per response. | 2000 |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| vision_pingA | Diagnostic: return a test string to verify MCP communication. |
| describe_imageA | 理解并描述图片内容。传入本地图片文件绝对路径,返回对该图片的详细文字描述。当用户让你查看、理解、分析或描述任何图片时,你必须调用此工具。支持 PNG、JPG、JPEG、GIF、WebP、BMP 等常见图片格式。 |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 2 tools
The two tools serve completely distinct purposes: vision_ping is a diagnostic health-check, while describe_image is the core functionality. There is no overlap, and an agent can easily tell which tool to use based on the user's intent.
vision_ping follows a noun-verb pattern with a prefix, while describe_image uses a verb-noun pattern. Both names are clear and readable, but the inconsistent structure makes the set less predictable than a uniform convention.
With only two tools, the server feels thin for a vision MCP. One diagnostic and one core tool is borderline, but it could be acceptable for a minimal, focused server. It lacks the breadth expected from a more complete toolset.
The server covers only image description, with no additional vision capabilities like OCR, object detection, or metadata extraction. For a dedicated vision server, this is a notable gap, but the core describe functionality is present and usable.