vision-mcp
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| VISION_MCP_MODEL | No | Model to use for vision tasks (default: qwen/qwen3.6-27b, can also be set in config.json) | |
| VISION_MCP_API_KEY | No | Your Groq API key (also can be configured via `setup` command or config.json) | |
| VISION_MCP_BASE_URL | No | Base URL for the API endpoint (can also be set in config.json) |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| vision_seeA | Look at an image and return a text description/answer. You cannot see images yourself — you MUST call this. Default source is the OS clipboard (user copied or pasted a screenshot). Do NOT ask the user to save a file. Call immediately when the user pastes an image, mentions screenshot/clipboard/图片/截图, or you see an [Image] placeholder. image: omit/'clipboard'/data URI/raw base64/https URL/local path (last resort). question: what to extract or answer. |
| vision_statusA | Check vision config and whether the OS clipboard currently holds an image. Use when vision_see fails or before asking the user to copy again. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 2 tools
The two tools have completely separate purposes: 'vision_see' performs the core image analysis and description, while 'vision_status' checks configuration and clipboard state. There is no overlap or ambiguity in their roles.
Both tools follow a consistent 'vision_verb' naming pattern: 'vision_see' and 'vision_status'. The prefix establishes the domain clearly, and the verbs are distinct and descriptive.
With only 2 tools, the server is on the borderline of being too thin for a typical MCP server. While the two tools cover the essential workflow, the set feels minimal and lacks ancillary tools that might be expected (e.g., configuration or image management).
The server covers the core use case of analyzing an image and checking readiness. Minor gaps exist (e.g., no explicit error recovery or format listing), but agents can work around these by combining the existing tools or relying on user assistance.