Vision MCP Server
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| VISION_MODEL | No | Model name | Qwen3-VL-32B |
| VISION_API_KEY | No | API key (optional for local models) | |
| VISION_BASE_URL | Yes | OpenAI-compatible chat completions endpoint | |
| VISION_MAX_TOKENS | No | Max response tokens | 4096 |
| VISION_TEMPERATURE | No | Sampling temperature | 0.7 |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": true
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| analyze_imageC | Analyze an image using a vision language model. Supports local file paths and URLs. |
| ocr_imageC | Extract text from an image using OCR. Supports plain text, Markdown, and JSON output formats. |
| compare_imagesA | Compare 2-4 images and describe differences/similarities. Supports local file paths and URLs. |
| analyze_videoB | Analyze video content using a vision language model. Requires a model with video support (e.g., Qwen3-VL). |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 4 tools
Each tool targets a distinct task: image content analysis, video analysis, image comparison, and text extraction. There is no overlap in purpose, making selection clear.
All tools follow a consistent verb_noun pattern (analyze_image, analyze_video, compare_images, ocr_image), with no mixing of styles.
With 4 tools, the set is focused and not overwhelming. It covers core vision tasks, though a few more (e.g., image generation) could be added for broader scope.
The tools cover image/video analysis, OCR, and comparison. Minor gaps exist (e.g., image metadata extraction), but the surface is sufficient for common vision use cases.