Vision MCP for Reasonix
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| VISION_MODEL | No | 模型名(覆盖 profile) | |
| VISION_API_KEY | Yes | API key,大多数profile需要 | |
| VISION_PROFILE | No | 预设供应商,可选值:zhipu, openai, qwen, local | local |
| VISION_BASE_URL | No | OpenAI 兼容端点(覆盖 profile) | |
| VISION_MAX_TOKENS | No | 最大响应 tokens | 4096 |
| VISION_TEMPERATURE | No | 采样温度 | 0.7 |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": true
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| analyze_imageC | Analyze an image using a vision language model. Supports local file paths and URLs. |
| ocr_imageC | Extract text from an image using OCR. Supports plain text, Markdown, and JSON output formats. |
| compare_imagesA | Compare 2-4 images and describe differences/similarities. Supports local file paths and URLs. |
| analyze_videoB | Analyze video content using a vision language model. Requires a model with video support (e.g., Qwen3-VL). |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 4 tools
Each tool targets a distinct visual task: single image analysis, video analysis, image comparison, and OCR. There is no overlap or ambiguity.
All tools follow a consistent verb_noun pattern with snake_case: analyze_image, analyze_video, compare_images, ocr_image.
Four tools cover the essential visual analysis tasks without being too few or excessive, fitting the server's scope well.
The set includes single image analysis, video analysis, image comparison, and OCR, covering key visual capabilities with no obvious gaps.