visual-intelligence-mcp
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| VI_MODEL | No | Model name (must support vision). Default: minimax-m3. | minimax-m3 |
| VI_API_KEY | Yes | Middleware API key (required). | |
| VI_BASE_URL | Yes | Middleware base URL, e.g., https://api.xxx.com/v1 (required). | |
| VI_MAX_TOKENS | No | Maximum response tokens. Default: 2048. | 2048 |
| VI_IMAGE_QUALITY | No | JPEG quality (1-100). Default: 80. | 80 |
| VI_MAX_IMAGE_SIZE | No | Max image width/height in pixels; larger images are scaled down. Default: 1280. | 1280 |
| VI_MAX_IMAGE_BYTES | No | Max compressed image size; errors if exceeded. Default: 10MB. | 10MB |
| VI_REQUEST_TIMEOUT_MS | No | Request timeout in milliseconds. Default: 60000. | 60000 |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| analyze_imageA | 分析本地图片(UI 自动化视觉辅助)。 当需要"查看屏幕/截图/界面"时使用:传入截图路径与问题,返回多模态模型的文字描述或结构化 JSON。 调用规范(必须遵守):
|
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 1 tool
With only one tool, there is no possibility of confusion or overlap. The tool's purpose—analyzing images to answer questions or return structured data—is clear and unambiguous.
The single tool name follows a clear verb_noun pattern (analyze_image) that is consistent and descriptive. There are no conflicting conventions to cause confusion.
A single tool feels thin for a server branded as 'visual-intelligence,' but the tool itself is versatile enough to handle various image analysis requests. The count is right at the borderline where it could use additional specialized tools, but it is not wholly inappropriate.
The tool covers the core need of analyzing local images and returning either descriptive text or structured JSON, including support for UI-related queries. Minor gaps exist, such as requiring local file paths and lacking support for direct image URLs, but these are workaroundable.