glm-vision-mcp-server
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| ZHIPUAI_API_KEY | Yes | 智谱 API key(环境变量,勿硬编码) | |
| GLM_VISION_MODEL | No | 底层视觉模型 ID(可选,默认 glm-4.6v-flash) | glm-4.6v-flash |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| understand_imageA | 使用智谱免费视觉模型 glm-4.6v-flash 理解一张或多张图片(本地文件路径或 http(s) 图片 URL)。单图自动详细描述;多图分别描述并指出异同。可附带具体问题 question(中文回答)。 |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 1 tool
With only one tool, there is no possibility of tool confusion. The tool's purpose is clearly defined for image understanding.
The single tool name follows a clear verb_noun pattern, indicating predictability, though with only one tool the pattern is not fully established.
The server has exactly one tool, which feels thin even for a narrow vision domain. The tool itself is comprehensive, but the overall surface is minimal. This is borderline, so score 3.
The tool covers image description, multi-image comparison, and Q&A, which are the core capabilities expected of an image understanding server. No obvious gaps within its intended scope.