glm-vision-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@glm-vision-mcpExtract text from this image: https://example.com/receipt.jpg"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
glm-vision-mcp
An MCP server that wraps the Zhipu GLM-4.6V-Flash (free vision model), exposing an analyze_image tool to any MCP client, with support for single/multi-image analysis, OCR, and multi-image comparison.
Features
Capability | Description |
Image analysis | Local paths / http(s) URLs / base64 data URIs are all supported, automatically converted to data URIs |
Multi-image comparison | Pass multiple images in a single call and compare them according to the prompt |
Rate-limit resilience | 429 / 1302 / 1305 exponential backoff retry → multi-key rotation → fallback to backup model |
Configuration self-check | The |
Related MCP server: vision-mcp
Requirements
Python >= 3.10
Zhipu Open Platform API Key (https://open.bigmodel.cn/usercenter/apikeys),
glm-4.6v-flashis freeTo further reduce the chance of rate limiting, you can register multiple accounts and obtain multiple Keys, separated by commas in the configuration
Installation
cd glm-vision-mcp
python -m venv .venv
.venv\Scripts\pip install -r requirements.txtStartup
# stdio 模式(MCP 客户端默认方式)
$env:ZHIPU_API_KEY = "你的Key"
.venv\Scripts\python server.py
# SSE 调试模式(无鉴权,仅限本机)
.venv\Scripts\python server.py --sse 8090Client Configuration
Codex (~/.codex/config.toml)
[mcp_servers.glm-vision]
command = "C:\\绝对路径\\glm-vision-mcp\\.venv\\Scripts\\python.exe"
args = ["C:\\绝对路径\\glm-vision-mcp\\server.py"]
[mcp_servers.glm-vision.env]
ZHIPU_API_KEY = "你的Key"
# GLM_VISION_MODELS = "glm-4.6v-flash"
# GLM_API_BASE = "https://open.bigmodel.cn/api/paas/v4/chat/completions"Claude Desktop (claude_desktop_config.json)
{
"mcpServers": {
"glm-vision": {
"command": "C:\\绝对路径\\glm-vision-mcp\\.venv\\Scripts\\python.exe",
"args": ["C:\\绝对路径\\glm-vision-mcp\\server.py"],
"env": { "ZHIPU_API_KEY": "你的Key" }
}
}
}Tool Interface
analyze_image(images, prompt, temperature, max_tokens, thinking)
Parameter | Type | Required | Description |
| string[] | Yes | Local path / http(s) URL / data URI |
| string | No | Analysis request, default "Please describe the content of this image in detail" |
| number | No | 0.0~1.0, default 0.7 |
| integer | No | Maximum output tokens, default 2048 |
| boolean | No | Deep thinking mode, default false |
Environment Variables
Variable | Required | Description |
| Yes | Zhipu API Key, comma-separated for multi-key rotation |
| No | Model priority, comma-separated, default |
| No | Override API endpoint |
Notes
Local images must be ≤ 10MB per image, supported formats: jpg/jpeg/png/webp/gif/bmp
Free models may be rate-limited during peak hours; an error is only raised when all Keys + all models are rate-limited. Wait 15~30 seconds and retry at off-peak times
Non-rate-limit errors such as 401 are not retried or downgraded; they are returned directly for easier troubleshooting
Verification
# 离线检查(不联网)
.venv\Scripts\python test_smoke.py
# 联网冒烟:MCP 握手 + analyze_image 真实调用
$env:ZHIPU_API_KEY = "你的Key"
.venv\Scripts\python test_smoke.py --liveAvailable Tools
2 toolsanalyze_imageA
调用智谱 GLM-4.6V-Flash 分析一张或多张图片。
适用:图片描述、OCR 文字提取、表格/图表解析、UI 截图理解、多图对比等。
Args: images: 图片列表,每项是本地文件路径、http(s):// URL 或 base64 data URI。 prompt: 对图片的提问或分析指令,如 “提取图中所有文字,保留排版顺序”。 temperature: 采样温度 0.0~1.0,越小越确定,默认 0.7。 max_tokens: 最大输出 token 数,默认 2048。 thinking: 是否开启深度思考(返回附带思考过程),默认 False。
Returns: dict,含 model / content(/ reasoning / elapsed_ms)字段;失败抛异常。
| Name | Required | Description | Default |
|---|---|---|---|
| images | Yes | ||
| prompt | No | 请详细描述这张图片的内容。 | |
| thinking | No | ||
| max_tokens | No | ||
| temperature | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
无注解,描述承担了行为透明度责任。它说明了返回 dict 含 model/content/reasoning/elapsed_ms 字段、失败时会抛异常、thinking 开启时附带思考过程。但未提及图片格式/大小限制、网络依赖或外部 API 调用风险,留有轻微缺口。
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
描述采用目的/适用/Args/Returns 分段,结构清晰、无冗余。示例 prompt“提取图中所有文字,保留排版顺序”提供了额外价值,且整体长度适中。
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
对于 5 参数、无输出 schema、无注解的工具,描述覆盖了所有参数、返回值、异常行为和使用场景,已相当完整。但缺少图片大小限制、网络连接要求等边界条件,存在小幅遗漏。
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema 描述覆盖率为 0%,描述完全补偿了这一点:images 明确支持本地路径、URL、base64;prompt 给出语义和示例;temperature 说明范围与确定性影响;max_tokens 和 thinking 均给出默认值和含义。参数解释远远超出 schema 标题。
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
描述明确以“调用智谱 GLM-4.6V-Flash 分析一张或多张图片”给出具体动词、资源和模型,并列出图片描述、OCR、表格解析等使用场景,与兄弟工具 check_config 区分明显。目的清晰且无歧义。
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
描述提供了“适用:图片描述、OCR 文字提取、表格/图表解析、UI 截图理解、多图对比等”场景列表,说明何时使用。但未明确点名替代工具或给出“不要用于……”的排除条件;不过兄弟工具仅 check_config,上下文已足够区分。
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_configA
检查 MCP 配置是否就绪(不泄露 Key 本体)。
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries full responsibility. It adds value by stating that the key itself is not leaked, which addresses a security concern. However, it does not clarify side effects, read-only nature, or error behavior, leaving some transparency gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, clear sentence with no unnecessary words. It efficiently conveys the core functionality and a key constraint, earning a perfect score for conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simplicity (no parameters, no output schema, low complexity), the description is reasonably complete. It states the purpose and a safety constraint, though it could mention what 'ready' means or the expected output format. Still, it suffices for a basic config check.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has no parameters, and the schema coverage is 100% (empty). According to the baseline for 0 parameters, a score of 4 is appropriate. The description does not need to add param details as there are none.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: checking MCP configuration readiness, while explicitly noting that it avoids leaking the key. It is distinct from the sibling tool 'analyze_image' which focuses on image analysis.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for verifying configuration, but it does not explicitly state when to use this tool over others (e.g., 'use when needing to verify config'). Given the minimal sibling set, it is somewhat clear, but lacks explicit guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v1.0.0- First observed
analyze_image - First observed
check_config
TDQS
Scored across 2 tools
The two tools are completely distinct: analyze_image handles all image analysis tasks while check_config verifies setup. There is zero overlap or ambiguity in their purposes.
Both tools follow the consistent verb_noun snake_case convention—analyze_image and check_config. The pattern is clean and predictable, matching the higher-scoring examples.
With only 2 tools in a server, the count falls into the 'feels thin' category. While the core vision analysis capability is covered by a single tool, the overall server feels minimal and could reasonably be expected to have a few more supporting tools to feel substantial.
The single analysis tool fully covers the core domain of image description, OCR, and chart parsing—no obvious dead ends for standard usage. The config check provides useful operational support. Minor gaps like batch processing or model inquiry are absent but not critical for the stated scope.
Maintenance
Related MCP Connectors
MCP server for Qwen Image 3 AI image generation
MCP server for GLM chat completions using Zhipu AI models via AceDataCloud
Focused MCP server for OpenAI image/audio generation (v2.0.0). Wraps endpoints via HAPI CLI.
OCR, transcription, file extraction, and image generation for AI agents via MCP.
Related MCP Servers
- AlicenseAqualityAmaintenanceMulti-model vision understanding MCP server that provides unified image analysis for AI assistants without native vision, supporting models like GLM-4.6V, DeepSeek-OCR, Qwen3-VL-Flash, and more.1367 npm113MIT
- AlicenseNot gradedqualityCmaintenanceMCP server for analyzing images using multiple vision LLM providers (OpenCode, OpenAI, Anthropic, Google, and custom OpenAI-compatible endpoints). Provides tools to analyze single or multiple images, list providers, and test vision capabilities.MIT
- FlicenseNot gradedqualityCmaintenanceAn MCP server that leverages Zhipu's free GLM-4.6V-Flash vision model to enable image, video, and file understanding (OCR, table parsing, defect detection, document Q&A, and more) across MCP-compatible clients like Codex and Claude Desktop.-
- FlicenseAqualityBmaintenanceA Model Context Protocol server that wraps the free GLM-4.6V-Flash vision model, enabling text-only LLM clients like Codex, Cursor, and Claude Desktop to analyze images, videos, and files (PDF/TXT) through standard MCP tools.32-