glm-vision-mcp
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@glm-vision-mcpExtract text from this image: https://example.com/receipt.jpg"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
glm-vision-mcp
An MCP server that wraps the Zhipu GLM-4.6V-Flash (free vision model), exposing an analyze_image tool to any MCP client, with support for single/multi-image analysis, OCR, and multi-image comparison.
Features
Capability | Description |
Image analysis | Local paths / http(s) URLs / base64 data URIs are all supported, automatically converted to data URIs |
Multi-image comparison | Pass multiple images in a single call and compare them according to the prompt |
Rate-limit resilience | 429 / 1302 / 1305 exponential backoff retry → multi-key rotation → fallback to backup model |
Configuration self-check | The |
Related MCP server: vision-mcp
Requirements
Python >= 3.10
Zhipu Open Platform API Key (https://open.bigmodel.cn/usercenter/apikeys),
glm-4.6v-flashis freeTo further reduce the chance of rate limiting, you can register multiple accounts and obtain multiple Keys, separated by commas in the configuration
Installation
cd glm-vision-mcp
python -m venv .venv
.venv\Scripts\pip install -r requirements.txtStartup
# stdio 模式(MCP 客户端默认方式)
$env:ZHIPU_API_KEY = "你的Key"
.venv\Scripts\python server.py
# SSE 调试模式(无鉴权,仅限本机)
.venv\Scripts\python server.py --sse 8090Client Configuration
Codex (~/.codex/config.toml)
[mcp_servers.glm-vision]
command = "C:\\绝对路径\\glm-vision-mcp\\.venv\\Scripts\\python.exe"
args = ["C:\\绝对路径\\glm-vision-mcp\\server.py"]
[mcp_servers.glm-vision.env]
ZHIPU_API_KEY = "你的Key"
# GLM_VISION_MODELS = "glm-4.6v-flash"
# GLM_API_BASE = "https://open.bigmodel.cn/api/paas/v4/chat/completions"Claude Desktop (claude_desktop_config.json)
{
"mcpServers": {
"glm-vision": {
"command": "C:\\绝对路径\\glm-vision-mcp\\.venv\\Scripts\\python.exe",
"args": ["C:\\绝对路径\\glm-vision-mcp\\server.py"],
"env": { "ZHIPU_API_KEY": "你的Key" }
}
}
}Tool Interface
analyze_image(images, prompt, temperature, max_tokens, thinking)
Parameter | Type | Required | Description |
| string[] | Yes | Local path / http(s) URL / data URI |
| string | No | Analysis request, default "Please describe the content of this image in detail" |
| number | No | 0.0~1.0, default 0.7 |
| integer | No | Maximum output tokens, default 2048 |
| boolean | No | Deep thinking mode, default false |
Environment Variables
Variable | Required | Description |
| Yes | Zhipu API Key, comma-separated for multi-key rotation |
| No | Model priority, comma-separated, default |
| No | Override API endpoint |
Notes
Local images must be ≤ 10MB per image, supported formats: jpg/jpeg/png/webp/gif/bmp
Free models may be rate-limited during peak hours; an error is only raised when all Keys + all models are rate-limited. Wait 15~30 seconds and retry at off-peak times
Non-rate-limit errors such as 401 are not retried or downgraded; they are returned directly for easier troubleshooting
Verification
# 离线检查(不联网)
.venv\Scripts\python test_smoke.py
# 联网冒烟:MCP 握手 + analyze_image 真实调用
$env:ZHIPU_API_KEY = "你的Key"
.venv\Scripts\python test_smoke.py --liveMaintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Tools
Related MCP Servers
- AlicenseAqualityAmaintenanceMulti-model vision understanding MCP server that provides unified image analysis for AI assistants without native vision, supporting models like GLM-4.6V, DeepSeek-OCR, Qwen3-VL-Flash, and more.13,789100MIT
- AlicenseNot gradedqualityCmaintenanceMCP server for analyzing images using multiple vision LLM providers (OpenCode, OpenAI, Anthropic, Google, and custom OpenAI-compatible endpoints). Provides tools to analyze single or multiple images, list providers, and test vision capabilities.MIT
- FlicenseNot gradedqualityCmaintenanceAn MCP server that leverages Zhipu's free GLM-4.6V-Flash vision model to enable image, video, and file understanding (OCR, table parsing, defect detection, document Q&A, and more) across MCP-compatible clients like Codex and Claude Desktop.
- FlicenseAqualityBmaintenanceA Model Context Protocol server that wraps the free GLM-4.6V-Flash vision model, enabling text-only LLM clients like Codex, Cursor, and Claude Desktop to analyze images, videos, and files (PDF/TXT) through standard MCP tools.32
Related MCP Connectors
MCP server for GLM chat completions using Zhipu AI models via AceDataCloud
OCR, transcription, file extraction, and image generation for AI agents via MCP.
MCP server for MiniMax H3 multimodal video generation
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/River831/glm-vision-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server