glm-vision-mcp
glm-vision-mcp
Servidor MCP que envuelve GLM-4.6V-Flash de 智谱 (modelo de visión gratuito), que expone una herramienta analyze_image a cualquier cliente MCP, con soporte para análisis de una o varias imágenes, OCR y comparación de varias imágenes.
Características
Capacidad | Descripción |
Análisis de imágenes | Ruta local / URL http(s) / data URI base64, se convierte automáticamente a data URI |
Comparación de varias imágenes | Pasa varias imágenes a la vez y compara según el prompt |
Resiliencia a la limitación de velocidad | Reintentos con retroceso exponencial 429 / 1302 / 1305 → sondeo de varias claves → degradación al modelo de respaldo |
Autoverificación de configuración | La herramienta |
Related MCP server: vision-mcp
Requisitos del entorno
Python >= 3.10
Clave API de la plataforma abierta de 智谱 (https://open.bigmodel.cn/usercenter/apikeys),
glm-4.6v-flashgratuitoPara reducir aún más la probabilidad de limitación, puedes registrarte en varias cuentas y obtener una clave de cada una, separadas por comas en la configuración.
Instalación
cd glm-vision-mcp
python -m venv .venv
.venv\Scripts\pip install -r requirements.txtInicio
# stdio 模式(MCP 客户端默认方式)
$env:ZHIPU_API_KEY = "你的Key"
.venv\Scripts\python server.py
# SSE 调试模式(无鉴权,仅限本机)
.venv\Scripts\python server.py --sse 8090Configuración del cliente
Codex (~/.codex/config.toml)
[mcp_servers.glm-vision]
command = "C:\\绝对路径\\glm-vision-mcp\\.venv\\Scripts\\python.exe"
args = ["C:\\绝对路径\\glm-vision-mcp\\server.py"]
[mcp_servers.glm-vision.env]
ZHIPU_API_KEY = "你的Key"
# GLM_VISION_MODELS = "glm-4.6v-flash"
# GLM_API_BASE = "https://open.bigmodel.cn/api/paas/v4/chat/completions"Claude Desktop (claude_desktop_config.json)
{
"mcpServers": {
"glm-vision": {
"command": "C:\\绝对路径\\glm-vision-mcp\\.venv\\Scripts\\python.exe",
"args": ["C:\\绝对路径\\glm-vision-mcp\\server.py"],
"env": { "ZHIPU_API_KEY": "你的Key" }
}
}
}Interfaz de la herramienta
analyze_image(images, prompt, temperature, max_tokens, thinking)
Parámetro | Tipo | Obligatorio | Descripción |
| string[] | Sí | Ruta local / URL http(s) / data URI |
| string | No | Requisitos de análisis, por defecto «describe detalladamente el contenido de esta imagen» |
| number | No | 0.0~1.0, por defecto 0.7 |
| integer | No | Máximo de tokens de salida, por defecto 2048 |
| boolean | No | Modo de pensamiento profundo, por defecto false |
Variables de entorno
Variable | Obligatorio | Descripción |
| Sí | Clave API de 智谱, separada por comas para soportar sondeo de varias claves |
| No | Prioridad de modelos, separados por comas, por defecto |
| No | Sobrescribir el endpoint de la API |
Notas
Imágenes locales de hasta 10 MB cada una, compatibles con jpg/jpeg/png/webp/gif/bmp
El modelo gratuito puede estar limitado en horas punta; solo se produce un error si todas las claves y todos los modelos están limitados. Espera de 15 a 30 segundos y reintenta en otro momento.
Los errores que no son de limitación, como 401/400, no se degradan ni se reintentan; se devuelven directamente para facilitar la localización de problemas de configuración.
Verificación
# 离线检查(不联网)
.venv\Scripts\python test_smoke.py
# 联网冒烟:MCP 握手 + analyze_image 真实调用
$env:ZHIPU_API_KEY = "你的Key"
.venv\Scripts\python test_smoke.py --liveMaintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Tools
Related MCP Servers
- AlicenseAqualityAmaintenanceMulti-model vision understanding MCP server that provides unified image analysis for AI assistants without native vision, supporting models like GLM-4.6V, DeepSeek-OCR, Qwen3-VL-Flash, and more.13,789100MIT
- AlicenseNot gradedqualityCmaintenanceMCP server for analyzing images using multiple vision LLM providers (OpenCode, OpenAI, Anthropic, Google, and custom OpenAI-compatible endpoints). Provides tools to analyze single or multiple images, list providers, and test vision capabilities.MIT
- FlicenseNot gradedqualityCmaintenanceAn MCP server that leverages Zhipu's free GLM-4.6V-Flash vision model to enable image, video, and file understanding (OCR, table parsing, defect detection, document Q&A, and more) across MCP-compatible clients like Codex and Claude Desktop.
- FlicenseAqualityBmaintenanceA Model Context Protocol server that wraps the free GLM-4.6V-Flash vision model, enabling text-only LLM clients like Codex, Cursor, and Claude Desktop to analyze images, videos, and files (PDF/TXT) through standard MCP tools.32
Related MCP Connectors
MCP server for GLM chat completions using Zhipu AI models via AceDataCloud
OCR, transcription, file extraction, and image generation for AI agents via MCP.
MCP server for MiniMax H3 multimodal video generation
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/River831/glm-vision-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server