glm4v-vision-mcp
This server provides AI-powered image analysis via the GLM-4V Flash model, supporting three main functions:
Analyze Image (
analyze_image): Send an image and a custom prompt to get natural language analysis, such as describing scenes, identifying objects, or answering questions. Supports PNG, JPG, JPEG, WebP, and GIF formats.Extract Text (
extract_text): Perform OCR to extract text from images, with language detection for Chinese, English, or automatic mode.Describe Image (
describe_image): Generate structured image descriptions in multiple styles — detailed, concise, poetic, or technical — ideal for accessibility and annotation.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@glm4v-vision-mcpExtract the text from the image at ./screenshot.png"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
GLM-4V Flash MCP Server
基于智谱 AI GLM-4V Flash 模型的图片识别 MCP (Model Context Protocol) 服务器。
✨ 功能
图片分析 (
analyze_image) - 使用自然语言分析图片内容文字提取 (
extract_text) - OCR 功能,支持中英文图片描述 (
describe_image) - 生成图片描述,支持多种风格
Related MCP server: image-mcp
📦 安装
前置要求
Node.js 18+
智谱 AI API Key (点此获取 API key)
快速开始
# 1. 克隆仓库
git clone https://github.com/GLM-4V-Flash-MCP/glm-4v-flash-mcp.git
cd glm-4v-flash-mcp
# 2. 安装依赖
npm install
# 3. 设置 API Key
export ZHIPU_API_KEY="your-api-key-here"
# 4. 运行服务器
npm start🔧 配置
集成到 Claude Code
在 ~/.claude/settings.json 中添加:
{
"mcpServers": {
"glm-4v-flash": {
"command": "node",
"args": ["/path/to/server.js"],
"env": {
"ZHIPU_API_KEY": "${ZHIPU_API_KEY}"
},
"workingDirectory": "/path/to"
}
}
}集成到 VS Code
在 .vscode/settings.json 中添加:
{
"mcpServers": {
"glm-4v-flash": {
"command": "node",
"args": ["server.js"],
"env": {
"ZHIPU_API_KEY": "${env:ZHIPU_API_KEY}"
}
}
}
}集成到其他 MCP 客户端
在 .mcp/mcp.json 中配置(已包含在项目中):
{
"mcpServers": {
"glm-4v-flash": {
"command": "node",
"args": ["server.js"],
"env": {
"ZHIPU_API_KEY": "${ZHIPU_API_KEY}"
},
"cwd": "${ZHIPU_MCP_DIR}"
}
}
}使用前设置环境变量:
export ZHIPU_API_KEY="your-api-key-here"
export ZHIPU_MCP_DIR="/path/to/glm-4v-flash-mcp"🛠️ 工具说明
analyze_image
分析图片内容,支持自定义提示词。
参数:
image_path(必填): 图片文件路径prompt(可选): 分析提示词
示例:
{
"name": "analyze_image",
"arguments": {
"image_path": "/path/to/image.jpg",
"prompt": "描述图片中的场景和人物"
}
}extract_text
从图片中提取文字(OCR)。
参数:
image_path(必填): 图片文件路径language(可选): 文字语言,可选chinese、english、auto
示例:
{
"name": "extract_text",
"arguments": {
"image_path": "/path/to/image.png",
"language": "chinese"
}
}describe_image
生成图片的详细描述。
参数:
image_path(必填): 图片文件路径style(可选): 描述风格,可选detailed、concise、poetic、technical
示例:
{
"name": "describe_image",
"arguments": {
"image_path": "/path/to/image.jpg",
"style": "detailed"
}
}📋 支持的图片格式
PNG
JPG / JPEG
WebP
GIF
BMP
🔒 环境变量
变量名 | 说明 | 必需 |
| 智谱 AI API Key | 是 |
🚀 开发
# 安装依赖
npm install
# 开发模式(自动重载)
npm run dev
# 运行测试
npm test📄 许可证
MIT License - 详见 LICENSE 文件
🔗 相关链接
关于作者
大强同学 — 科技博主,也是一名github开源作者,非科班出身,以实践驱动开发,践行Build in Public成长理念,深耕 Windows效率生态,擅长将AI Agent从构想转化为可落地的实用方案,我坚信AI与智能体将重塑个人做事方式,愿以自身技术积累,助力个体把握智能时代机遇,高效提升自身创作、办公与成长效率。
平台 | 链接 |
🌐 官网 | |
𝕏 Twitter | |
📺 B站 | |
▶️ YouTube | |
💬 公众号 | 微信搜「大强同学」或扫码关注 ↓ |

Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- Alicense-qualityDmaintenanceAn MCP server for analyzing images using ModelScope's vision models. Supports both local files and URLs, enabling image content description and question answering.Last updated558MIT
- AlicenseAqualityBmaintenanceMCP server for image recognition, supporting multiple vision backends (Anthropic, Zhipu, Ollama) to describe, answer questions, and analyze images.Last updated3471MIT
- Flicense-qualityCmaintenanceAn MCP server that leverages Zhipu's free GLM-4.6V-Flash vision model to enable image, video, and file understanding (OCR, table parsing, defect detection, document Q&A, and more) across MCP-compatible clients like Codex and Claude Desktop.Last updated
- Alicense-qualityCmaintenanceMCP server that provides a 'borrowed eye' for text-only LLMs, enabling them to identify and describe local images via the Qwen VL vision model, including face recognition, scene description, OCR, and targeted visual questioning.Last updated1Apache 2.0
Related MCP Connectors
MCP server for GLM chat completions using Zhipu AI models via AceDataCloud
MCP server for ByteDance Seedream AI image generation
MCP server for Hailuo (MiniMax) AI video generation
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/dqtx760/glm4v-vision-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server