visual-understanding
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@visual-understandingDescribe this image: https://example.com/photo.jpg"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
visual-understanding
多提供商视觉理解工具——MCP 服务器 + CLI 双模式。通过统一的接口调用 智谱 GLM-V、OpenAI GPT-4o、Anthropic Claude 或任何 OpenAI 兼容端点, 完成图片/视频/文档的多模态理解与目标定位。
🚀 零安装:
uvx visual-understanding <command>直接运行(已发布 PyPI)🔌 双模式:MCP 服务器(
serve)+ CLI(analyze/ground/list-providers)🌐 多提供商:Zhipu / OpenAI / Anthropic / 任意 OpenAI 兼容端点,YAML 配置即插即用
👁️ 核心能力:图片/视频/文档理解 + 目标定位(bounding box)
功能
能力 | 说明 | 支持的输入 |
多模态理解 ( | 图片描述、OCR、视觉问答、文档解读、多图对比 | 图片 URL/路径/base64、视频 URL、文档 URL |
目标定位 ( | 定位图像中的目标,输出归一化坐标,可选画框可视化 | 图片 URL/路径/base64 |
提供商查询 ( | 查看已配置的提供商、模型、能力 | — |
Related MCP server: imagine-mcp
快速开始
安装
方式一:uvx 运行(推荐,零安装) —— 已发布到 PyPI:
# 直接运行,无需安装(uv 自动缓存)
uvx visual-understanding list-providers
# 如果默认镜像(如清华 TUNA)尚未同步新包,可临时指定官方索引:
uvx --default-index https://pypi.org/simple visual-understanding list-providers方式二:本地安装:
cd mcp-servers/visual-understanding
pip install -e .配置 API Key
至少设置一个提供商的 API Key(环境变量):
# 智谱(推荐——支持原生定位、视频、文件)
export ZHIPU_API_KEY="your_key" # https://bigmodel.cn/usercenter/proj-mgmt/apikeys
# OpenAI
export OPENAI_API_KEY="your_key" # https://platform.openai.com/api-keys
# Anthropic
export ANTHROPIC_API_KEY="your_key" # https://console.anthropic.com/settings/keys验证安装
visual-understanding list-providers模式一:MCP 服务器
在 ZCode / Claude Desktop / Cursor 等 MCP 客户端中注册:
ZCode (~/.zcode/cli/config.json):
{
"mcpServers": {
"visual-understanding": {
"command": "uvx",
"args": ["visual-understanding", "serve"]
}
}
}💡 如果默认 PyPI 镜像(清华/中科大等)尚未同步最新版本,加上官方索引参数:
{ "command": "uvx", "args": ["--default-index", "https://pypi.org/simple", "visual-understanding", "serve"] }
Claude Desktop (claude_desktop_config.json):
{
"mcpServers": {
"visual-understanding": {
"command": "visual-understanding",
"args": ["serve"]
}
}
}注册后即可在对话中直接使用 vision_analyze、vision_ground、list_providers 工具。
模式二:CLI / Skill
# 描述图片
visual-understanding analyze --images photo.jpg
# OCR 文字提取
visual-understanding analyze --images scan.png --prompt "Extract all text"
# 视觉问答
visual-understanding analyze --images photo.jpg --prompt "What color is the car?"
# 目标定位 + 画框
visual-understanding ground --image photo.jpg --prompt "all people" --visualize --save-path result.png
# 使用特定提供商/模型
visual-understanding analyze --images photo.jpg --provider openai --model gpt-4oAgent 可通过
SKILL.md中的指引调用 CLI。两种模式共享同一套核心逻辑。
提供商配置
内置默认
不创建配置文件时,内置三个提供商:
提供商 | 类型 | 模型 | 视频 | 文件 | 原生定位 |
| OpenAI 兼容 | GLM-V 系列 | ✅ | ✅ | ✅ |
| OpenAI 兼容 | GPT-4o 系列 | ❌ | ❌ | ❌ |
| Anthropic | Claude 系列 | ❌ | ❌ | ❌ |
自定义配置
# 方式一:环境变量指定路径
export VISUAL_UNDERSTANDING_CONFIG=/path/to/config.yaml
# 方式二:默认路径
mkdir -p ~/.config/visual-understanding
cp config.example.yaml ~/.config/visual-understanding/config.yaml配置文件格式见 config.example.yaml。
添加自定义 OpenAI 兼容端点
任何 OpenAI 兼容的视觉模型服务都可以添加(vLLM、Ollama、Together、Azure 等):
providers:
my-vlm:
type: openai_compat
api_key_env: MY_API_KEY # 环境变量名
base_url: http://localhost:8080/v1 # API 地址
chat_models: [qwen-vl-max]
default_chat_model: qwen-vl-max
max_images: 10然后:
export MY_API_KEY="your_key_or_dummy"
visual-understanding analyze --images photo.jpg --provider my-vlm架构
┌─────────────┐
│ config │ YAML / env vars
└──────┬──────┘
│
┌────────────┼────────────┐
▼ ▼ ▼
┌──────────┐ ┌──────────┐ ┌──────────┐
│ ops.py │ │ ops.py │ │ ops.py │ ← 共享业务逻辑
│ do_analyze│ │ do_ground│ │ do_list │
└────┬─────┘ └────┬─────┘ └──────────┘
│ │
┌───────┴──────┐ │
▼ ▼ ▼
┌────────┐ ┌────────┐ ┌──────────┐
│server.py│ │ cli.py │ │grounding │
│ (MCP) │ │ (CLI) │ │ .py │
└───┬────┘ └────────┘ └──────────┘
│
▼
┌──────────────────────────────────┐
│ providers/ │
│ ┌────────────┐ ┌────────────┐ │
│ │openai_compat│ │ anthropic │ │
│ │(OpenAI/智谱)│ │ (Claude) │ │
│ └────────────┘ └────────────┘ │
└──────────────────────────────────┘config.py— Pydantic 配置模型 + YAML 加载(三级查找)media.py— 输入解析(URL/路径/base64 归一化 + SSRF 防护)providers/— 提供商抽象 + 实现(OpenAI 兼容、Anthropic)grounding.py— 定位 prompt 构造、坐标解析、Pillow 画框ops.py— 共享操作逻辑(MCP 工具与 CLI 子命令的唯一调用入口)server.py— FastMCP 服务器(3 个 MCP 工具)cli.py— CLI 入口(4 个子命令:analyze / ground / list-providers / serve)
安全设计
API 密钥始终通过环境变量名引用(
api_key_env),配置文件中不出现明文密钥base_url仅在配置中指定,工具参数不接受覆盖(防止密钥泄露到恶意端点)URL 输入仅允许 http/https 公网地址,拒绝 localhost/内网 IP(防 SSRF)
.gitignore排除config.yaml、.env
技术栈
MCP Python SDK (FastMCP v1.x)
httpx异步 HTTPpydantic配置校验pyyaml配置文件pillowgrounding 画框可视化
License
MIT
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- AlicenseAqualityAmaintenanceMulti-model vision understanding MCP server that provides unified image analysis for AI assistants without native vision, supporting models like GLM-4.6V, DeepSeek-OCR, Qwen3-VL-Flash, and more.13,789100MIT
- AlicenseBqualityAmaintenanceProduction-grade MCP server for image and video understanding and generation across Gemini, OpenAI, and Grok.54Apache 2.0
- AlicenseAqualityBmaintenanceMCP server for image recognition, supporting multiple vision backends (Anthropic, Zhipu, Ollama) to describe, answer questions, and analyze images.3461MIT
- Alicense-qualityBmaintenanceLocal MCP server that provides multi-modal vision capabilities to single-modal base models via API, supporting multi-turn iterative image recognition and document image parsing.5Apache 2.0
Related MCP Connectors
MCP server for MiniMax H3 multimodal video generation
MCP server for Wan AI video generation
Multimodal video analysis MCP — transcription, vision, and OCR for any video URL.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/JayceVane/visual-understanding'
If you have feedback or need assistance with the MCP directory API, please join our Discord server