vlm-mcp
VLM-MCP
VLM-basierter MCP-Server für Bildverständnis. Unterstützt lokales llama.cpp und Online-VLMs (z. B. Qwen3-VL-Flash) über eine einheitliche OpenAI-kompatible API.
English
Funktionen
Doppeltes Backend: lokales llama.cpp + Online-Qwen3-VL-Flash, einheitliche OpenAI-kompatible API
Dreistufiger Cache: L1-Bildkodierungs-Cache, L2-Antwort-Cache (mit TTL), L3-llama-server-KV-Cache
Sitzungsverwaltung: Kontext für mehrteilige Gespräche, automatische Räumung und Timeout-Bereinigung
Prompt-Vorlagen: integriert describe / ocr / chart / translate / qa
Lebenszyklusverwaltung: llama-server-Unterprozess startet/stoppt automatisch mit MCP, keine manuelle Verwaltung erforderlich
Backend-Integrität: Backends bei API-Key-Fehlern automatisch deaktivieren, manuelle Aktivierung/Deaktivierung unterstützt
Bilder aus mehreren Quellen: lokaler Pfad, HTTP-URL, Base64-Data-URI, reines Base64-Fallback
Architektur
MCP Client (SSE :11432)
│
▼
server.py ── tool layer (analyze_image / create_session / ...)
│
├── session_manager.py ── session lifecycle
├── cache.py ── L1 image cache + L2 response cache
├── image_utils.py ── image parsing (path/URL/Base64)
│
▼
providers/ ── OpenAI-compatible interface
│
├── llama-cpp (localhost:11433) ← auto-launched by llama_launcher.py
└── qwen-vl (dashscope API)Schnellstart
Voraussetzungen
Komponente | Hinweise |
Python 3.11+ | Laufzeit |
Paketmanager | |
Native Binärdatei ( | |
Qwen3-VL-8B GGUF | Sprachmodell + visueller Projektor |
Hinweis: Dieses Projekt verwendet die native llama.cpp-Binärdatei (
llama-server), NICHTllama-cpp-python. Keine Python-Bindungen erforderlich – laden Sie einfach die ausführbare Datei von llama.cpp herunter.
Empfohlenes Modell: Laden Sie zwei Dateien von Qwen3-VL-8B-Instruct-GGUF herunter:
Datei | Empfohlen | Hinweise |
Vision-Modell |
| Q4_K_M-Quantisierung, ausgewogenes Verhältnis von Geschwindigkeit und Genauigkeit |
Vision-Projektor |
| Muss F16 sein, nicht quantisieren |
8 GB VRAM reichen aus. Der reine Onlinemodus (nur qwen-vl-Backend) kann llama.cpp und GGUF-Modelle überspringen.
Installation
git clone https://github.com/YC-CLT/VLM-mcp.git
cd VLM-mcp
uv syncKonfiguration
cp config.example.json config.jsonBearbeiten Sie config.json:
{
"backends": {
"llama-cpp": {
"enabled": true,
"base_url": "http://localhost:11433/v1",
"api_key": "sk-no-key-required",
"model_name": "qwen3-vl"
},
"qwen-vl": {
"enabled": false,
"base_url": "https://dashscope.aliyuncs.com/compatible-mode/v1",
"api_key": "your-dashscope-api-key",
"model_name": "qwen-vl-flash"
}
},
"default_backend": "llama-cpp",
"cache_enabled": true,
"llama": {
"server_exe": "llama-server",
"model": "D:/path/to/Qwen3VL-8B-Instruct-Q4_K_M.gguf",
"mmproj": "D:/path/to/mmproj-Qwen3VL-8B-Instruct-F16.gguf",
"ngl": 99
}
}Wichtige Felder:
backends.<name>.enabled: setzen Siefalse, um ein Backend manuell zu deaktivierenllama.model/llama.mmproj: absolute Pfade zu den Modelldateien (erforderlich)llama.ngl: GPU-Ebenen,99= alles auf GPU,0= nur CPUllama.server_exe: ausführbare Datei von llama-server, standardmäßig wird im PATH gesucht
Ausführen
uv run main.pyDer llama-server-Unterprozess startet und stoppt automatisch mit MCP. Keine manuelle Verwaltung erforderlich.
MCP-SSE-Endpunkt: http://127.0.0.1:11432/sse
Von einem beliebigen Verzeichnis ausführen:
uv run --directory D:\CodeFile\VLM-mcp main.py
MCP-Client-Konfiguration
Fügen Sie Ihrer MCP-Client-Konfiguration Folgendes hinzu:
{
"mcpServers": {
"vlm-mcp": {
"url": "http://127.0.0.1:11432/sse"
}
}
}MCP-Tools
Tool | Parameter | Beschreibung |
|
| Bild mit Vorlagen- und Sitzungsunterstützung analysieren |
|
| Sitzung für mehrteilige Gespräche erstellen |
|
| Sitzung schließen |
| — | Alle aktiven Sitzungen auflisten |
| — | Backends und deren Status auflisten |
| — | Verfügbare Prompt-Vorlagen auflisten |
Vorlagen
Vorlage | Parameter | Beschreibung |
| — | Allgemeine Bildbeschreibung |
| — | Textextraktion |
| — | Diagrammanalyse |
|
| Bildübersetzung (Standard: zh) |
|
| Bild-Fragen & Antworten |
Beispiele
// Single analysis
{
"tool": "analyze_image",
"args": {
"image": "D:/photos/cat.png",
"prompt": "What is in this image?"
}
}
// Using template
{
"tool": "analyze_image",
"args": {
"image": "https://example.com/chart.png",
"template": "chart"
}
}
// Multi-turn session
{ "tool": "create_session", "args": { "backend": "llama-cpp" } }
// → { "session_id": "xxx" }
{ "tool": "analyze_image", "args": { "image": "...", "prompt": "...", "session_id": "xxx" } }
{ "tool": "analyze_image", "args": { "prompt": "Tell me more", "session_id": "xxx" } }
{ "tool": "close_session", "args": { "session_id": "xxx" } }Konfigurationskonstanten
Nicht sensible Konstanten in config.py:
Konstante | Standard | Beschreibung |
| 20 | Maximale Bildgröße |
| 10 | Bilddownload-Timeout (s) |
| 100 | L1-Cache-Limit |
| 500 | L2-Cache-Limit |
| 3600 | Online-Backend-Cache-TTL (s) |
| 1800 | Lokales Backend-Cache-TTL (s) |
| 1800 | Sitzungs-Timeout (s) |
| 5 | Maximale Sitzungen pro Backend |
| "INFO" | Protokollebene |
Entwicklung
uv sync --dev
uv run pytest tests/ -vHäufig gestellte Fragen (FAQ)
llama-server läuft auf der CPU?
Überprüfen Sie llama.ngl in config.json – 99 = alles auf GPU, 0 = nur CPU.
llama-server startet nicht?
Stellen Sie sicher, dass server_exe ausführbar ist und die Pfade model/mmproj existieren. Prüfen Sie llama_server.log.
Online-Backend gibt 401 zurück?
Ungültiger API-Key deaktiviert das Backend automatisch. Legen Sie einen gültigen Schlüssel fest und starten Sie neu. Oder setzen Sie "enabled": false, um es zu überspringen.
Portkonflikt?
MCP-Port 11432, llama-server-Port 11433. Ändern Sie llama.port in config.json oder den Port in server.py.
Related MCP server: MCP Vision Server
中文
特性
双后端支持:本地 llama.cpp + 在线 Qwen3-VL-Flash,统一 OpenAI 兼容 API
三层缓存:L1 图片编码缓存、L2 响应缓存(带 TTL)、L3 llama-server KV Cache
会话管理:多轮对话上下文保持,自动淘汰与超时清理
提示词模板:内置 describe / ocr / chart / translate / qa 模板
生命周期管理:llama-server 子进程与 MCP 同起同停,启动即用,无需手动管理
后端健康:API Key 错误自动禁用后端,支持手动启用/禁用
多图片来源:本地路径、HTTP URL、Base64 Data URI、纯 Base64 回退
架构
MCP Client (SSE :11432)
│
▼
server.py ── 工具层 (analyze_image / create_session / ...)
│
├── session_manager.py ── 会话生命周期
├── cache.py ── L1 图片缓存 + L2 响应缓存
├── image_utils.py ── 图片解析 (路径/URL/Base64)
│
▼
providers/ ── OpenAI 兼容接口
│
├── llama-cpp (localhost:11433) ← llama_launcher.py 自动启动
└── qwen-vl (dashscope API)快速开始
环境要求
注意:本项目使用 llama.cpp 原生二进制(
llama-server),不是llama-cpp-python。 无需安装 Python 绑定(即无需llama-cpp-python,这个和单llama.cpp相互独立),只需下载 llama.cpp 可执行文件即可。
推荐模型下载:从 Qwen3-VL-8B-Instruct-GGUF 下载两个文件:
文件 | 推荐 | 说明 |
视觉模型 |
| Q4_K_M 量化,平衡速度与精度 |
图像编码器 |
| 建议 F16,不必量化 |
这样8G显存就可以跑
纯在线模式(仅用 qwen-vl 后端)可跳过 llama.cpp 和 GGUF 模型。
安装
git clone https://github.com/YC-CLT/VLM-mcp.git
cd VLM-mcp
uv sync配置
cp config.example.json config.json编辑 config.json:
{
"backends": {
"llama-cpp": {
"enabled": true,
"base_url": "http://localhost:11433/v1",
"api_key": "sk-no-key-required",
"model_name": "qwen3-vl"
},
"qwen-vl": {
"enabled": false,
"base_url": "https://dashscope.aliyuncs.com/compatible-mode/v1",
"api_key": "your-dashscope-api-key",
"model_name": "qwen-vl-flash"
}
},
"default_backend": "llama-cpp",
"cache_enabled": true,
"llama": {
"server_exe": "llama-server",
"model": "D:/path/to/Qwen3VL-8B-Instruct-Q4_K_M.gguf",
"mmproj": "D:/path/to/mmproj-Qwen3VL-8B-Instruct-F16.gguf",
"ngl": 99
}
}关键字段:
backends.<name>.enabled:设为false可手动禁用后端llama.model/llama.mmproj:本地模型文件绝对路径(必填)llama.ngl:GPU 层数,99表示全部 offload 到 GPU,0为纯 CPUllama.server_exe:llama-server 可执行文件,默认从 PATH 查找
运行
uv run main.py启动后会自动拉起 llama-server 子进程,MCP 退出时自动停止。无需手动管理 llama-server。
MCP SSE 端点:http://127.0.0.1:11432/sse
从任意目录运行:
uv run --directory D:\CodeFile\VLM-mcp main.py
MCP 客户端配置
在你的 MCP 客户端配置文件中添加:
{
"mcpServers": {
"vlm-mcp": {
"url": "http://127.0.0.1:11432/sse"
}
}
}MCP 工具
工具 | 参数 | 说明 |
|
| 分析图片,支持模板和会话 |
|
| 创建多轮对话会话 |
|
| 关闭会话 |
| — | 列出所有活跃会话 |
| — | 列出后端及其状态 |
| — | 列出可用提示词模板 |
模板
模板 | 参数 | 说明 |
| — | 通用图片描述 |
| — | 文字提取 |
| — | 图表分析 |
|
| 图片翻译(默认中文) |
|
| 图片问答 |
使用示例
// 单次分析
{
"tool": "analyze_image",
"args": {
"image": "D:/photos/cat.png",
"prompt": "这张图片里有什么?"
}
}
// 使用模板
{
"tool": "analyze_image",
"args": {
"image": "https://example.com/chart.png",
"template": "chart"
}
}
// 多轮会话
{ "tool": "create_session", "args": { "backend": "llama-cpp" } }
// → { "session_id": "xxx" }
{ "tool": "analyze_image", "args": { "image": "...", "prompt": "...", "session_id": "xxx" } }
{ "tool": "analyze_image", "args": { "prompt": "继续分析", "session_id": "xxx" } }
{ "tool": "close_session", "args": { "session_id": "xxx" } }配置常量
非敏感常量集中于 config.py,可在代码中直接修改:
常量 | 默认值 | 说明 |
| 20 | 图片最大体积 |
| 10 | 图片下载超时(秒) |
| 100 | L1 缓存上限 |
| 500 | L2 缓存上限 |
| 3600 | 在线后端缓存 TTL(秒) |
| 1800 | 本地后端缓存 TTL(秒) |
| 1800 | 会话超时(秒) |
| 5 | 每后端最大会话数 |
| "INFO" | 日志级别 |
开发
uv sync --dev
uv run pytest tests/ -v常见问题
llama-server 跑在 CPU 上?
检查 config.json 中 llama.ngl 是否为 99(全 GPU),0 为纯 CPU。
llama-server 启动失败?
确认 server_exe 可执行(PATH 中或绝对路径),model/mmproj 路径存在。查看 llama_server.log。
在线后端 401 错误?
API Key 无效时会自动禁用该后端,设好 Key 后重启即可恢复。也可手动设 "enabled": false 跳过。
端口被占用?
MCP 端口 11432,llama-server 端口 11433。修改 config.json 中 llama.port 或 server.py 中端口号。
许可
MIT
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables image analysis and understanding using Vision Language Models through OpenAI-compatible APIs. Supports analyzing images from URLs or local files with custom prompts.12MIT
- AlicenseAqualityDmaintenanceProvides advanced image analysis capabilities including object recognition, OCR text extraction, and multi-turn visual dialogues using OpenAI-compatible APIs. It supports both local files and Base64 inputs with additional features for session persistence and web-based configuration management.3MIT
- FlicenseNot gradedqualityCmaintenanceEnables image recognition using vision models via OpenAI-compatible APIs, supporting multiple platforms like OpenAI, DeepSeek, and Ollama.
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to analyze images using any OpenAI-compatible vision API, providing tools for image analysis, OCR, error diagnosis, diagram understanding, and chart analysis.MIT
Related MCP Connectors
LLM chat, text summarization and AI image generation
OCR, transcription, file extraction, and image generation for AI agents via MCP.
Universal AI API Orchestrator — 1,554 tools, 96 services. One install.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/YC-CLT/VLM-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server