Skip to main content
Glama
README.md
<div align="center">

# 🖼️ vision-mcp

**自托管多模态 VLM 图片识别 MCP 服务器**

TUI 终端粘贴图片 → AI 客户端自动识别返回 · 数据不出内网

[![MCP](https://img.shields.io/badge/Model_Context_Protocol-server-6f42c1?style=flat-square)](https://modelcontextprotocol.io)
[![TypeScript](https://img.shields.io/badge/TypeScript-5.5-3178c6?style=flat-square&logo=typescript&logoColor=white)](https://www.typescriptlang.org/)
[![Node](https://img.shields.io/badge/node-%E2%89%A518-339933?style=flat-square&logo=node.js&logoColor=white)](https://nodejs.org/)
[![Tests](https://img.shields.io/badge/tests-35%2F35-brightgreen?style=flat-square)](#开发)
[![Build](https://img.shields.io/badge/build-tsup-orange?style=flat-square)](#开发)
[![License: MIT](https://img.shields.io/badge/license-MIT-blue?style=flat-square)](./LICENSE)
[![Transport](https://img.shields.io/badge/transport-stdio-4A4A4A?style=flat-square)](#)

<sub>Claude Code · Codex · OpenCode · 任何 MCP 兼容客户端</sub>

</div>

---

## ✨ 为什么用它

| | 优势 | 说明 |
|---|---|---|
| 🔒 | **私有部署,数据不出网** | 直连你自托管的 VLM,图片不经过第三方云 |
| 🔌 | **OpenAI 兼容,后端可换** | vLLM / Ollama / GLM-4V / Qwen-VL 任选,换 base URL 即可,不改代码 |
| 🖼️ | **TUI 粘图即用** | 终端粘贴图片,客户端自动调工具识别,体验对齐智谱图片识别 MCP |
| 🧩 | **四个专用工具** | 通用理解 / OCR / 图表理解 / UI 转码,各带预设 system prompt 与结构化输出 |
| 📥 | **三种图片输入** | 本地路径 · http(s) URL · `data:` URI,客户端给哪种收哪种 |
| 🛡️ | **错误不泄漏** | 错误串仅静态/状态码,绝不把 VLM 响应体或栈泄漏给客户端 |
| ⚡ | **轻量单进程** | stdio,客户端按需拉起子进程,无常驻、无服务端状态 |
| 🔁 | **内置韧性** | 5xx/超时自动重试一次、4xx 不重试、请求超时、图片大小上限 |
| ✅ | **TDD 全覆盖** | 35 个测试 + 端到端往返(假 VLM + InMemoryTransport) |

## 📐 架构

```mermaid
flowchart LR
    A["🖥️ TUI 客户端<br/>(Claude Code / Codex / OpenCode)"] -- stdio JSON-RPC --> B
    subgraph B["vision-mcp (Node, stdio)"]
        direction TB
        C["tools ×4<br/>analyze_image / extract_text /<br/>understand_diagram / ui_to_code"]
        C --> D["analyze()<br/>共享核心"]
        D --> E["imageSource<br/>路径/URL/data-URI → 归一化"]
        D --> F["vlmClient<br/>OpenAI 兼容 + 重试"]
    end
    F -- HTTPS chat/completions --> G["🧠 自托管 VLM<br/>(qwen-vl / glm-4v / ...)"]
    G -- JSON --> B
    B -- tool result --> A
```

## 🛠️ 工具

全部共享 `image_source`(本地路径 | http(s) URL | `data:` URI)。

| 工具 | 专有参数 | 输出 |
|---|---|---|
| `analyze_image` | `prompt`(必填) | 自然语言描述 / 问答 |
| `extract_text` | `prompt?`、`programming_language?` | OCR 文本(代码截图带语言标注) |
| `understand_diagram` | `diagram_type?`(省略或 `auto`)、`prompt?` | 结构化描述 + mermaid/markdown 复刻 |
| `ui_to_code` | `output_type`(`code`/`spec`/`description`)、`framework?`(`html`/`react-tailwind`)、`prompt?` | 对应 code/spec/description |

## 🚀 快速开始

### 克隆并构建

```bash
git clone https://github.com/skyone123/vision-mcp.git
cd vision-mcp
npm install
npm run build      # 产出 dist/index.js + dist/index.d.ts
npm test           # 可选:35/35 测试
```

客户端只用到 `dist/index.js`,**记下它的绝对路径**(下文记作 `$DIST`),配置里要用。

> 例:Linux/macOS `/home/you/vision-mcp/dist/index.js`;Windows `D:/git/vision-mcp/dist/index.js`。

### 环境变量

| 变量 | 默认 | 必填 | 说明 |
|---|---|:---:|---|
| `VLM_BASE_URL` | — | ✅ | OpenAI 兼容 base,如 `http://localhost:8000/v1`(带 `/v1`) |
| `VLM_MODEL` | `qwen-vl-max` | — | 模型名 |
| `VLM_API_KEY` | `""` | — | Bearer token;后端要鉴权才填,留空不带 `Authorization` 头 |
| `VLM_TIMEOUT_MS` | `60000` | — | 单次请求超时 |
| `VLM_MAX_IMAGE_BYTES` | `10485760` | — | 图片上限 10MB |
| `VLM_MAX_TOKENS` | `2048` | — | 返回 token 上限 |

> 缺 `VLM_BASE_URL` 启动即报错退出,不会静默失败。

## 🔧 配置

### 第 1 步 · 判断后端要不要 API key

```bash
curl http://localhost:8000/v1/models
```

- `200` + 模型列表 → **不用 key**
- `401/403` → **要 key**,带 key 再试:`curl http://localhost:8000/v1/models -H "Authorization: Bearer 你的token"`

模型名从返回里挑视觉模型:

```bash
curl -s http://localhost:8000/v1/models | grep '"id"'
```

实测视觉能力能吃图(最关键):

```bash
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer 你的token" \
  -d '{
    "model": "qwen-vl-max",
    "messages": [{"role":"user","content":[
      {"type":"text","text":"一句话描述这张图"},
      {"type":"image_url","image_url":{"url":"https://upload.wikimedia.org/wikipedia/commons/thumb/4/47/PNG_transparency_demonstration_1.png/640px-PNG_transparency_demonstration_1.png"}}
    ]}]
  }'
```

返回正常文字 → 端点可用,照搬这些值填进 `env`。

### 第 2 步 · 写进客户端

> 把下面的 `$DIST` 换成上一步记下的 `dist/index.js` 绝对路径,`command` 用 `node`。

<details>
<summary><b>Claude Code(CLI)</b></summary>

```bash
claude mcp add vision-mcp --scope user \
  --env VLM_BASE_URL=http://localhost:8000/v1 \
  --env VLM_MODEL=qwen-vl-max \
  -- node "$DIST"
```

要 key 就再加一行 `--env VLM_API_KEY=你的token`。

</details>

<details>
<summary><b>cc-switch / 标准 MCP JSON(单条对象格式)</b></summary>

```json
{
  "command": "node",
  "args": ["/absolute/path/to/vision-mcp/dist/index.js"],
  "env": {
    "VLM_BASE_URL": "http://localhost:8000/v1",
    "VLM_MODEL": "qwen-vl-max"
  }
}
```

带 key 就在 `env` 加 `"VLM_API_KEY": "你的token"`。

</details>

<details>
<summary><b>Claude Code 手改配置文件(<code>~/.claude.json</code>)</b></summary>

```jsonc
{
  "mcpServers": {
    "vision-mcp": {
      "command": "node",
      "args": ["/absolute/path/to/vision-mcp/dist/index.js"],
      "env": { "VLM_BASE_URL": "http://localhost:8000/v1", "VLM_MODEL": "qwen-vl-max" }
    }
  }
}
```

</details>

<details>
<summary><b>Codex(<code>~/.codex/config.toml</code>)</b></summary>

```toml
[mcp_servers.vision-mcp]
command = "node"
args = ["/absolute/path/to/vision-mcp/dist/index.js"]
env = { VLM_BASE_URL = "http://localhost:8000/v1", VLM_MODEL = "qwen-vl-max" }
```

</details>

<details>
<summary><b>OpenCode(<code>~/.config/opencode/opencode.json</code> 或项目根 <code>.opencode.json</code>)</b></summary>

```json
{
  "mcp": {
    "vision-mcp": {
      "type": "local",
      "command": ["node", "/absolute/path/to/vision-mcp/dist/index.js"],
      "environment": {
        "VLM_BASE_URL": "http://localhost:8000/v1",
        "VLM_MODEL": "qwen-vl-max"
      }
    }
  }
}
```

> OpenCode 不同版本字段名可能微调,若工具不出现对照其官方 MCP 文档。

</details>

### 第 3 步 · 验证

```bash
claude mcp list          # 应看到 vision-mcp,状态 connected
```

MCP server 无需手动常驻——客户端按需拉起子进程。然后在对话里粘贴一张图问"图里有什么",客户端自动调 `analyze_image`;或显式:

> 用 analyze_image 工具看一下这张图:<粘贴图片>

## 💻 开发

```bash
npm run dev              # tsx 直接跑源码(开发期)
npm run build            # tsup 打包 dist/index.js
npm test                 # vitest,35/35
npx tsc --noEmit         # 类型检查
```

源码结构:

```
src/
  config.ts          # env → VlmConfig
  imageSource.ts     # loadImage: 路径/URL/data-URI 归一化
  vlmClient.ts       # complete: 调 OpenAI 兼容端点 + 重试/超时
  analyze.ts         # 共享核心: loadImage + complete
  server.ts          # McpServer 注册 + stdio + main
  index.ts           # #!/usr/bin/env node 入口
  tools/
    analyzeImage.ts
    extractText.ts
    understandDiagram.ts
    uiToCode.ts
```

每个文件单一职责,可独立测试;四个工具是 `analyze()` 的薄封装,各烘焙自己的 system prompt。

## 🗺️ 路线图(可选扩展)

当前范围:仅 stdio · 单后端 · 单图 · 无持久化。以下为按需扩展项:

| 候选 | 价值 | 建议 |
|---|---|---|
| **流式输出** | `ui_to_code` 输出可能很长,流式能边出边看 | 👍 值得做,UX 提升 |
| **图片预处理** | 发送前按长边缩放/压缩,省 token、降超时 | 👍 值得做,降本 |
| **结构化输出** | `extract_text`/`understand_diagram` 返回 JSON | 🤔 看场景 |
| **HTTP/SSE 传输** | 多客户端共享、远程部署 | 🤔 当前 stdio 够用,按需 |
| **多后端路由** | 不同任务路由到不同 VLM | ❌ YAGNI |
| **视频/多图批处理** | — | ❌ 超出当前定位 |
| **服务端缓存** | 相同图重复识别 | ❌ YAGNI |

## 📄 许可证

[MIT](./LICENSE) © 2026 luyuxin

---

<div align="center">
<sub>自托管 · OpenAI 兼容 · TDD 35/35</sub>
</div>