Skip to main content
Glama
README.md
# Vision-Bridge MCP Server 🧿

[English](./README.en.md) | 中文

给**不支持多模态的模型**(如 DeepSeek、纯文本模型)加一双"眼睛":通过 MCP 暴露
**看图 / OCR** 工具,底层调用任意 **OpenAI 兼容** 的视觉大模型(智谱、千问、OpenAI 等),
把图片理解成文字返回给 AI。

```
你的模型(无视觉)--调用MCP工具--> Vision-Bridge Server --图片--> 视觉大模型
                                          ^ 理解图片          |
用户发图片给你 <--返回文字描述<-----------┘<------------------┘
```

## ✨ 特性

- **通用 OpenAI 兼容**:一套代码,可切换智谱 / 千问 / OpenAI / 任何兼容端点,只改配置不改代码
- **两个工具**:`look_at_image`(看图理解)+ `extract_text_from_image`(OCR 逐字转录)
- **图片预处理**:自动修正方向、压缩限长、转 JPEG,大图不糊细节
- **深度思考**:智谱模型默认开启 thinking 模式,精度优先
- **限流重试**:429 自动指数退避重试,API 错误原样透传
- **零成本可选**:智谱 `glm-4.6v-flash` 或千问新用户 100 万 token 均可免费使用
- **一键安装**:`pip install` 直装,装完即可用 `vision-bridge` 命令

## 🚀 快速开始

### 方式一:pip 一键安装(推荐)

```bash
pip install "vision-bridge-mcp @ git+https://github.com/zgz518/vision-bridge-mcp.git"
```

然后在**你的工作目录**创建 `.env`(见[配置](#-配置)),验证:

```bash
vision-bridge --test                                # 自检:查看配置和工具
vision-bridge --once /path/to/image.png "有什么"     # 命令行单图联调
```

### 方式二:uv

```bash
uv tool install git+https://github.com/zgz518/vision-bridge-mcp
vision-bridge --test
```

### 方式三:克隆 + 虚拟环境

```bash
git clone https://github.com/zgz518/vision-bridge-mcp.git
cd vision-bridge-mcp
python -m venv .venv
# Windows: .venv\Scripts\activate   macOS/Linux: source .venv/bin/activate
pip install -r requirements.txt
cp .env.example .env
python server.py --test
```

### 方式四:直接把链接发给 AI 助手(最省事)

如果你在用 WorkBuddy / Claude Code / Cursor 等 AI 助手,什么都不用自己做——把下面这条消息发给它即可:

> **帮我安装并使用这个 MCP 服务器:https://github.com/zgz518/vision-bridge-mcp**

AI 会自动完成克隆 / 安装依赖 / 生成配置,并引导你提供 API Key(智谱/千问/OpenAI 任选),最后注册到你的 MCP 客户端。你只需要把 Key 交给它。

## ⚙️ 配置

**复制下面的内容,在你的运行目录新建 `.env` 文件并粘贴**(选一家服务商即可,默认智谱免费版):

```bash
# 方式一:智谱(免费,推荐)
# 申请 Key:https://bigmodel.cn 控制台 -> API Keys
VISION_BASE_URL=https://open.bigmodel.cn/api/paas/v4
VISION_API_KEY=你的智谱key
VISION_MODEL=glm-4.6v-flash

# 方式二:阿里云百炼·千问(新用户 100 万 token 免费)
# VISION_BASE_URL=https://dashscope.aliyuncs.com/compatible-mode/v1
# VISION_API_KEY=你的百炼key
# VISION_MODEL=qwen-vl-max

# 方式三:OpenAI
# VISION_BASE_URL=https://api.openai.com/v1
# VISION_API_KEY=sk-xxxx
# VISION_MODEL=gpt-4o-mini
```

服务商速览:

| 服务商 | VISION_BASE_URL | 模型示例 | 费用 |
|--------|-----------------|----------|------|
| 智谱(默认) | `https://open.bigmodel.cn/api/paas/v4` | `glm-4.6v-flash` | 免费 |
| 阿里云百炼·千问 | `https://dashscope.aliyuncs.com/compatible-mode/v1` | `qwen-vl-max` | 新用户 100 万 token 免费 |
| OpenAI | `https://api.openai.com/v1` | `gpt-4o-mini` | 付费 |

> pip/uv 安装时,`.env` 放在你运行 `vision-bridge` 命令的目录;clone 时放在项目根目录。
> 项目内也附有 `.env.example` 模板,clone 用户可直接 `cp .env.example .env` 使用。

## 🔌 接入 MCP 客户端

### WorkBuddy

编辑 `%USERPROFILE%\.workbuddy\mcp.json`(**不带点前缀**),加入:

```json
{
  "mcpServers": {
    "vision-bridge-mcp": {
      "type": "stdio",
      "command": "vision-bridge",
      "args": []
    }
  }
}
```

(clone 方式则 command 填 `.venv/Scripts/python.exe`、args 填 `server.py` 绝对路径)

保存后:连接器管理 → 右上角【自定义连接器】→ 找到 `vision-bridge-mcp` → 点「信任」启用。

### Claude Code / Cline / Cherry Studio 等

以 Claude Code 为例(`.mcp.json` 或项目配置):

```json
{
  "mcpServers": {
    "vision-bridge-mcp": {
      "command": "vision-bridge"
    }
  }
}
```

## 💡 使用技巧(重要)

纯文本模型(尤其是 DeepSeek 系)有一个通病:**上传图片后不主动调工具,反而凭图片路径幻觉出内容**。服务器已内置「严禁幻觉」指令,但如果你的模型仍不自动触发,请在客户端的**全局指令/人设**里加一条常驻规则:

> 当用户上传或粘贴图片,或消息中出现以 .png/.jpg/.jpeg/.webp/.gif 结尾的本地文件路径时,你必须立即调用 look_at_image 工具,把图片路径作为 image 参数传入。在工具返回结果之前,禁止描述任何图片内容(未调工具就描述 = 幻觉)。用户要求提取图片中的文字时,改调 extract_text_from_image。

或者更简单:每次贴图后附带一句 **"看这张图"**,即可触发。

## 🛠 工具说明

| 工具 | 参数 | 用途 |
|------|------|------|
| `look_at_image` | `image`: 本地路径或 http(s) URL;`prompt`: 可选关注点 | 看图,返回结构化中文描述 |
| `extract_text_from_image` | `image`: 本地路径或 http(s) URL | OCR,逐字转录图片中的文字 |

## ⚠️ 注意事项

- **密钥安全**:`.env` 已被 gitignore 排除,**不要**把真实 Key 写进任何会提交的文件;
- **限流**:免费档模型有速率限制(429),已内置重试;高频使用建议充值升级档位或换用付费模型;
- **精度**:免费版为轻量模型,对细节/计数要求高的场景,换旗舰模型(智谱 `glm-4.6v`、千问 `qwen-vl-max`)。

## 📄 License

MIT

TDQS

A4.2/5.0

Scored across 2 tools

Disambiguation4/5

The two tools have clearly distinct primary purposes: look_at_image for general image understanding and extract_text_from_image for exact OCR transcription. There is a minor overlap when an image contains text, but the explicit OCR tool removes most ambiguity.

Naming Consistency5/5

Both tool names follow a consistent verb_noun pattern with lowercase snake_case (look_at_image, extract_text_from_image). The naming is predictable and clearly reflects each tool's function.

Tool Count3/5

With only two tools, the server feels minimally scoped. While the narrow focus on vision tasks is reasonable, the count is at the lower boundary and leaves the set feeling thin rather than comprehensive.

Completeness4/5

The server covers the two most essential vision bridge capabilities—image understanding and text extraction. However, other potentially useful operations like object detection or image comparison are absent, leaving minor gaps in the surface.

Maintenance

ActivitySlowing
ResponsivenessNo issues