MiniMax Vision MCP
by wenjiaqi8255
README.md
# MiniMax Vision MCP
将 MiniMax VL 多模态能力通过 MCP 暴露给 Claude Code(或其他 MCP 客户端),补足 DeepSeek v4 flash 等非多模态模型的读图能力。
## 功能
- **`understand_image`** — 分析单张或多张图片,返回文字描述
- 支持本地文件路径、http(s) URL、base64 data URL
- 格式:JPEG、PNG、WebP、GIF
- 支持两种 MiniMax API 模式(通过 `MINIMAX_ENDPOINT` 切换)
## 前置条件
- Python 3.10+
- MiniMax API Key(开通了 VL 模型的 token plan)
## 快速开始
### 1. 安装
```bash
cd ~/Downloads/minimax-vision-mcp
pip install -e .
```
或用 `uv`(推荐):
```bash
cd ~/Downloads/minimax-vision-mcp
uv pip install -e .
```
### 2. 配置 Claude Code
编辑 `~/.claude/settings.local.json`(或项目的 `.claude/settings.json`),添加:
```json
{
"mcpServers": {
"minimax-vision": {
"command": "uv",
"args": [
"run",
"--directory",
"/Users/wenjiaqi/Downloads/minimax-vision-mcp",
"minimax-vision-mcp"
],
"env": {
"MINIMAX_API_KEY": "your-api-key-here",
"MINIMAX_ENDPOINT": "chat_completion"
}
}
}
}
```
> **Claude Desktop 也兼容**:上述 stdio 配置格式可直接用于 `claude_desktop_config.json`。
### 3. 重启 Claude Code
重启后,在对话中发送图片或图片路径,Claude 会自动调用 `understand_image` 工具来分析图片。
## 环境变量
| 变量 | 必需 | 默认值 | 说明 |
|------|------|--------|------|
| `MINIMAX_API_KEY` | ✅ | — | MiniMax API 密钥 |
| `MINIMAX_API_HOST` | ❌ | `https://api.minimax.chat` | API 地址(中国区用 `https://api.minimaxi.com`) |
| `MINIMAX_MODEL` | ❌ | `minimax-vl-01` | VL 模型名(仅 `chat_completion` 模式) |
| `MINIMAX_ENDPOINT` | ❌ | `chat_completion` | API 模式:`chat_completion` 或 `coding_plan` |
### 两种 API 模式
**`chat_completion`(默认)**
通用 MiniMax VL API,通过 `/v1/text/chatcompletion_v2` 调用,支持多图、system prompt、temperature 等参数。适用于标准 token plan。
**`coding_plan`**
MiniMax Coding Plan 专有端点 `/v1/coding_plan/vlm`,仅支持单图 + prompt。如果你是 Coding Plan 用户可用此模式。
> **不确定用哪个?** 先试试 `chat_completion`(默认)。如果返回 404 或 auth 错误,切到 `coding_plan`。
## 使用示例
```python
# 分析本地图片
understand_image(
prompt="这张图片里有什么?请详细描述。",
image_path="/Users/wenjiaqi/Downloads/photo.png"
)
# 分析网络图片
understand_image(
prompt="Extract text from this image",
image_url="https://example.com/screenshot.jpg"
)
# 多图对比
understand_image(
prompt="Compare these two UI designs",
image_paths=["/path/to/design1.png", "/path/to/design2.png"]
)
```
## 项目结构
```
minimax-vision-mcp/
├── pyproject.toml
├── README.md
└── src/
└── minimax_vision_mcp/
├── __init__.py
└── server.py # MCP 服务器主文件
```
## License
MIT
TDQS
B3.4/5.0
Scored across 1 tool
Disambiguation5/5
Only one tool exists, so there is no ambiguity between tools. The tool's purpose is clearly defined.
Naming Consistency5/5
The single tool name 'understand_image' follows a clear verb_noun pattern and is descriptive of its function.
Tool Count2/5
With only one tool, the server feels too limited for the implied scope of 'MiniMax Vision MCP'. A vision server would typically offer multiple capabilities.
Completeness2/5
The tool provides only image-to-text analysis. No other operations (e.g., object detection, metadata retrieval) are available, leaving significant gaps for a vision server.
Maintenance
ActivityStale
ResponsivenessNo issues