Skip to main content
Glama
README.md
# MiniMax Vision MCP

将 MiniMax VL 多模态能力通过 MCP 暴露给 Claude Code(或其他 MCP 客户端),补足 DeepSeek v4 flash 等非多模态模型的读图能力。

## 功能

- **`understand_image`** — 分析单张或多张图片,返回文字描述
- 支持本地文件路径、http(s) URL、base64 data URL
- 格式:JPEG、PNG、WebP、GIF
- 支持两种 MiniMax API 模式(通过 `MINIMAX_ENDPOINT` 切换)

## 前置条件

- Python 3.10+
- MiniMax API Key(开通了 VL 模型的 token plan)

## 快速开始

### 1. 安装

```bash
cd ~/Downloads/minimax-vision-mcp
pip install -e .
```

或用 `uv`(推荐):

```bash
cd ~/Downloads/minimax-vision-mcp
uv pip install -e .
```

### 2. 配置 Claude Code

编辑 `~/.claude/settings.local.json`(或项目的 `.claude/settings.json`),添加:

```json
{
  "mcpServers": {
    "minimax-vision": {
      "command": "uv",
      "args": [
        "run",
        "--directory",
        "/Users/wenjiaqi/Downloads/minimax-vision-mcp",
        "minimax-vision-mcp"
      ],
      "env": {
        "MINIMAX_API_KEY": "your-api-key-here",
        "MINIMAX_ENDPOINT": "chat_completion"
      }
    }
  }
}
```

> **Claude Desktop 也兼容**:上述 stdio 配置格式可直接用于 `claude_desktop_config.json`。

### 3. 重启 Claude Code

重启后,在对话中发送图片或图片路径,Claude 会自动调用 `understand_image` 工具来分析图片。

## 环境变量

| 变量 | 必需 | 默认值 | 说明 |
|------|------|--------|------|
| `MINIMAX_API_KEY` | ✅ | — | MiniMax API 密钥 |
| `MINIMAX_API_HOST` | ❌ | `https://api.minimax.chat` | API 地址(中国区用 `https://api.minimaxi.com`) |
| `MINIMAX_MODEL` | ❌ | `minimax-vl-01` | VL 模型名(仅 `chat_completion` 模式) |
| `MINIMAX_ENDPOINT` | ❌ | `chat_completion` | API 模式:`chat_completion` 或 `coding_plan` |

### 两种 API 模式

**`chat_completion`(默认)**
通用 MiniMax VL API,通过 `/v1/text/chatcompletion_v2` 调用,支持多图、system prompt、temperature 等参数。适用于标准 token plan。

**`coding_plan`**
MiniMax Coding Plan 专有端点 `/v1/coding_plan/vlm`,仅支持单图 + prompt。如果你是 Coding Plan 用户可用此模式。

> **不确定用哪个?** 先试试 `chat_completion`(默认)。如果返回 404 或 auth 错误,切到 `coding_plan`。

## 使用示例

```python
# 分析本地图片
understand_image(
    prompt="这张图片里有什么?请详细描述。",
    image_path="/Users/wenjiaqi/Downloads/photo.png"
)

# 分析网络图片
understand_image(
    prompt="Extract text from this image",
    image_url="https://example.com/screenshot.jpg"
)

# 多图对比
understand_image(
    prompt="Compare these two UI designs",
    image_paths=["/path/to/design1.png", "/path/to/design2.png"]
)
```

## 项目结构

```
minimax-vision-mcp/
├── pyproject.toml
├── README.md
└── src/
    └── minimax_vision_mcp/
        ├── __init__.py
        └── server.py          # MCP 服务器主文件
```

## License

MIT

TDQS

B3.4/5.0

Scored across 1 tool

Disambiguation5/5

Only one tool exists, so there is no ambiguity between tools. The tool's purpose is clearly defined.

Naming Consistency5/5

The single tool name 'understand_image' follows a clear verb_noun pattern and is descriptive of its function.

Tool Count2/5

With only one tool, the server feels too limited for the implied scope of 'MiniMax Vision MCP'. A vision server would typically offer multiple capabilities.

Completeness2/5

The tool provides only image-to-text analysis. No other operations (e.g., object detection, metadata retrieval) are available, leaving significant gaps for a vision server.

Maintenance

ActivityStale
ResponsivenessNo issues