GLM Vision MCP
# GLM Vision MCP
<p align="center">
<strong>Give any text-only reasoning model a pair of eyes.</strong>
</p>
<p align="center">
An MCP (Model Context Protocol) server that wraps vision-language models — like <strong>GLM-4.6V-Flash</strong> — as callable tools, so text-only LLMs (DeepSeek-R1, Qwen3-Thinking, etc.) can <em>see</em> images inside Cursor, Trae CN, Cline, Claude Desktop, and any MCP-compatible agent.
</p>
---
## How It Works
```
┌─────────────────────────────────────────────────────────┐
│ Your IDE / Agent │
│ │
│ Text-only model (e.g. DeepSeek-R1) │
│ │ │
│ │ "What's in this screenshot?" │
│ ▼ │
│ MCP Tool: see_image(path, question) │
│ │ │
│ ▼ (stdio / MCP protocol) │
│ ┌──────────────────────────────────────────────────┐ │
│ │ GLM Vision MCP Server │ │
│ │ │ │
│ │ Reads image → base64 → sends to vision model │ │
│ │ (GLM-4.6V-Flash / GPT-4o / any VLM) │ │
│ │ │ │
│ │ Returns: "A login error dialog showing..." │ │
│ └──────────────────────────────────────────────────┘ │
│ │ │
│ ▼ │
│ Text-only model continues reasoning with vision data │
└─────────────────────────────────────────────────────────┘
```
You configure **3 things**: model provider, model ID, and API key. The server handles everything else — image reading, base64 encoding, API calls, error handling.
---
## Quick Start
### 1. Install
```bash
# From PyPI (once published) or from source:
pip install -e .
```
Or use directly with `uv` / `pipx` without installing:
```bash
# Using uv (recommended for MCP)
uv run glm-vision-mcp
```
### 2. Configure
The server reads configuration from **environment variables**:
| Variable | Required | Default | Description |
|---|---|---|---|
| `VISION_API_KEY` | ✅ Yes | — | Your model provider's API key |
| `VISION_MODEL_ID` | ✅ Yes | `glm-4.6v-flash` | The vision model to use |
| `VISION_MODEL_PROVIDER` | ✅ Yes | `zhipu` | Provider name (see table below) |
| `VISION_BASE_URL` | ❌ No | *provider default* | Custom API base URL |
| `VISION_MAX_TOKENS` | ❌ No | `2048` | Max response tokens |
| `VISION_TEMPERATURE` | ❌ No | `0.4` | Sampling temperature |
### 3. Add to Your IDE
#### Cursor
Add to `~/.cursor/mcp.json` (or `.cursor/mcp.json` in your project):
```json
{
"mcpServers": {
"glm-vision": {
"command": "python",
"args": ["-m", "glm_vision_mcp"],
"env": {
"VISION_MODEL_PROVIDER": "zhipu",
"VISION_MODEL_ID": "glm-4.6v-flash",
"VISION_API_KEY": "your-zhipu-api-key-here"
}
}
}
}
```
#### Trae CN
Add to Trae's MCP settings (`设置 → MCP`):
```json
{
"mcpServers": {
"glm-vision": {
"command": "python",
"args": ["-m", "glm_vision_mcp"],
"env": {
"VISION_MODEL_PROVIDER": "zhipu",
"VISION_MODEL_ID": "glm-4.6v-flash",
"VISION_API_KEY": "your-zhipu-api-key-here"
}
}
}
}
```
#### Cline (VS Code)
Add to `~/.cline/mcp_settings.json`:
```json
{
"mcpServers": {
"glm-vision": {
"command": "python",
"args": ["-m", "glm_vision_mcp"],
"env": {
"VISION_MODEL_PROVIDER": "zhipu",
"VISION_MODEL_ID": "glm-4.6v-flash",
"VISION_API_KEY": "your-zhipu-api-key-here"
},
"disabled": false,
"autoApprove": []
}
}
}
```
#### Claude Desktop
Add to `claude_desktop_config.json`:
```json
{
"mcpServers": {
"glm-vision": {
"command": "python",
"args": ["-m", "glm_vision_mcp"],
"env": {
"VISION_MODEL_PROVIDER": "zhipu",
"VISION_MODEL_ID": "glm-4.6v-flash",
"VISION_API_KEY": "your-zhipu-api-key-here"
}
}
}
}
```
#### Generic MCP Client (any MCP-compatible tool)
```json
{
"mcpServers": {
"glm-vision": {
"command": "python",
"args": ["-m", "glm_vision_mcp"],
"env": {
"VISION_MODEL_PROVIDER": "zhipu",
"VISION_MODEL_ID": "glm-4.6v-flash",
"VISION_API_KEY": "your-api-key"
}
}
}
}
```
> **Tip:** If you installed via `uv`, use `"command": "uv"` and `"args": ["run", "glm-vision-mcp"]` instead.
---
## Supported Providers
| Provider | `VISION_MODEL_PROVIDER` | Default Base URL | Example Models |
|---|---|---|---|
| **Zhipu (智谱)** | `zhipu` | `https://open.bigmodel.cn/api/paas/v4` | `glm-4.6v-flash`, `glm-4v-plus` |
| **OpenAI** | `openai` | `https://api.openai.com/v1` | `gpt-4o`, `gpt-4o-mini` |
| **DeepSeek** | `deepseek` | `https://api.deepseek.com/v1` | `deepseek-vl` |
| **Moonshot (Kimi)** | `moonshot` | `https://api.moonshot.cn/v1` | `moonshot-v1-8k-vision` |
| **SiliconFlow** | `siliconflow` | `https://api.siliconflow.cn/v1` | `Qwen/Qwen2-VL-72B` |
| **Custom** | `custom` | *(you set `VISION_BASE_URL`)* | Any OpenAI-compatible VLM |
### Using a Custom Endpoint
Set `VISION_BASE_URL` to point to your own server (vLLM, Ollama, LM Studio, etc.):
```json
{
"env": {
"VISION_MODEL_PROVIDER": "custom",
"VISION_MODEL_ID": "your-model-name",
"VISION_API_KEY": "any-or-empty",
"VISION_BASE_URL": "http://localhost:8000/v1"
}
}
```
---
## Tools
The server exposes 4 tools. Your IDE's agent will automatically call them when it needs vision:
### `see_image` — Core Vision Tool
Ask any question about an image.
```
see_image(image, question="What is in this image?")
```
- **image**: File path, URL, or base64 string
- **question**: What you want to know (default: "What is in this image?")
### `describe_image` — Image Description
Generate a text description/caption.
```
describe_image(image, detail_level="detailed")
```
- **detail_level**: `"brief"` | `"detailed"` | `"exhaustive"` (default: `"detailed"`)
### `extract_text` — OCR
Extract all visible text from an image.
```
extract_text(image, language_hint="Chinese")
```
- **language_hint**: Optional — e.g. `"Chinese"`, `"English"`, `"mixed"`
### `analyze_chart` — Chart & Diagram Analysis
Analyze charts, graphs, architecture diagrams, or UI screenshots.
```
analyze_chart(image, question="")
```
- **question**: Optional specific question (default: general analysis)
---
## Usage Example
Once configured, just talk to your IDE's agent normally:
> **You:** "Look at the screenshot at `/tmp/error.png` — what's wrong?"
The agent will:
1. Call `see_image("/tmp/error.png", "What error is shown?")`
2. The MCP server sends the image to GLM-4.6V-Flash
3. GLM returns: *"The dialog shows a 'Connection Refused' error..."*
4. Your text-only model uses that answer to help you
---
## Getting an API Key
### Zhipu (智谱) — Free Tier Available
1. Visit [https://open.bigmodel.cn](https://open.bigmodel.cn)
2. Sign up / log in
3. Go to API Keys → Create new key
4. Copy the key (format: `xxxxxxxx.xxxxxxxx`)
> GLM-4.6V-Flash offers free quota — great for testing!
### OpenAI
1. Visit [https://platform.openai.com/api-keys](https://platform.openai.com/api-keys)
2. Create a new key
---
## Development
### Project Structure
```
glm-vision-mcp/
├── src/glm_vision_mcp/
│ ├── __init__.py
│ ├── __main__.py # python -m glm_vision_mcp
│ ├── server.py # MCP server + tool definitions
│ ├── config.py # Env-var config loader
│ ├── utils.py # Image encoding utilities
│ └── providers/
│ ├── base.py # VisionProvider (shared HTTP logic)
│ ├── zhipu.py # Zhipu GLM provider
│ ├── openai_compat.py # Generic OpenAI-compatible provider
│ └── registry.py # Provider name → class mapping
├── examples/
│ └── quickstart.py
├── pyproject.toml
├── requirements.txt
└── README.md
```
### Running Tests
```bash
pip install -e ".[dev]"
pytest
```
### Adding a New Provider
1. Create `src/glm_vision_mcp/providers/my_provider.py`:
```python
from glm_vision_mcp.providers.base import VisionProvider
class MyProvider(VisionProvider):
def _build_headers(self):
# Custom auth if needed
return {"X-Api-Key": self.config.api_key}
```
2. Register in `providers/registry.py`:
```python
_PROVIDERS["my_provider"] = MyProvider
```
---
## FAQ
**Q: Can I use this with a non-vision model?**
No — the configured model must support vision input (images). If you're unsure, GLM-4.6V-Flash is a good free option.
**Q: Does it work with local images?**
Yes. Pass a file path and the server will read and base64-encode it automatically.
**Q: How fast is it?**
Depends on the model provider. GLM-4.6V-Flash is very fast (typically 1-3 seconds per image).
**Q: Can multiple images be analyzed at once?**
Currently each tool call handles one image. For multi-image comparison, call `see_image` multiple times or extend the tools.
---
## License
[MIT](LICENSE) © 2026 xiayuyang750
TDQS
Scored across 4 tools
Most tools have distinct purposes: see_image for Q&A, extract_text for OCR, analyze_chart for charts/diagrams, and describe_image for general descriptions. However, see_image and describe_image overlap somewhat, as see_image with a default question essentially generates a description. The descriptions help clarify intent, so the ambiguity is limited.
All tool names follow a consistent verb_noun pattern in lowercase with underscores: see_image, extract_text, analyze_chart, describe_image. This makes the API predictable and easy to navigate.
With 4 tools, the server is well-scoped for a vision MCP. Each tool covers a distinct core capability (Q&A, OCR, chart analysis, description) without unnecessary bloat. The count is appropriate for the domain.
The tool surface covers the major vision use cases: answering questions, extracting text, analyzing charts, and generating descriptions. Minor gaps exist, such as no explicit image comparison or object detection, but these are not core for a generic vision server. Overall, the coverage is solid with no dead ends.