vision-opencode-mcp
by sddzwxy
README.md
# vision-opencode-mcp
Give opencode (or any MCP client) visual reference by calling an
OpenAI-compatible vision model API. Exposes one tool:
- **`vision_describe(image_path?, question?)`** — read a local image and return
what the vision model sees. When `image_path` is omitted, it automatically
picks the **newest image** (by `LastWriteTime`) from the pasted-images dir,
which is where opencode stores clipboard-pasted images.
Requires only `python>=3.10` and the `mcp` SDK (installed automatically).
## Requirements
- Python 3.10+
- A `VISION_API_KEY` for an OpenAI-compatible vision endpoint
(base URL and model are baked in as defaults; override if needed)
## Install
```bash
# from this repo
python -m venv .venv
# Windows:
.venv\Scripts\pip install -e .
# macOS/Linux:
# .venv/bin/pip install -e .
# or with uv (faster)
uv sync
```
This installs the `vision-mcp` console script and all dependencies
(`mcp`, which pulls in `mcp-types`).
## Configure for opencode
Add an `mcp` entry to your opencode config
(`opencode.json` or `~/.config/opencode/opencode.json`):
```jsonc
{
"mcp": {
"vision": {
"type": "local",
"command": [
"C:\\path\\to\\vision-opencode-mcp\\.venv\\Scripts\\python.exe",
"C:\\path\\to\\vision-opencode-mcp\\server.py"
],
"enabled": true,
"environment": {
"VISION_API_KEY": "your-api-key-here"
}
}
}
}
```
If you installed with `uv sync`, you can also point at the `vision-mcp`
console script instead of the venv python:
```jsonc
{
"mcp": {
"vision": {
"type": "local",
"command": ["C:\\path\\to\\vision-opencode-mcp\\.venv\\Scripts\\vision-mcp.exe"],
"enabled": true,
"environment": {
"VISION_API_KEY": "your-api-key-here"
}
}
}
}
```
Restart opencode. The `vision_describe` tool then appears under the `vision`
MCP server.
## Environment variables
| Variable | Default | Required | Description |
|---|---|---|---|
| `VISION_API_KEY` | *(none)* | **yes** | API key for the vision endpoint. Never hardcoded. |
| `VISION_API_BASE` | `http://opencode.ai/zen/v1` | no | Base URL of the OpenAI-compatible API. |
| `VISION_MODEL` | `mimo-v2.5-free` | no | Vision model name. |
| `VISION_USER_AGENT` | Chrome UA | no | Sent as `User-Agent`; some endpoints (Cloudflare-protected) return 403 without a browser UA. |
| `VISION_TIMEOUT` | `180000` | no | HTTP timeout in ms. |
| `VISION_MAX_TOKENS` | `8192` | no | `max_tokens` for the completion (thinking models need headroom). |
| `VISION_PASTED_DIR` | `%TEMP%\opencode\pasted_images` | no | Directory scanned for the newest image when `image_path` is empty. |
## Usage
With `image_path`:
```
call vision_describe image_path="D:\pics\shot.png" question="这个截图里报了什么错?"
```
Without `image_path` (auto-picks the newest pasted image):
```
call vision_describe question="这张图片里有什么?"
```
## CLI (debug)
The repo also ships `vision.py`, a small command-line version of the same call:
```bash
python vision.py <image_path> [question]
```
Prints the model text to stdout; exits `0` on success, `1` on error
(prints `ERROR: ...`). It reads the same environment variables as the server.
## Notes
- The API key is read from the environment only — it never appears in the
code, so this repo is safe to share.
- For clipboard-pasted images in opencode desktop, the tool auto-picks the
newest file under the pasted-images dir, so you can just say "看这张图"
and the tool picks it up.
TDQS
A4.1/5.0
Scored across 1 tool
Disambiguation5/5
With only one tool, there is no possibility of confusion or overlap. The tool's purpose is clearly described and distinct by virtue of being the only option.
Naming Consistency5/5
The single tool name 'vision_describe' follows a consistent verb_noun pattern. With only one tool, internal naming consistency is trivially maintained.
Tool Count3/5
One tool is borderline thin for a server whose name suggests broader vision capabilities. However, the narrow focus on image description makes the count acceptable, if minimal.
Completeness4/5
The tool covers the core need of interpreting local image files, but lacks support for remote images or additional vision operations (e.g., OCR, image metadata). These are minor gaps for the stated purpose.
Maintenance
ActivitySlowing
ResponsivenessNo issues