Skip to main content
Glama
sddzwxy
by sddzwxy

vision-opencode-mcp

Give opencode (or any MCP client) visual reference by calling an OpenAI-compatible vision model API. Exposes one tool:

  • vision_describe(image_path?, question?) — read a local image and return what the vision model sees. When image_path is omitted, it automatically picks the newest image (by LastWriteTime) from the pasted-images dir, which is where opencode stores clipboard-pasted images.

Requires only python>=3.10 and the mcp SDK (installed automatically).

Requirements

  • Python 3.10+

  • A VISION_API_KEY for an OpenAI-compatible vision endpoint (base URL and model are baked in as defaults; override if needed)

Related MCP server: vision-mcp

Install

# from this repo
python -m venv .venv
# Windows:
.venv\Scripts\pip install -e .
# macOS/Linux:
# .venv/bin/pip install -e .

# or with uv (faster)
uv sync

This installs the vision-mcp console script and all dependencies (mcp, which pulls in mcp-types).

Configure for opencode

Add an mcp entry to your opencode config (opencode.json or ~/.config/opencode/opencode.json):

{
  "mcp": {
    "vision": {
      "type": "local",
      "command": [
        "C:\\path\\to\\vision-opencode-mcp\\.venv\\Scripts\\python.exe",
        "C:\\path\\to\\vision-opencode-mcp\\server.py"
      ],
      "enabled": true,
      "environment": {
        "VISION_API_KEY": "your-api-key-here"
      }
    }
  }
}

If you installed with uv sync, you can also point at the vision-mcp console script instead of the venv python:

{
  "mcp": {
    "vision": {
      "type": "local",
      "command": ["C:\\path\\to\\vision-opencode-mcp\\.venv\\Scripts\\vision-mcp.exe"],
      "enabled": true,
      "environment": {
        "VISION_API_KEY": "your-api-key-here"
      }
    }
  }
}

Restart opencode. The vision_describe tool then appears under the vision MCP server.

Environment variables

Variable

Default

Required

Description

VISION_API_KEY

(none)

yes

API key for the vision endpoint. Never hardcoded.

VISION_API_BASE

http://opencode.ai/zen/v1

no

Base URL of the OpenAI-compatible API.

VISION_MODEL

mimo-v2.5-free

no

Vision model name.

VISION_USER_AGENT

Chrome UA

no

Sent as User-Agent; some endpoints (Cloudflare-protected) return 403 without a browser UA.

VISION_TIMEOUT

180000

no

HTTP timeout in ms.

VISION_MAX_TOKENS

8192

no

max_tokens for the completion (thinking models need headroom).

VISION_PASTED_DIR

%TEMP%\opencode\pasted_images

no

Directory scanned for the newest image when image_path is empty.

Usage

With image_path:

call vision_describe  image_path="D:\pics\shot.png"  question="这个截图里报了什么错?"

Without image_path (auto-picks the newest pasted image):

call vision_describe  question="这张图片里有什么?"

CLI (debug)

The repo also ships vision.py, a small command-line version of the same call:

python vision.py <image_path> [question]

Prints the model text to stdout; exits 0 on success, 1 on error (prints ERROR: ...). It reads the same environment variables as the server.

Notes

  • The API key is read from the environment only — it never appears in the code, so this repo is safe to share.

  • For clipboard-pasted images in opencode desktop, the tool auto-picks the newest file under the pasted-images dir, so you can just say "看这张图" and the tool picks it up.

Available Tools

1 tool
vision_describeA

Give the model visual reference: read a local image file and return what a vision-capable model sees/understands about it. Use whenever you need to see or interpret an image (screenshots, diagrams, photos).

ParametersJSON Schema
NameRequiredDescriptionDefault
questionNoOptional free-form question about the image. When empty, the model produces a general detailed description.
image_pathNoAbsolute path to a local image file (png/jpg/jpeg/webp/bmp/gif). When omitted/empty, automatically picks the newest image by LastWriteTime under the pasted-images dir (VISION_PASTED_DIR).

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the core behavior (reading a local file and returning a vision-based interpretation) and implies a non-destructive read operation. However, it doesn't mention potential limitations such as unsupported formats (covered in schema), file size constraints, or behavior when the file is missing, leaving some gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no superfluous content. The first sentence identifies the action and outcome, the second gives a concise usage directive. It is well-structured and front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only tool with two optional parameters and no output schema, the description adequately covers purpose and usage. The schema handles parameter details. The output is hinted as 'what a vision-capable model sees/understands' but not fully specified, which is a minor gap. Overall, it is complete enough for an agent to invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema provides 100% coverage with clear descriptions for both 'question' and 'image_path', including default behavior for the latter. The tool description does not add additional meaning beyond the schema, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: 'read a local image file and return what a vision-capable model sees/understands about it.' It uses a specific verb ('read') and resource ('local image file'), making the purpose unambiguous. Since there are no sibling tools, distinguishing is not needed, but the description stands on its own.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'Use whenever you need to see or interpret an image' and provides examples like screenshots, diagrams, and photos. This gives clear context for when to use the tool, though it doesn't mention alternatives because none exist. It's a straightforward usage instruction without exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

TDQS

A4.1/5.0
Disambiguation5/5

With only one tool, there is no possibility of confusion or overlap. The tool's purpose is clearly described and distinct by virtue of being the only option.

Naming Consistency5/5

The single tool name 'vision_describe' follows a consistent verb_noun pattern. With only one tool, internal naming consistency is trivially maintained.

Tool Count3/5

One tool is borderline thin for a server whose name suggests broader vision capabilities. However, the narrow focus on image description makes the count acceptable, if minimal.

Completeness4/5

The tool covers the core need of interpreting local image files, but lacks support for remote images or additional vision operations (e.g., OCR, image metadata). These are minor gaps for the stated purpose.

Maintenance

ActivityMaintained
ResponsivenessSyncing

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/sddzwxy/vision-opencode-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server