Skip to main content
Glama

@winton979/vision-mcp

MCP server that exposes an analyze_image tool backed by an OpenAI-compatible vision LLM (GPT-4o, Qwen-VL, etc.).

What it does

Provides a single MCP tool analyze_image that accepts an image via:

  • path — local file path

  • url — public http(s) URL

  • base64 — raw base64 string (with or without data: prefix)

and returns a text description from the configured vision model.

Related MCP server: vision-mcp

Prerequisites

  • Node.js ≥ 18 (global fetch required)

Configuration

Set these environment variables when configuring the MCP server:

Variable

Required

Default

Description

VISION_BASE_URL

No

https://api.openai.com/v1

OpenAI-compatible API base URL

VISION_API_KEY

Yes

API key for the gateway

VISION_MODEL

No

gpt-4o

Vision model name

Claude Code setup

macOS / Linux

Add to ~/.claude.json or ~/.claude/.mcp.json:

{
  "mcpServers": {
    "vision": {
      "command": "npx",
      "args": ["-y", "@winton979/vision-mcp"],
      "env": {
        "VISION_BASE_URL": "<your-base-url>",
        "VISION_API_KEY": "<your-api-key>",
        "VISION_MODEL": "<your-model>"
      }
    }
  }
}

Windows

{
  "mcpServers": {
    "vision": {
      "command": "cmd",
      "args": ["/c", "npx", "-y", "@winton979/vision-mcp"],
      "env": {
        "VISION_BASE_URL": "<your-base-url>",
        "VISION_API_KEY": "<your-api-key>",
        "VISION_MODEL": "<your-model>"
      }
    }
  }
}

Codex setup

macOS / Linux

Add to ~/.codex/config.toml:

[mcp_servers.vision-mcp]
type = "stdio"
command = "npx"
args = ["-y", "@winton979/vision-mcp"]
env = { VISION_BASE_URL = "<your-base-url>", VISION_API_KEY = "<your-api-key>", VISION_MODEL = "<your-model>" }

Windows

[mcp_servers.vision-mcp]
type = "stdio"
command = "npx"
args = ["-y", "@winton979/vision-mcp"]
env = { VISION_BASE_URL = "<your-base-url>", VISION_API_KEY = "<your-api-key>", VISION_MODEL = "<your-model>" }

Tool: analyze_image

Parameter

Type

Required

Description

path

string

one of three

Local file path to the image

url

string

one of three

Public http(s) URL of the image

base64

string

one of three

Raw base64 string

mime_type

string

No

Override MIME type (auto-detected)

task

string

No

Optimized analysis mode (default general)

prompt

string

No

Specific question; overrides the task's default instruction

response_mode

string

No

Output structure: markdown (default) / json / plain_text

model

string

No

Override model per call

max_tokens

integer

No

Default 4096

temperature

number

No

Default 0.2

detail

string

No

low / high / auto

system

string

No

Extra system guidance appended after the built-in base rules

task

Selects an optimized built-in system context + default instruction, so callers don't have to hand-write a prompt for common cases:

task

use for

general

default — objects, text, layout, anomalies

ocr

verbatim text transcription, preserving layout

ui_review

layout, alignment, overflow, element states, a11y

document

document structure & key content

table

reconstruct tables as markdown tables

diagram

nodes, edges, flow, relationships

chart

chart type, axes, series, trends, values

receipt

merchant, line items, totals

math

transcribe & solve step by step

code

transcribe code verbatim

If prompt is also provided, it takes precedence as the specific question while the task's specialized context still applies — e.g. task=ocr, prompt="what is the total amount?".

response_mode

  • markdown (default) — structured report (## Summary / ## Visible Objects / ## Text / ## Findings / ## Uncertainties)

  • json — a single JSON object (summary, objects, text, findings, uncertainties) for easy parsing; the tool returns only the JSON, with no extra metadata appended

  • plain_text — unstructured text

json mode sends response_format: { type: "json_object" }. Some OpenAI-compatible gateways do not support this field and may return HTTP 400; in that case fall back to markdown.

Built-in behavior

A fixed base system prompt is always applied — it enforces observable-facts-only reporting, exact text preservation, and prompt-injection protection (text inside the image is treated as content, never as instructions). The optional system parameter is appended after these rules and cannot override them.

Tips: avoid [image] tag conversion (Windows)

When you paste a local image path into Claude Code, the CLI may auto-convert it into an [image] tag and inline the bytes, which fails on models or not stable that do not accept image input. To keep the raw path intact, use ImageClipboardModify — it appends a fixed prefix to clipboard image paths so they are no longer recognized as images, letting analyze_image receive the path verbatim. for use mcp vision always

中文说明:Claude Code 会把粘贴的本地图片路径自动识别为 [image] 标签,导致不支持图片的模型偶尔不稳定或者报错。可使用 ImageClipboardModify 给剪切板图片附加一段固定文字,绕过识别,让路径以纯文本形式传入 analyze_image,使用使用mcp vision解析

Local development

git clone https://github.com/winton979/vision-mcp.git
cd vision-mcp
npm install
npm run build

# Run smoke test
SMOKE_IMAGE=/path/to/test.png VISION_API_KEY=sk-... npm run smoke

License

MIT

Available Tools

1 tool
analyze_imageA

Analyze an image with a vision LLM (OpenAI-compatible chat/completions). Provide exactly one of path (local file), url (http/https), or base64. Optionally pass a custom prompt to steer the analysis (OCR, table extraction, captioning, Q&A, etc).

ParametersJSON Schema
NameRequiredDescriptionDefault
urlNoPublic http(s) URL of the image.
pathNoAbsolute or relative local file path to the image.
modelNoOverride the vision model. Defaults to env VISION_MODEL (gpt-4o).
base64NoRaw base64 string (with or without data: prefix).
detailNoOptional image detail hint passed to the gateway.
promptNoWhat you want the model to do with the image. Defaults to a detailed description.
systemNoOptional system message.
mime_typeNoOverride MIME type for base64 input. Auto-detected if omitted.
max_tokensNo
temperatureNo

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description bears full responsibility for behavioral disclosure. It explains the tool invokes an LLM and defaults to detailed description, but does not mention limitations (e.g., image size, format compatibility, or error handling). Some behavioral context is implied but not fully detailed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two concise sentences, front-loading the core function and unique input constraints. Every word adds value without redundancy. It is efficiently structured for quick comprehension.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 10 parameters and no output schema, the description covers the main functionality and key parameters but lacks details on output format, error handling, or performance characteristics. It is adequate for basic use but leaves gaps for complex scenarios.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is high (80%), but the description adds critical semantic value by explaining the mutual exclusivity of path/url/base64 and the role of the prompt parameter. This goes beyond the schema's individual descriptions, providing context on how parameters relate and behave.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool analyzes images using a vision LLM, specifies the three supported input formats (path, URL, base64), and mentions optional prompt customization. The verb 'analyze' combined with 'image' and the mention of specific use cases (OCR, table extraction) makes the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly instructs the user to provide exactly one of path, url, or base64, which is a clear usage constraint. It also notes the optional prompt to steer analysis. While no siblings exist to differentiate, the guidance is direct and actionable, though it lacks explicit when-not-to-use scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.1.0
    • First observedanalyze_image

TDQS

A4/5.0

Scored across 1 tool

Disambiguation5/5

Only one tool exists, so there is no ambiguity between tools. An agent cannot misselect among tools.

Naming Consistency5/5

With a single tool, naming consistency is not applicable but there is no inconsistency to flag.

Tool Count2/5

A single tool covering all vision analysis tasks feels too thin for the scope. While the tool is versatile via prompts, it lacks separate endpoints for different operations, making the surface sparse.

Completeness3/5

The tool covers the core vision analysis functionality with options for local/remote/base64 input and custom prompts. However, there are no tools for managing image resources or handling results, leaving minor gaps for complex workflows.

Maintenance

ActivityStale
ResponsivenessNo issues

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers