Skip to main content
Glama

Vision MCP Server

A Model Context Protocol (MCP) server that gives vision understanding to agents connected to non-multimodal models (DeepSeek, older GPT-4, local small models, etc.): the agent hands an image to the MCP tool, the server calls a vision model, and returns text.

Supports major providers in China and the US plus any OpenAI-compatible endpoint. Official SDKs first, abstraction before implementation, zero-intrusion provider additions.

中文文档见 README.zh-CN.md

Features

  • 4 tools: analyze_image / describe_image / ocr_image / list_providers, all returning plain Markdown text

  • 13 built-in providers: OpenAI / Anthropic / Google Gemini / Qwen (DashScope) / Zhipu / Doubao (Volcengine) / ERNIE (Qianfan) / StepFun / Ollama / Alibaba Bailian / SiliconFlow / OpenRouter / custom OpenAI-compatible endpoint

  • Up to 9 images per call (configurable via VISION_MCP_MAX_IMAGES): local path / http(s) URL / base64 (data URI or raw base64), auto-sniffed, types mixable

  • Three-tier fallback chain: official SDK → OpenAI-compatible endpoint → native fetch (see SPEC §1)

  • Stateless: every call is independent; images and results are never cached; keys are read from environment variables only

Related MCP server: vision-mcp

Quick start

Option A: npx (published to npm, no repo needed)

npx -y @inferai/vision-mcp

Option B: local build

git clone <repo> && cd vision-mcp
pnpm install
pnpm build
node dist/index.js

MCP configuration examples (stdio)

The server speaks the stdio transport: the MCP client spawns the process and exchanges JSON-RPC messages over stdin/stdout. Configure it wherever your client defines MCP servers:

  • Claude Code: project-level .mcp.json or user-level ~/.claude.json (mcpServers key)

  • Claude Desktop: claude_desktop_config.json

  • Any MCP client (Cursor, self-built agents, etc.): same structure

Windows

On Windows npx resolves to npx.cmd, and MCP clients that spawn processes without a shell can't run it directly — wrap it in cmd /c:

{
  "mcpServers": {
    "vision-mcp": {
      "command": "cmd",
      "args": ["/c", "npx", "-y", "@inferai/vision-mcp"],
      "env": {
        "OPENAI_API_KEY": "sk-...",
        "DASHSCOPE_API_KEY": "sk-..."
      }
    }
  }
}

Local development (adjust the path; --env-file-if-exists=.env loads .env natively):

{
  "mcpServers": {
    "vision-mcp": {
      "command": "node",
      "args": [
        "--env-file-if-exists=.env",
        "C:\\path\\to\\vision-mcp\\dist\\index.js"
      ],
      "env": {
        "OPENAI_API_KEY": "sk-..."
      }
    }
  }
}

Linux / macOS

npx runs directly:

{
  "mcpServers": {
    "vision-mcp": {
      "command": "npx",
      "args": ["-y", "@inferai/vision-mcp"],
      "env": {
        "OPENAI_API_KEY": "sk-...",
        "DASHSCOPE_API_KEY": "sk-..."
      }
    }
  }
}

Local development (adjust the path; --env-file-if-exists=.env loads .env natively):

{
  "mcpServers": {
    "vision-mcp": {
      "command": "node",
      "args": [
        "--env-file-if-exists=.env",
        "/absolute/path/to/vision-mcp/dist/index.js"
      ],
      "env": {
        "OPENAI_API_KEY": "sk-..."
      }
    }
  }
}

With startup arguments (override provider defaults via argv, see below; on Windows prefix the command/args with cmd /c):

{
  "mcpServers": {
    "vision-mcp": {
      "command": "npx",
      "args": [
        "-y",
        "@inferai/vision-mcp",
        "--default-provider=dashscope",
        "--siliconflow-api-key=sk-...",
        "--siliconflow-model=Qwen/Qwen2.5-VL-7B-Instruct"
      ],
      "env": {
        "DASHSCOPE_API_KEY": "sk-..."
      }
    }
  }
}

stdio notes:

  • stdout carries the MCP protocol only — the server never prints logs there; diagnostics go to stderr

  • the client manages the process lifecycle (spawn on start, kill on exit); no daemon needed

  • first npx run downloads the package and may take a few seconds

  • env variables can also come from the shell environment if the client inherits it (no env block needed)

Debug with MCP Inspector:

pnpm dlx @modelcontextprotocol/inspector node dist/index.js --xxx-api-key=xxx --xxx2-api-key=xxx

Setting variables

  1. MCP config env block (recommended, most reliable across platforms) — write the variables into the env object above

  2. .env file (local development) — copy .env.example to .env, fill it in, then node --env-file-if-exists=.env dist/index.js (Node 22 native, no dotenv needed)

  3. Shell exportexport OPENAI_API_KEY=sk-xxx then run

Providers without keys show as unavailable in list_providers and report the missing variable when called.

Publishing (before npx works)

pnpm publish          # or pnpm release (changeset flow)

Environment variables

Every provider's API_KEY, BASE_URL, and MODEL support environment overrides (convention: <PROVIDER_PREFIX>_API_KEY / <PROVIDER_PREFIX>_BASE_URL / <PROVIDER_PREFIX>_MODEL):

Provider

Environment variables

Default model

OpenAI

OPENAI_API_KEY, OPENAI_BASE_URL, OPENAI_MODEL

gpt-4o

Anthropic

ANTHROPIC_API_KEY, ANTHROPIC_BASE_URL, ANTHROPIC_MODEL

claude-sonnet-4-5

Google Gemini

GEMINI_API_KEY, GEMINI_BASE_URL, GEMINI_MODEL

gemini-2.5-flash

Alibaba DashScope

DASHSCOPE_API_KEY, DASHSCOPE_BASE_URL, DASHSCOPE_MODEL

qwen-vl-max

Zhipu

ZHIPU_API_KEY, ZHIPU_BASE_URL, ZHIPU_MODEL

glm-4v-flash (free)

Volcengine Doubao

VOLCENGINE_ARK_API_KEY, VOLCENGINE_ARK_BASE_URL, VOLCENGINE_ARK_MODEL

doubao-1.5-vision-pro

Baidu Qianfan

QIANFAN_API_KEY, QIANFAN_SECRET_KEY, QIANFAN_BASE_URL, QIANFAN_MODEL

ernie-4.5-vl-8k

StepFun

STEPFUN_API_KEY, STEPFUN_BASE_URL, STEPFUN_MODEL

step-1v

Ollama (local)

OLLAMA_BASE_URL, OLLAMA_MODEL

— (no built-in default; endpoint and model must be set)

Alibaba Bailian

BAILIAN_API_KEY, BAILIAN_BASE_URL (default DashScope compatible mode), BAILIAN_MODEL

qwen-vl-max

SiliconFlow

SILICONFLOW_API_KEY, SILICONFLOW_BASE_URL (default https://api.siliconflow.cn/v1), SILICONFLOW_MODEL

Qwen/Qwen2.5-VL-72B-Instruct

OpenRouter

OPENROUTER_API_KEY, OPENROUTER_BASE_URL (default https://openrouter.ai/api/v1), OPENROUTER_MODEL

openai/gpt-4o

Custom compatible

OPENAI_COMPAT_BASE_URL, OPENAI_COMPAT_API_KEY?, OPENAI_COMPAT_MODEL

? = optional (has a built-in default); * = required.

Global configuration:

Environment variable

Default

Description

VISION_MCP_DEFAULT_PROVIDER

first available

Default provider

VISION_MCP_DEFAULT_MODEL

provider default

Default model

VISION_MCP_PROVIDER_PRIORITY

table order

Provider priority (comma-separated, high first, e.g. openai,dashscope,zhipu)

VISION_MCP_MAX_RETRIES

0 (off)

Per-provider retry count before falling back

VISION_MCP_MAX_FALLBACKS

0 (off)

Max provider fallbacks before giving up

VISION_MCP_MAX_IMAGE_BYTES

20 MB

Image size limit

VISION_MCP_MAX_IMAGES

9

Max images per tool call

VISION_MCP_TIMEOUT_MS

60000

Download & request timeout (ms)

Fallback chain

When multiple providers are available, calls walk the priority chain: configured default → VISION_MCP_PROVIDER_PRIORITY list → table order (unavailable providers are skipped).

  • each provider is retried up to VISION_MCP_MAX_RETRIES times on provider errors (upstream failures, timeouts)

  • after a provider exhausts its retries, the next available provider in the chain is tried, up to VISION_MCP_MAX_FALLBACKS fallbacks

  • only provider errors trigger retry/fallback; config or image errors fail fast

  • an explicitly requested provider argument is tried alone (no fallback)

  • when everything fails, the error lists every provider attempted and its last error

Also available as argv: --provider-priority=..., --max-retries=N, --max-fallbacks=N (beat env vars).

MCP startup arguments (argv)

Every provider's apiKey / baseUrl / model can be overridden via startup arguments (higher priority than environment variables), format --<provider>-<field>:

node dist/index.js \
  --openai-api-key=sk-xxx \
  --openai-base-url=https://my-gateway.example.com/v1 \
  --openai-model=gpt-4o-mini \
  --dashscope-api-key=sk-xxx \
  --default-provider=dashscope
  • Global: --default-provider <name> / --default-model <name>

  • Per provider: --<provider>-api-key, --<provider>-base-url, --<provider>-model (equals or space form both work)

  • Any OpenAI-compatible third-party service: wire it up in one line with --openai-compat-base-url + --openai-compat-api-key + --openai-compat-model; or point any built-in provider's base-url at a mirror/proxy

Priority: tool args provider/model > startup args (per-provider > global default) > environment variables > provider built-in defaults.

Tools

Tool

Arguments

Description

analyze_image

images*, prompt?, provider?, model?

General image analysis

describe_image

images*, provider?, model?

Describe image content (default instruction)

ocr_image

images*, language? (auto/zh/en/zh-en), provider?, model?

OCR, preserving layout

list_providers

Provider list and configuration status

images accepts a single image or an array (up to VISION_MCP_MAX_IMAGES, default 9): local path / http(s):// URL / data: URI / raw base64, auto-sniffed. Multiple images are seen by the model in the given order (compare, diff, or combine them).

Security note: URL downloads are SSRF-protected — every hop (including redirects) is validated and URLs resolving to loopback, private, or link-local addresses are blocked (hint in the error explains why).

Provider integration (three-tier fallback chain)

provider

Integration

Notes

openai / stepfun / ollama / bailian / siliconflow / openrouter / openai-compat

OpenAI-compatible adapter (openai SDK)

One adapter, configurable baseURL

anthropic

Official SDK @anthropic-ai/sdk

messages + image content block

gemini

Official SDK @google/generative-ai

generateContent + inlineData

dashscope

Native fetch

official npm package has no vision; direct multimodal-generation API

zhipu

Native fetch

official SDK accepts string content only; direct v4 API

volcengine

Native fetch

official openapi is a management plane; direct Ark API

qianfan

Native fetch

official SDK is string-only; AK/SK → token → v2 API

Adding a provider: for OpenAI-compatible endpoints, add one row to RULES in src/core/config.ts plus one mapping in the factory table in src/index.ts — zero new code. Official SDK or native fetch implementations: see SPEC §1.

Development

pnpm check        # biome checks
pnpm test         # rstest unit tests (injected mocks, no network)
pnpm build        # rslib build

Real-call smoke tests (only run against providers whose keys are configured; skipped otherwise):

OPENAI_API_KEY=sk-... pnpm exec rstest tests/e2e

Architecture

src/
├── index.ts            # Entry: composition root, stdio startup
├── core/               # Abstraction: interfaces / image loading / config / registry
├── providers/          # Adapters: official SDK or compatible endpoints, protocol conversion only
└── server/tools.ts     # MCP tool layer: zod validation + error mapping

Full spec: SPEC.md.

Available Tools

4 tools
analyze_imageA

Analyze one or more images with a vision model and return the text result. Accepts 1 to 9 images (local path, URL, data URI, or base64; types can be mixed). Pass several images to compare, diff, or combine them — the model sees them in the given order. Use for reading screenshots, photos, charts, UI states, document pages, etc.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoModel name, overrides the provider default model
imagesYesOne image, or an array of up to 9 images. Each entry: local path / http(s):// URL / data: URI / raw base64 string; types can be mixed. Pass multiple images to compare, diff, or combine them (e.g. before/after pairs, several charts) — order matters.
promptNoAnalysis instruction; defaults to a detailed description of the image(s)
providerNoProvider name (e.g. openai / dashscope / zhipu / ollama); defaults to the configured default

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It does a good job: it mentions the ability to handle 1-9 images, supports multiple formats (local path, URL, data URI, base64), and clarifies that order is preserved and images can be compared/diffed. This goes beyond what one might assume and helps the agent anticipate multi-image behavior. However, it does not mention any potential failure modes, rate limits, or other operational details that could be relevant, though these may not be necessary for this tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is exceptionally concise—two sentences that efficiently convey the core purpose, input capabilities, and typical use cases. It is front-loaded with the primary function and does not contain redundant or explanatory fluff. Every clause adds value, and the structure flows naturally from what it does to how it behaves with multiple inputs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with four parameters, no nested objects, and no output schema, the description covers the essential aspects: input types, count limits, ordering, and example use cases. It gives the agent enough to invoke the tool correctly for common scenarios. It does not explain when to prefer this over describe_image or ocr_image, which could be considered a completeness gap given those siblings exist, but the description still stands on its own as a complete usage guide for the tool itself.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100%, and the schema itself already thoroughly explains the parameters, including the images format, the order-preservation nuance, and the default prompt behavior. The tool description adds little beyond what the schema states. For instance, the description's note about 'compare, diff, or combine' is a rephrasing of the schema's 'compare, diff, or combine' text. Since the schema covers this ground, the description meets the baseline but does not add substantial new meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Analyze one or more images with a vision model and return the text result.' It specifies a verb (analyze) and a resource (images), and indicates the output (text). However, it does not differentiate from the siblings describe_image and ocr_image, which likely serve overlapping purposes. For example, it does not explain how analyze_image is distinct from describe_image beyond the multi-image capability.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use the tool, listing use cases such as 'reading screenshots, photos, charts, UI states, document pages, etc.' This gives an agent a good sense of appropriate scenarios. However, it does not explicitly mention alternatives or exclusions, such as 'use describe_image for single images' or 'use ocr_image for text extraction only,' so the guidance is strong but not exhaustive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

describe_imageA

Describe the image content: subject, layout, text, details. Accepts 1 to 9 images; each image is described in the given order. Equivalent to analyze_image with the default description instruction.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoModel name, overrides the provider default model
imagesYesOne image, or an array of up to 9 images. Each entry: local path / http(s):// URL / data: URI / raw base64 string; types can be mixed. Pass multiple images to compare, diff, or combine them (e.g. before/after pairs, several charts) — order matters.
providerNoProvider name (e.g. openai / dashscope / zhipu / ollama); defaults to the configured default

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description provides limited behavioral info. It states that it accepts 1–9 images and preserves order, which is useful, but does not mention side effects, authentication, rate limits, or error handling. Since it's a descriptive read operation, the risk is low, but transparency is incomplete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences convey the core function, parameter limits, and ordering behavior without fluff. The information is front-loaded and efficiently structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, but the description implies text output via 'Describe the image content'. It covers the input constraints and purpose adequately. While it doesn't specify return format, that is not critical for a description task, and the context is sufficiently complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds value by explaining the images parameter: 'Pass multiple images to compare, diff, or combine them (e.g. before/after pairs, several charts) — order matters.' This enriches understanding beyond the schema's generic descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Describe') and resource ('image content') with explicit scope (subject, layout, text, details). It also distinguishes from siblings by noting it is 'Equivalent to analyze_image with the default description instruction', which clarifies its relationship to a similar tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by mentioning equivalence to analyze_image, suggesting that custom instructions would require analyze_image. However, it does not explicitly state when to use this tool over ocr_image or list_providers, leaving some ambiguity. Still, the core use case is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_providersA

List all registered vision model providers, their default models, and configuration status. Providers without keys will error when called.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description carries full burden. It mentions that providers without keys will error when called, which is a useful behavioral warning. However, it doesn't disclose potential side effects (though listing is read-only, it could be mistaken), nor does it explain implications of configuration status, but the warning is a positive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, extremely concise, and front-loaded with the core action (list all providers) and output details. Every sentence provides value: the second clarifies error behavior. No redundancy or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (no parameters, no output schema, no annotations), the description covers key aspects: what it lists, the error condition. It lacks explicit mention of the exact return format or whether it includes provider IDs, but for a list tool, this is largely sufficient. A minor gap is not specifying if it returns configuration status as structured data or human-readable text.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so there are no parameter semantics to explain. The description focuses on the output content (default models, configuration status), which is sufficient. With no parameters, a baseline of 4 is appropriate as the description adds meaning about the result.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool lists all registered vision model providers, including their default models and configuration status. It distinguishes itself from sibling tools (analyze/describe/ocr) which are image-related operations, though it doesn't explicitly compare.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description doesn't explicitly state when to use this tool vs alternatives, but the purpose is clear enough that an agent would know it's for listing providers, not image analysis. It conveys usage context implicitly by listing what it returns, but no explicit exclusions or alternative tool references.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ocr_imageA

OCR: transcribe all text in the image(s), preserving reading order and paragraph structure. Accepts 1 to 9 images; transcripts follow the given order. Suitable for screenshots, scans, invoices, slides, etc.

ParametersJSON Schema
NameRequiredDescriptionDefault
modelNoModel name, overrides the provider default model
imagesYesOne image, or an array of up to 9 images. Each entry: local path / http(s):// URL / data: URI / raw base64 string; types can be mixed. Pass multiple images to compare, diff, or combine them (e.g. before/after pairs, several charts) — order matters.
languageNoRecognition language; defaults to auto
providerNoProvider name (e.g. openai / dashscope / zhipu / ollama); defaults to the configured default

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and does disclose meaningful behavior: preserving reading order and paragraph structure, accepting 1–9 images, and returning transcripts in the given order. It does not mention output format or error behavior, but the core read-only transcription behavior is well covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences with the core action front-loaded. Every sentence earns its place, and there is no redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers purpose, input limits, ordering, and typical use cases, which is sufficient for an agent to select and invoke the tool. Since there is no output schema, a brief note on return format would improve completeness, but this is a minor gap for a simple OCR tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description reinforces image order and count, which the schema already documents, but adds no parameter-level meaning for model, language, or provider beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('transcribe all text') and resource ('image(s)'), and adds output characteristics (reading order, paragraph structure). This clearly distinguishes it from sibling tools like analyze_image and describe_image.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives concrete use cases (screenshots, scans, invoices, slides) but does not explicitly state when to choose this tool over analyze_image or describe_image, nor when not to use it. The differentiation is implied by the OCR action rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv1.1.1
    • Changedanalyze_image4 fields changed
      • removedInput schema / properties / image
        Removed value: -{
        -  "description": "Image: local path / http(s):// URL / data: URI / raw base64 string",
        -  "minLength": 1,
        -  "type": "string"
        -}
      • addedInput schema / properties / images
        Added value: +{
        +  "anyOf": [
        +    {
        +      "description": "local path / http(s):// URL / data: URI / raw base64 string",
        +      "minLength": 1,
        +      "type": "string"
        +    },
        +    {
        +      "description": "Array of 9 images max; order is preserved when the model sees them",
        +      "items": {
        +        "description": "local path / http(s):// URL / data: URI / raw base64 string",
        +        "minLength": 1,
        +        "type": "string"
        +      },
        +      "maxItems": 9,
        +      "minItems": 1,
        +      "type": "array"
        +    }
        +  ],
        +  "description": "One image, or an array of up to 9 images. Each entry: local path / http(s):// URL / data: URI / raw base64 string; types can be mixed. Pass multiple images to compare, diff, or combine them (e.g. before/after pairs, several charts) — order matters."
        +}
      • changedInput schema / properties / prompt / description
        Previous value: -"Analysis instruction; defaults to a detailed description"New value: +"Analysis instruction; defaults to a detailed description of the image(s)"
      • changedInput schema / required
        Previous value: -[
        -  "image"
        -]New value: +[
        +  "images"
        +]
    • Changeddescribe_image3 fields changed
      • removedInput schema / properties / image
        Removed value: -{
        -  "description": "Image: local path / http(s):// URL / data: URI / raw base64 string",
        -  "minLength": 1,
        -  "type": "string"
        -}
      • addedInput schema / properties / images
        Added value: +{
        +  "anyOf": [
        +    {
        +      "description": "local path / http(s):// URL / data: URI / raw base64 string",
        +      "minLength": 1,
        +      "type": "string"
        +    },
        +    {
        +      "description": "Array of 9 images max; order is preserved when the model sees them",
        +      "items": {
        +        "description": "local path / http(s):// URL / data: URI / raw base64 string",
        +        "minLength": 1,
        +        "type": "string"
        +      },
        +      "maxItems": 9,
        +      "minItems": 1,
        +      "type": "array"
        +    }
        +  ],
        +  "description": "One image, or an array of up to 9 images. Each entry: local path / http(s):// URL / data: URI / raw base64 string; types can be mixed. Pass multiple images to compare, diff, or combine them (e.g. before/after pairs, several charts) — order matters."
        +}
      • changedInput schema / required
        Previous value: -[
        -  "image"
        -]New value: +[
        +  "images"
        +]
    • Changedocr_image3 fields changed
      • removedInput schema / properties / image
        Removed value: -{
        -  "description": "Image: local path / http(s):// URL / data: URI / raw base64 string",
        -  "minLength": 1,
        -  "type": "string"
        -}
      • addedInput schema / properties / images
        Added value: +{
        +  "anyOf": [
        +    {
        +      "description": "local path / http(s):// URL / data: URI / raw base64 string",
        +      "minLength": 1,
        +      "type": "string"
        +    },
        +    {
        +      "description": "Array of 9 images max; order is preserved when the model sees them",
        +      "items": {
        +        "description": "local path / http(s):// URL / data: URI / raw base64 string",
        +        "minLength": 1,
        +        "type": "string"
        +      },
        +      "maxItems": 9,
        +      "minItems": 1,
        +      "type": "array"
        +    }
        +  ],
        +  "description": "One image, or an array of up to 9 images. Each entry: local path / http(s):// URL / data: URI / raw base64 string; types can be mixed. Pass multiple images to compare, diff, or combine them (e.g. before/after pairs, several charts) — order matters."
        +}
      • changedInput schema / required
        Previous value: -[
        -  "image"
        -]New value: +[
        +  "images"
        +]
  2. 4 tool updatesv1.0.0
    • First observedanalyze_image
    • First observeddescribe_image
    • First observedlist_providers
    • First observedocr_image

TDQS

A3.7/5.0

Scored across 4 tools

Disambiguation2/5

analyze_image, describe_image, and ocr_image all accept the same input and return text results. describe_image is explicitly described as equivalent to analyze_image with a default instruction, and ocr_image is just a specialized prompt variant. This creates significant overlap and makes it unclear when to choose one over another. Only list_providers is clearly distinct.

Naming Consistency4/5

All tool names use snake_case with a verb-first pattern: analyze_image, list_providers, describe_image, ocr_image. The only minor deviation is ocr_image using an acronym instead of a plain verb, but it still fits the pattern. Overall, naming is predictable and consistent.

Tool Count3/5

With only 4 tools, the server is on the low end of the typical range. However, 3 of the 4 tools essentially perform the same task with different prompt variations, so the effective functionality is even more limited. The count feels padded rather than well-scoped, and could be reduced to just analyze_image and list_providers without loss.

Completeness4/5

The server covers the core functionality of image analysis, including general analysis, description, and OCR, plus provider list management. Since analyze_image is generic and accepts multiple images for comparison, it covers most basic vision tasks. Minor gaps include lack of explicit model management or configuration tools, but list_providers partially addresses this.

Maintenance

ActivityMaintained
ResponsivenessSyncing

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    MCP server for analyzing images using multiple vision LLM providers (OpenCode, OpenAI, Anthropic, Google, and custom OpenAI-compatible endpoints). Provides tools to analyze single or multiple images, list providers, and test vision capabilities.
    MIT