Skip to main content
Glama

analyze-image-mcp

Минимальный MCP-сервер (пример подключения ниже — для OpenCode, но подойдёт любому MCP-клиенту), который даёт text-only модели доступ к отдельной vision-модели через один инструмент — analyze_image. Работает с любым OpenAI-compatible vision API (включая провайдеров с ограниченной поддержкой tool calling), так как обращается к нему напрямую, минуя механизм tools/tool_choice самого OpenCode.

Зачем это нужно

Если у тебя:

  • основная модель — только текстовая;

  • вторая модель — с поддержкой изображений, но у провайдера не включён --enable-auto-tool-choice (частая ошибка "auto" tool choice requires --enable-auto-tool-choice and --tool-call-parser to be set);

то стандартная отправка картинки напрямую в OpenCode ломается. Этот MCP-сервер решает проблему: он вызывает vision-модель отдельным чистым HTTP-запросом без полей tools/tool_choice, а результат (текстовое описание) отдаёт основной модели как обычный текст.

Related MCP server: VisionToolMCP

Как это работает

Пользователь прикладывает изображение
        │
        ▼
Основная (text-only) модель видит инструкцию:
"если есть изображение — вызови analyze_image"
        │
        ▼
MCP-сервер (analyze_image):
  1. читает файл / URL / data-URI
  2. кодирует в data:image/...;base64,...
  3. отправляет чистый OpenAI-compatible запрос
     (без tools, без tool_choice) в vision-провайдера
        │
        ▼
Текстовое описание возвращается основной модели
        │
        ▼
Основная модель отвечает пользователю

Установка

Нужен Node.js 18+ (используется встроенный fetch) и git. Пакет в npm не публикуется — ставится из GitHub. Выбери один из вариантов.

Вариант 1. Клонирование (рекомендуется)

git clone https://github.com/iljyxa/analyze-image-mcp.git ~/analyze-image-mcp
cd ~/analyze-image-mcp
npm install --omit=dev

npm install скачивает @modelcontextprotocol/sdk — официальную библиотеку протокола MCP (JSON-RPC через stdio). Папка node_modules/ должна оставаться рядом с index.js, иначе OpenCode не сможет запустить сервер.

Обновление:

cd ~/analyze-image-mcp && git pull && npm install --omit=dev

Вариант 2. Без клонирования, через npx

Ничего ставить заранее не нужно: npx сам скачает репозиторий из GitHub и зависимости при первом запуске (кэшируется). Команда для конфига приведена ниже. Чтобы зафиксировать версию, добавь тег или коммит: github:iljyxa/analyze-image-mcp#v1.0.0.

Конфигурация

Сервер настраивается переменными окружения:

Переменная

Описание

VISION_BASE_URL

Базовый URL OpenAI-compatible API, например https://your-provider.com/v1

VISION_API_KEY

API-ключ провайдера

VISION_MODEL

ID vision-модели

Сервер сам не читает .env — значения передаются через блок environment в конфиге OpenCode (см. ниже).

Подключение к OpenCode

В opencode.json (глобальный конфиг OpenCode — ~/.config/opencode/opencode.json):

{
  "$schema": "https://opencode.ai/config.json",
  "mcp": {
    "vision": {
      "type": "local",
      "command": ["node", "/home/<пользователь>/analyze-image-mcp/index.js"],
      "environment": {
        "VISION_BASE_URL": "https://твой-провайдер.com/v1",
        "VISION_API_KEY": "твой-ключ",
        "VISION_MODEL": "твоя-vision-модель"
      }
    }
  }
}

Путь в command — абсолютный (~ в конфиге не раскрывается). Для варианта 2 (npx) замени command на:

"command": ["npx", "-y", "github:iljyxa/analyze-image-mcp"]

Инструкция для агента

Чтобы основная модель сама вызывала инструмент при появлении изображения, добавь правило в системный промпт / AGENTS.md / настройки агента:

Если пользователь прикладывает изображение, а активная модель не может видеть изображения напрямую, всегда сначала вызывай инструмент analyze_image с путём к изображению, затем формируй ответ на основе полученного описания.

Проверка работы

  1. Перезапусти OpenCode после изменения конфига.

  2. Спроси у агента: "есть ли у тебя инструмент analyze_image?"

  3. Приложи изображение и задай вопрос по нему — модель должна вызвать analyze_image и ответить на основе полученного описания.

Известные ограничения

  • MCP-инструмент не перехватывает изображение автоматически — модель должна сама решить вызвать analyze_image. Надёжность зависит от инструкции в системном промпте и от того, насколько хорошо основная модель следует правилам вызова инструментов.

  • Поддерживаемые форматы изображений: PNG, JPEG, GIF, WebP.

  • Максимальный размер ответа vision-модели ограничен max_tokens: 2048 — при необходимости увеличь в index.js.

Лицензия

MIT

Available Tools

1 tool
analyze_imageA

Analyzes an image using a dedicated vision model and returns a detailed text description. Use this whenever the user attaches or references an image and the active model cannot see images itself. Accepts a local file path, file:// URI, http(s) URL, or data: URL.

ParametersJSON Schema
NameRequiredDescriptionDefault
questionNoOptional specific question about the image. If omitted, a full generic description is returned.
image_pathYesLocal path, file:// URI, http(s) URL, or data: URL of the image to analyze.

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full disclosure burden. It usefully reveals that analysis is delegated to an external vision model and that output is prose text, but it says nothing about latency, cost, image size/format limits, or failure behavior when a path or URL is unreachable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with the core purpose front-loaded, followed by the usage condition and input formats. Every sentence carries distinct information and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly compensates by stating that a detailed text description is returned, and it covers triggering conditions and accepted input forms. It is close to complete; the remaining gap is operational behavior (limits, cost/latency, error handling) for an external-model call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented. The description's enumeration of accepted input forms (local path, file:// URI, http(s) URL, data: URL) restates the image_path schema text rather than adding new meaning, and it says nothing about the optional 'question' parameter's effect beyond what the schema states. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb ('Analyzes'), the resource ('an image'), the mechanism ('a dedicated vision model'), and the result ('a detailed text description'). There are no sibling tools to disambiguate from, so the statement is fully self-contained and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit trigger ('whenever the user attaches or references an image') plus a gating condition that functions as a when-not rule ('the active model cannot see images itself'). No alternatives exist to name, so nothing further is required for correct selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev1.0.0
    • First observedanalyze_image

TDQS

A4.3/5.0

Scored across 1 tool

Disambiguation5/5

There is only one tool, so there is no possibility of confusing it with another. Its purpose—analyze an image and return a text description—is unambiguous, and the description clearly scopes when to use it.

Naming Consistency5/5

The single tool name 'analyze_image' follows a clean verb_noun snake_case convention. With no other names to conflict with, consistency is trivially perfect.

Tool Count4/5

The server's scope is narrowly a single capability—image analysis—so one tool is largely justified and each tool earns its place. It sits below the typical 3-15 range, and optional additions like batch analysis or multi-image comparison could be argued for, keeping it just short of ideal.

Completeness4/5

The tool covers the core lifecycle for its purpose: it accepts local paths, file URIs, http(s) URLs, and data URLs, so most image-referencing workflows are reachable. Minor gaps exist around batch/multiple-image handling and structured output options, but no dead ends for the primary use case.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers