analyze-image-mcp
Sends clean OpenAI-compatible vision API requests to a configured vision model provider, allowing text-only models to analyze images and receive text descriptions without using tool calling.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@analyze-image-mcpanalyze ~/Pictures/error.png and tell me what the error says"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
analyze-image-mcp
Минимальный MCP-сервер (пример подключения ниже — для OpenCode, но подойдёт любому MCP-клиенту), который даёт text-only модели доступ к отдельной vision-модели через один инструмент — analyze_image. Работает с любым OpenAI-compatible vision API (включая провайдеров с ограниченной поддержкой tool calling), так как обращается к нему напрямую, минуя механизм tools/tool_choice самого OpenCode.
Зачем это нужно
Если у тебя:
основная модель — только текстовая;
вторая модель — с поддержкой изображений, но у провайдера не включён
--enable-auto-tool-choice(частая ошибка"auto" tool choice requires --enable-auto-tool-choice and --tool-call-parser to be set);
то стандартная отправка картинки напрямую в OpenCode ломается. Этот MCP-сервер решает проблему: он вызывает vision-модель отдельным чистым HTTP-запросом без полей tools/tool_choice, а результат (текстовое описание) отдаёт основной модели как обычный текст.
Related MCP server: VisionToolMCP
Как это работает
Пользователь прикладывает изображение
│
▼
Основная (text-only) модель видит инструкцию:
"если есть изображение — вызови analyze_image"
│
▼
MCP-сервер (analyze_image):
1. читает файл / URL / data-URI
2. кодирует в data:image/...;base64,...
3. отправляет чистый OpenAI-compatible запрос
(без tools, без tool_choice) в vision-провайдера
│
▼
Текстовое описание возвращается основной модели
│
▼
Основная модель отвечает пользователюУстановка
Нужен Node.js 18+ (используется встроенный fetch) и git. Пакет в npm не публикуется — ставится из GitHub. Выбери один из вариантов.
Вариант 1. Клонирование (рекомендуется)
git clone https://github.com/iljyxa/analyze-image-mcp.git ~/analyze-image-mcp
cd ~/analyze-image-mcp
npm install --omit=devnpm install скачивает @modelcontextprotocol/sdk — официальную библиотеку протокола MCP (JSON-RPC через stdio). Папка node_modules/ должна оставаться рядом с index.js, иначе OpenCode не сможет запустить сервер.
Обновление:
cd ~/analyze-image-mcp && git pull && npm install --omit=devВариант 2. Без клонирования, через npx
Ничего ставить заранее не нужно: npx сам скачает репозиторий из GitHub и зависимости при первом запуске (кэшируется). Команда для конфига приведена ниже. Чтобы зафиксировать версию, добавь тег или коммит: github:iljyxa/analyze-image-mcp#v1.0.0.
Конфигурация
Сервер настраивается переменными окружения:
Переменная | Описание |
| Базовый URL OpenAI-compatible API, например |
| API-ключ провайдера |
| ID vision-модели |
Сервер сам не читает .env — значения передаются через блок environment в конфиге OpenCode (см. ниже).
Подключение к OpenCode
В opencode.json (глобальный конфиг OpenCode — ~/.config/opencode/opencode.json):
{
"$schema": "https://opencode.ai/config.json",
"mcp": {
"vision": {
"type": "local",
"command": ["node", "/home/<пользователь>/analyze-image-mcp/index.js"],
"environment": {
"VISION_BASE_URL": "https://твой-провайдер.com/v1",
"VISION_API_KEY": "твой-ключ",
"VISION_MODEL": "твоя-vision-модель"
}
}
}
}Путь в command — абсолютный (~ в конфиге не раскрывается). Для варианта 2 (npx) замени command на:
"command": ["npx", "-y", "github:iljyxa/analyze-image-mcp"]Инструкция для агента
Чтобы основная модель сама вызывала инструмент при появлении изображения, добавь правило в системный промпт / AGENTS.md / настройки агента:
Если пользователь прикладывает изображение, а активная модель не может видеть изображения напрямую, всегда сначала вызывай инструмент analyze_image с путём к изображению, затем формируй ответ на основе полученного описания.Проверка работы
Перезапусти OpenCode после изменения конфига.
Спроси у агента: "есть ли у тебя инструмент analyze_image?"
Приложи изображение и задай вопрос по нему — модель должна вызвать
analyze_imageи ответить на основе полученного описания.
Известные ограничения
MCP-инструмент не перехватывает изображение автоматически — модель должна сама решить вызвать
analyze_image. Надёжность зависит от инструкции в системном промпте и от того, насколько хорошо основная модель следует правилам вызова инструментов.Поддерживаемые форматы изображений: PNG, JPEG, GIF, WebP.
Максимальный размер ответа vision-модели ограничен
max_tokens: 2048— при необходимости увеличь вindex.js.
Лицензия
Available Tools
1 toolanalyze_imageA
Analyzes an image using a dedicated vision model and returns a detailed text description. Use this whenever the user attaches or references an image and the active model cannot see images itself. Accepts a local file path, file:// URI, http(s) URL, or data: URL.
| Name | Required | Description | Default |
|---|---|---|---|
| question | No | Optional specific question about the image. If omitted, a full generic description is returned. | |
| image_path | Yes | Local path, file:// URI, http(s) URL, or data: URL of the image to analyze. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full disclosure burden. It usefully reveals that analysis is delegated to an external vision model and that output is prose text, but it says nothing about latency, cost, image size/format limits, or failure behavior when a path or URL is unreachable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with the core purpose front-loaded, followed by the usage condition and input formats. Every sentence carries distinct information and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly compensates by stating that a detailed text description is returned, and it covers triggering conditions and accepted input forms. It is close to complete; the remaining gap is operational behavior (limits, cost/latency, error handling) for an external-model call.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are already documented. The description's enumeration of accepted input forms (local path, file:// URI, http(s) URL, data: URL) restates the image_path schema text rather than adding new meaning, and it says nothing about the optional 'question' parameter's effect beyond what the schema states. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Analyzes'), the resource ('an image'), the mechanism ('a dedicated vision model'), and the result ('a detailed text description'). There are no sibling tools to disambiguate from, so the statement is fully self-contained and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit trigger ('whenever the user attaches or references an image') plus a gating condition that functions as a when-not rule ('the active model cannot see images itself'). No alternatives exist to name, so nothing further is required for correct selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v1.0.0- First observed
analyze_image
TDQS
Scored across 1 tool
There is only one tool, so there is no possibility of confusing it with another. Its purpose—analyze an image and return a text description—is unambiguous, and the description clearly scopes when to use it.
The single tool name 'analyze_image' follows a clean verb_noun snake_case convention. With no other names to conflict with, consistency is trivially perfect.
The server's scope is narrowly a single capability—image analysis—so one tool is largely justified and each tool earns its place. It sits below the typical 3-15 range, and optional additions like batch analysis or multi-image comparison could be argued for, keeping it just short of ideal.
The tool covers the core lifecycle for its purpose: it accepts local paths, file URIs, http(s) URLs, and data URLs, so most image-referencing workflows are reachable. Minor gaps exist around batch/multiple-image handling and structured output options, but no dead ends for the primary use case.
Maintenance
Related MCP Connectors
OCR, transcription, file extraction, and image generation for AI agents via MCP.
Create images & video from any MCP agent — 17 models, spend limits, one URL.
Image & PDF tools for AI agents: compress, convert, resize, PDF, AI vision, pipeline.
Multimodal video analysis MCP — transcription, vision, and OCR for any video URL.
Related MCP Servers
- AlicenseAqualityDmaintenanceBridges a vision model to enable text-only models like DeepSeek to describe images, extract text, and compare images via MCP tools.534 npm9MIT
- FlicenseAqualityBmaintenanceEnables text-only agents to process images by accepting image files, base64 data, or URLs, sending them to multimodal models, and returning structured text results via MCP.41-
- AlicenseAqualityBmaintenanceEnables non-multimodal models to see images by providing MCP tools for image understanding and OCR, backed by any OpenAI-compatible vision model.2MIT
- FlicenseNot gradedqualityCmaintenanceProvides an MCP tool that analyzes images from local paths, URLs, or data URLs via a vision language model, returning structured descriptions (brief, detailed, summary) so text-only LLMs can understand image content.-