MCP OpenVision

MCP OpenVision
Обзор
MCP OpenVision — это сервер Model Context Protocol (MCP), который предоставляет возможности анализа изображений на основе моделей зрения OpenRouter. Он позволяет помощникам ИИ анализировать изображения через простой интерфейс в экосистеме MCP.
Related MCP server: MCP Read Images
Установка
Установка через Smithery
Чтобы автоматически установить mcp-openvision для Claude Desktop через Smithery :
npx -y @smithery/cli install @Nazruden/mcp-openvision --client claudeИспользование пипа
pip install mcp-openvisionИспользование УФ (рекомендуется)
uv pip install mcp-openvisionКонфигурация
MCP OpenVision требует API-ключа OpenRouter и может быть настроен с помощью переменных среды:
OPENROUTER_API_KEY (обязательно): Ваш ключ API OpenRouter
OPENROUTER_DEFAULT_MODEL (необязательно): модель зрения, которую следует использовать
Модели видения OpenRouter
MCP OpenVision работает с любой моделью OpenRouter, которая поддерживает возможности Vision. Модель по умолчанию — qwen/qwen2.5-vl-32b-instruct:free , но вы можете указать любую другую совместимую модель.
Некоторые популярные модели машинного зрения, доступные через OpenRouter, включают:
qwen/qwen2.5-vl-32b-instruct:free(по умолчанию)anthropic/claude-3-5-sonnetanthropic/claude-3-opusanthropic/claude-3-sonnetopenai/gpt-4o
Вы можете указать пользовательские модели, установив переменную среды OPENROUTER_DEFAULT_MODEL или передав параметр model непосредственно в функцию image_analysis .
Использование
Тестирование с помощью MCP Inspector
Самый простой способ протестировать MCP OpenVision — использовать инструмент MCP Inspector:
npx @modelcontextprotocol/inspector uvx mcp-openvisionИнтеграция с Claude Desktop или Cursor
Отредактируйте файл конфигурации MCP:
Windows:
%USERPROFILE%\.cursor\mcp.jsonmacOS:
~/.cursor/mcp.jsonили~/Library/Application Support/Claude/claude_desktop_config.json
Добавьте следующую конфигурацию:
{
"mcpServers": {
"openvision": {
"command": "uvx",
"args": ["mcp-openvision"],
"env": {
"OPENROUTER_API_KEY": "your_openrouter_api_key_here",
"OPENROUTER_DEFAULT_MODEL": "anthropic/claude-3-sonnet"
}
}
}
}Локальный запуск для разработки
# Set the required API key
export OPENROUTER_API_KEY="your_api_key"
# Run the server module directly
python -m mcp_openvisionФункции
MCP OpenVision предоставляет следующий основной инструмент:
image_analysis : Анализ изображений с помощью моделей зрения, поддерживающих различные параметры:
image: Может быть предоставлено как:Данные изображения, закодированные в Base64
URL-адрес изображения (http/https)
Локальный путь к файлу
query: Инструкция пользователя для задачи анализа изображенияsystem_prompt: Инструкции, определяющие роль и поведение модели (необязательно)model: модель видения для использованияtemperature: контролирует случайность (0,0-1,0)max_tokens: Максимальная длина ответа
Создание эффективных запросов
Параметр query имеет решающее значение для получения полезных результатов анализа изображения. Хорошо составленный запрос предоставляет контекст о:
Цель : Почему вы анализируете это изображение
Области внимания : Конкретные элементы или детали, на которые следует обратить внимание.
Требуемая информация : тип информации, которую вам необходимо извлечь.
Настройки формата : как вы хотите структурировать результаты
Примеры эффективных запросов
Базовый запрос | Расширенный запрос |
«Опишите это изображение» | «Определите все розничные товары, которые видны на этом изображении полки магазина, и оцените их ценовой диапазон» |
«Что на этом изображении?» | «Проанализируйте это медицинское сканирование на предмет отклонений, сосредоточившись на выделенной области и поставив возможные диагнозы» |
«Проанализируйте эту диаграмму» | «Извлеките числовые данные из этой гистограммы, показывающей квартальные продажи, и определите ключевые тенденции на 2022–2023 годы» |
«Прочитай текст» | «Перепишите весь видимый текст в меню этого ресторана, сохранив названия блюд, описания и цены» |
Предоставляя контекст относительно того, зачем вам нужен анализ и какую конкретную информацию вы ищете, вы помогаете модели сосредоточиться на важных деталях и вырабатывать более ценную информацию.
Пример использования
# Analyze an image from a URL
result = await image_analysis(
image="https://example.com/image.jpg",
query="Describe this image in detail"
)
# Analyze an image from a local file with a focused query
result = await image_analysis(
image="path/to/local/image.jpg",
query="Identify all traffic signs in this street scene and explain their meanings for a driver education course"
)
# Analyze with a base64-encoded image and a specific analytical purpose
result = await image_analysis(
image="SGVsbG8gV29ybGQ=...", # base64 data
query="Examine this product packaging design and highlight elements that could be improved for better visibility and brand recognition"
)
# Customize the system prompt for specialized analysis
result = await image_analysis(
image="path/to/local/image.jpg",
query="Analyze the composition and artistic techniques used in this painting, focusing on how they create emotional impact",
system_prompt="You are an expert art historian with deep knowledge of painting techniques and art movements. Focus on formal analysis of composition, color, brushwork, and stylistic elements."
)Типы входных изображений
Инструмент image_analysis принимает несколько типов входных изображений:
Строки в кодировке Base64
URL-адреса изображений должны начинаться с http:// или https://
Пути к файлам :
Абсолютные пути : полные пути, начинающиеся с / (Unix) или буквы диска (Windows)
Относительные пути : пути относительно текущего рабочего каталога.
Относительные пути с project_root : используйте параметр
project_rootдля указания базового каталога.
Использование относительных путей
При использовании относительных путей к файлам (например, «examples/image.jpg») у вас есть два варианта:
Путь должен быть относительным к текущему рабочему каталогу, в котором запущен сервер.
Или вы можете указать параметр
project_root:
# Example with relative path and project_root
result = await image_analysis(
image="examples/image.jpg",
project_root="/path/to/your/project",
query="What is in this image?"
)Это особенно полезно в приложениях, где текущий рабочий каталог может быть непредсказуемым или когда вы хотите ссылаться на файлы, используя пути относительно определенного каталога.
Разработка
Настройка среды разработки
# Clone the repository
git clone https://github.com/modelcontextprotocol/mcp-openvision.git
cd mcp-openvision
# Install development dependencies
pip install -e ".[dev]"Форматирование кода
Этот проект использует Black для автоматического форматирования кода. Форматирование осуществляется через GitHub Actions:
Весь код, отправленный в репозиторий, автоматически форматируется черным цветом.
Для запросов на извлечение от участников репозитория Блэк форматирует код и фиксирует его непосредственно в ветке PR.
Для запросов на извлечение из форков Блэк создает новый PR с отформатированным кодом, который можно объединить с исходным PR.
Вы также можете запустить Black локально, чтобы отформатировать свой код перед фиксацией:
# Format all Python code in the src and tests directories
black src testsПроведение тестов
pytestПроцесс выпуска
В этом проекте используется автоматизированный процесс выпуска:
Обновите версию в
pyproject.toml, следуя принципам семантического версионирования.Вы можете использовать вспомогательный скрипт:
python scripts/bump_version.py [major|minor|patch]
Обновите
CHANGELOG.mdподробностями о новой версии.Скрипт также создает шаблонную запись в CHANGELOG.md, которую вы можете заполнить.
Зафиксируйте и отправьте эти изменения в
mainветку.Рабочий процесс GitHub Actions будет:
Обнаружить изменение версии
Автоматически создавать новый релиз GitHub
Запустите рабочий процесс публикации, который будет публиковаться в PyPI
Такая автоматизация помогает поддерживать согласованный процесс выпуска и гарантирует, что каждый выпуск имеет надлежащую версию и документируется.
Поддерживать
Если вы считаете этот проект полезным, рассмотрите возможность угостить меня кофе, чтобы поддержать текущую разработку и поддержку.
Лицензия
Данный проект лицензирован по лицензии MIT — подробности см. в файле LICENSE .
Available Tools
1 toolimage_analysisA
Analyze an image using OpenRouter's vision capabilities.
This tool allows you to send an image to OpenRouter's vision models for analysis.
You provide a query to guide the analysis and can optionally customize the system prompt
for more control over the model's behavior.
Args:
image: The image as a base64-encoded string, URL, or local file path
query: Text prompt to guide the image analysis. For best results, provide context
about why you're analyzing the image and what specific information you need.
Including details about your purpose and required focus areas leads to more
relevant and useful responses.
system_prompt: Instructions for the model defining its role and behavior
model: The vision model to use (defaults to the value set by OPENROUTER_DEFAULT_MODEL)
max_tokens: Maximum number of tokens in the response (100-4000)
temperature: Temperature parameter for generation (0.0-1.0)
top_p: Optional nucleus sampling parameter (0.0-1.0)
presence_penalty: Optional penalty for new tokens based on presence in text so far (0.0-2.0)
frequency_penalty: Optional penalty for new tokens based on frequency in text so far (0.0-2.0)
project_root: Optional root directory to resolve relative image paths against
Returns:
The analysis result as text
Examples:
Basic usage with a file path:
image_analysis(image="path/to/image.jpg", query="Describe this image in detail")
Basic usage with an image URL:
image_analysis(image="https://example.com/image.jpg", query="Describe this image in detail")
Basic usage with a relative path and project root:
image_analysis(image="examples/image.jpg", project_root="/path/to/project", query="Describe this image in detail")
Usage with a detailed contextual query:
image_analysis(
image="path/to/image.jpg",
query="Analyze this product packaging design for a fitness supplement. Identify all nutritional claims,
certifications, and health icons. Assess the visual hierarchy and how the key selling points
are communicated. This is for a competitive analysis project."
)
Usage with custom system prompt:
image_analysis(
image="path/to/image.jpg",
query="What objects can you see in this image?",
system_prompt="You are an expert at identifying objects in images. Focus on listing all visible objects."
)
| Name | Required | Description | Default |
|---|---|---|---|
| image | Yes | ||
| query | No | Describe this image in detail | |
| system_prompt | No | You are an expert vision analyzer with exceptional attention to detail. Your purpose is to provide accurate, comprehensive descriptions of images that help AI agents understand visual content they cannot directly perceive. Focus on describing all relevant elements in the image - objects, people, text, colors, spatial relationships, actions, and context. Be precise but concise, organizing information from most to least important. Avoid making assumptions beyond what's visible and clearly indicate any uncertainty. When text appears in images, transcribe it verbatim within quotes. Respond only with factual descriptions without subjective judgments or creative embellishments. Your descriptions should enable an agent to make informed decisions based solely on your analysis. | |
| model | No | ||
| max_tokens | No | ||
| temperature | No | ||
| top_p | No | ||
| presence_penalty | No | ||
| frequency_penalty | No | ||
| project_root | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden. It explains the core behavior (image analysis via OpenRouter's vision models) and mentions customization options, but doesn't disclose important behavioral traits like rate limits, authentication requirements, error conditions, or what happens with invalid inputs. The examples help but don't cover edge cases.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with clear sections (overview, args, returns, examples) and front-loads the core purpose. While comprehensive, some sentences could be more concise, particularly in the parameter explanations where some details are repeated across multiple examples.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (10 parameters, no annotations, no output schema), the description provides substantial context through detailed parameter explanations and multiple examples. However, it lacks information about return format details beyond 'text' and doesn't cover error handling or operational constraints that would be important for a vision analysis tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description provides extensive parameter documentation beyond the schema, which has 0% description coverage. It explains each parameter's purpose, format requirements (base64, URL, file path), ranges (max_tokens 100-4000), defaults, and provides detailed guidance for the query parameter. This fully compensates for the schema's lack of descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool 'analyzes an image using OpenRouter's vision capabilities' and specifies it's for sending images to vision models for analysis. It provides a specific verb ('analyze') and resource ('image'), but since there are no sibling tools, it doesn't need to differentiate from alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides implied usage through examples showing different scenarios (basic usage, detailed contextual queries, custom system prompts). However, it lacks explicit guidance on when to use this tool versus alternatives or any prerequisites for successful invocation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
With only one tool, there is no possibility of ambiguity or overlap between tools. The single tool 'image_analysis' has a clearly defined purpose that cannot be confused with any other tool in this server.
The single tool follows a clear verb_noun pattern ('image_analysis'), and with only one tool, there is no inconsistency to evaluate. The naming is straightforward and descriptive.
A single tool is too few for a server named 'MCP OpenVision' that implies broader vision capabilities. While the tool is well-described, the server feels thin and limited in scope, lacking complementary tools like image generation, comparison, or batch processing that would make it more complete.
The server is severely incomplete for a vision domain. It only provides image analysis, missing essential operations like image generation, editing, transformation, or multi-image processing. Agents will hit dead ends when needing to perform common vision tasks beyond analysis.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
A comprehensive Model Context Protocol (MCP) server that enables AI assistants to interact with yo…
The OpenRouter MCP server plugs OpenRouter into the AI tools you already use. Once connected, your assistant can pull live OpenRouter data (models, prices, your credits, rankings, and docs) and send quick test messages, all without leaving your editor.
A Model Context Protocol server for Wix AI tools
MCP server for building and testing AI agents with multi-model experimentation and insights.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceA powerful server that integrates the Moondream vision model to enable advanced image analysis, including captioning, object detection, and visual question answering, through the Model Context Protocol, compatible with AI assistants like Claude and Cline.19Apache 2.0
- AlicenseNot gradedqualityDmaintenanceAn MCP server for analyzing images using OpenRouter vision models, offering capabilities like automatic image resizing, model configuration, and handling custom queries about images.10MIT
- AlicenseAqualityDmaintenanceMCP OpenVision is a Model Context Protocol (MCP) server that provides image analysis capabilities powered by OpenRouter vision models. It enables AI assistants to analyze images via a simple interface within the MCP ecosystem.116MIT
- AlicenseNot gradedqualityCmaintenanceA Model Context Protocol server that provides multimodal vision tools such as image description, OCR, visual Q&A, and object detection, powered by any vision model via OpenRouter.MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/mikeysrecipes/mcp-openvision'
If you have feedback or need assistance with the MCP directory API, please join our Discord server