Skip to main content
Glama

MCP オープンビジョン

CI PyPIバージョン Pythonのバージョン ライセンス: MIT コーヒーを買ってください 鍛冶屋のバッジ

概要

MCP OpenVisionは、OpenRouterビジョンモデルを活用した画像解析機能を提供するモデルコンテキストプロトコル(MCP)サーバーです。MCPエコシステム内のシンプルなインターフェースを介して、AIアシスタントが画像を解析できるようになります。

Related MCP server: MCP OpenVision

インストール

Smithery経由でインストール

Smithery経由で Claude Desktop 用の mcp-openvision を自動的にインストールするには:

npx -y @smithery/cli install @Nazruden/mcp-openvision --client claude

pipの使用

pip install mcp-openvision

UVの使用(推奨)

uv pip install mcp-openvision

構成

MCP OpenVision には OpenRouter API キーが必要であり、環境変数を通じて設定できます。

  • OPENROUTER_API_KEY (必須): OpenRouter APIキー

  • OPENROUTER_DEFAULT_MODEL (オプション): 使用するビジョンモデル

OpenRouter ビジョンモデル

MCP OpenVisionは、ビジョン機能をサポートするあらゆるOpenRouterモデルで動作します。デフォルトのモデルはqwen/qwen2.5-vl-32b-instruct:freeですが、互換性のある他のモデルを指定することもできます。

OpenRouter で利用できる一般的なビジョン モデルには次のようなものがあります。

  • qwen/qwen2.5-vl-32b-instruct:free (デフォルト)

  • anthropic/claude-3-5-sonnet

  • anthropic/claude-3-opus

  • anthropic/claude-3-sonnet

  • openai/gpt-4o

OPENROUTER_DEFAULT_MODEL環境変数を設定するか、 modelパラメータをimage_analysis関数に直接渡すことで、カスタム モデルを指定できます。

使用法

MCP Inspectorによるテスト

MCP OpenVision をテストする最も簡単な方法は、MCP Inspector ツールを使用することです。

npx @modelcontextprotocol/inspector uvx mcp-openvision

Claude DesktopまたはCursorとの統合

  1. MCP 構成ファイルを編集します。

    • Windows: %USERPROFILE%\.cursor\mcp.json

    • macOS: ~/.cursor/mcp.jsonまたは~/Library/Application Support/Claude/claude_desktop_config.json

  2. 次の構成を追加します。

{
  "mcpServers": {
    "openvision": {
      "command": "uvx",
      "args": ["mcp-openvision"],
      "env": {
        "OPENROUTER_API_KEY": "your_openrouter_api_key_here",
        "OPENROUTER_DEFAULT_MODEL": "anthropic/claude-3-sonnet"
      }
    }
  }
}

開発のためにローカルで実行する

# Set the required API key
export OPENROUTER_API_KEY="your_api_key"

# Run the server module directly
python -m mcp_openvision

特徴

MCP OpenVision は次のコア ツールを提供します。

  • image_analysis : さまざまなパラメータをサポートするビジョンモデルを使用して画像を分析します。

    • image : 次のように提供できます:

      • Base64エンコードされた画像データ

      • 画像URL(http/https)

      • ローカルファイルパス

    • query : 画像解析タスクのユーザー指示

    • system_prompt : モデルの役割と動作を定義する指示(オプション)

    • model : 使用するビジョンモデル

    • temperature : ランダム性を制御します (0.0-1.0)

    • max_tokens : 最大レスポンス長

効果的なクエリの作成

queryパラメータは、画像分析から有用な結果を得るために不可欠です。適切に作成されたクエリは、以下のコンテキストを提供します。

  1. 目的: この画像を分析する理由

  2. 焦点領域: 注目すべき特定の要素または詳細

  3. 必要な情報: 抽出する必要がある情報の種類

  4. フォーマット設定: 結果をどのように構造化するか

効果的なクエリの例

基本クエリ

拡張クエリ

「この画像を説明してください」

「この店舗の棚の画像に表示されているすべての小売製品を識別し、その価格帯を推定してください」

「この画像には何があるの?」

「この医療スキャンを分析して異常がないか調べ、強調表示された領域に焦点を当て、考えられる診断を提供します。」

「このチャートを分析してください」

「四半期ごとの売上を示すこの棒グラフから数値データを抽出し、2022年から2023年の主要な傾向を特定します。」

「テキストを読む」

「このレストランのメニューに表示されているすべてのテキストを、品名、説明、価格を残して書き写してください」

分析が必要な理由や、求めている具体的な情報についてのコンテキストを提供することで、モデルが関連する詳細に焦点を合わせ、より価値のある洞察を生み出すのに役立ちます。

使用例

# Analyze an image from a URL
result = await image_analysis(
    image="https://example.com/image.jpg",
    query="Describe this image in detail"
)

# Analyze an image from a local file with a focused query
result = await image_analysis(
    image="path/to/local/image.jpg",
    query="Identify all traffic signs in this street scene and explain their meanings for a driver education course"
)

# Analyze with a base64-encoded image and a specific analytical purpose
result = await image_analysis(
    image="SGVsbG8gV29ybGQ=...",  # base64 data
    query="Examine this product packaging design and highlight elements that could be improved for better visibility and brand recognition"
)

# Customize the system prompt for specialized analysis
result = await image_analysis(
    image="path/to/local/image.jpg",
    query="Analyze the composition and artistic techniques used in this painting, focusing on how they create emotional impact",
    system_prompt="You are an expert art historian with deep knowledge of painting techniques and art movements. Focus on formal analysis of composition, color, brushwork, and stylistic elements."
)

画像入力タイプ

image_analysisツールは、いくつかの種類の画像入力を受け入れます。

  1. Base64エンコードされた文字列

  2. 画像の URL - http:// または https:// で始まる必要があります

  3. ファイルパス:

    • 絶対パス: / (Unix) またはドライブ文字 (Windows) で始まる完全なパス

    • 相対パス: 現在の作業ディレクトリからの相対パス

    • project_root を使用した相対パス: project_rootパラメータを使用してベースディレクトリを指定します。

相対パスの使用

相対ファイル パス (「examples/image.jpg」など) を使用する場合は、次の 2 つのオプションがあります。

  1. パスは、サーバーが動作している現在の作業ディレクトリからの相対パスでなければなりません。

  2. または、 project_rootパラメータを指定することもできます。

# Example with relative path and project_root
result = await image_analysis(
    image="examples/image.jpg",
    project_root="/path/to/your/project",
    query="What is in this image?"
)

これは、現在の作業ディレクトリが予測できないアプリケーションや、特定のディレクトリに対する相対パスを使用してファイルを参照する場合に特に便利です。

発達

開発環境のセットアップ

# Clone the repository
git clone https://github.com/modelcontextprotocol/mcp-openvision.git
cd mcp-openvision

# Install development dependencies
pip install -e ".[dev]"

コードのフォーマット

このプロジェクトでは、Blackを使って自動コードフォーマットを行っています。フォーマットはGitHub Actionsを通じて強制されます。

  • リポジトリにプッシュされたすべてのコードは自動的に黒でフォーマットされます

  • リポジトリの協力者からのプルリクエストの場合、ブラックはコードをフォーマットし、PRブランチに直接コミットします。

  • フォークからのプルリクエストの場合、ブラックは元のPRにマージできるフォーマットされたコードを含む新しいPRを作成します。

コミットする前に、Black をローカルで実行してコードをフォーマットすることもできます。

# Format all Python code in the src and tests directories
black src tests

テストを実行する

pytest

リリースプロセス

このプロジェクトでは、自動化されたリリース プロセスを使用します。

  1. セマンティックバージョニングの原則に従ってpyproject.tomlのバージョンを更新します。

    • ヘルパースクリプトを使用できます: python scripts/bump_version.py [major|minor|patch]

  2. CHANGELOG.md新しいバージョンの詳細で更新します。

    • このスクリプトはCHANGELOG.mdにテンプレートエントリを作成し、それを入力することができます。

  3. これらの変更をコミットしてmainブランチにプッシュします

  4. GitHub Actions ワークフローは次のようになります。

    • バージョンの変更を検出する

    • 新しいGitHubリリースを自動的に作成する

    • PyPIに公開する公開ワークフローをトリガーする

この自動化により、一貫したリリース プロセスが維持され、すべてのリリースが適切にバージョン管理され、文書化されることが保証されます。

サポート

このプロジェクトが役に立つと思われる場合は、進行中の開発とメンテナンスをサポートするために私にコーヒーを買っていただけると幸いです。

ライセンス

このプロジェクトは MIT ライセンスに基づいてライセンスされています - 詳細についてはLICENSEファイルを参照してください。

Available Tools

1 tool
image_analysisA
Analyze an image using OpenRouter's vision capabilities.

This tool allows you to send an image to OpenRouter's vision models for analysis.
You provide a query to guide the analysis and can optionally customize the system prompt
for more control over the model's behavior.

Args:
    image: The image as a base64-encoded string, URL, or local file path
    query: Text prompt to guide the image analysis. For best results, provide context
           about why you're analyzing the image and what specific information you need.
           Including details about your purpose and required focus areas leads to more
           relevant and useful responses.
    system_prompt: Instructions for the model defining its role and behavior
    model: The vision model to use (defaults to the value set by OPENROUTER_DEFAULT_MODEL)
    max_tokens: Maximum number of tokens in the response (100-4000)
    temperature: Temperature parameter for generation (0.0-1.0)
    top_p: Optional nucleus sampling parameter (0.0-1.0)
    presence_penalty: Optional penalty for new tokens based on presence in text so far (0.0-2.0)
    frequency_penalty: Optional penalty for new tokens based on frequency in text so far (0.0-2.0)
    project_root: Optional root directory to resolve relative image paths against

Returns:
    The analysis result as text

Examples:
    Basic usage with a file path:
        image_analysis(image="path/to/image.jpg", query="Describe this image in detail")

    Basic usage with an image URL:
        image_analysis(image="https://example.com/image.jpg", query="Describe this image in detail")

    Basic usage with a relative path and project root:
        image_analysis(image="examples/image.jpg", project_root="/path/to/project", query="Describe this image in detail")

    Usage with a detailed contextual query:
        image_analysis(
            image="path/to/image.jpg",
            query="Analyze this product packaging design for a fitness supplement. Identify all nutritional claims,
                  certifications, and health icons. Assess the visual hierarchy and how the key selling points
                  are communicated. This is for a competitive analysis project."
        )

    Usage with custom system prompt:
        image_analysis(
            image="path/to/image.jpg",
            query="What objects can you see in this image?",
            system_prompt="You are an expert at identifying objects in images. Focus on listing all visible objects."
        )
ParametersJSON Schema
NameRequiredDescriptionDefault
imageYes
queryNoDescribe this image in detail
system_promptNoYou are an expert vision analyzer with exceptional attention to detail. Your purpose is to provide accurate, comprehensive descriptions of images that help AI agents understand visual content they cannot directly perceive. Focus on describing all relevant elements in the image - objects, people, text, colors, spatial relationships, actions, and context. Be precise but concise, organizing information from most to least important. Avoid making assumptions beyond what's visible and clearly indicate any uncertainty. When text appears in images, transcribe it verbatim within quotes. Respond only with factual descriptions without subjective judgments or creative embellishments. Your descriptions should enable an agent to make informed decisions based solely on your analysis.
modelNo
max_tokensNo
temperatureNo
top_pNo
presence_penaltyNo
frequency_penaltyNo
project_rootNo

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Without annotations, the description carries full burden for behavioral disclosure. It explains that the tool uses OpenRouter's vision models and returns text, and it lists default parameter values. However, it lacks information about external API dependencies, potential latency, failure modes, or rate limits, which are important for an agent to understand. The description is adequate but not comprehensive in this regard.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with a clear introductory sentence, parameter list, return value, and examples. While it is verbose in parts (e.g., the query parameter explanation is lengthy), every sentence adds value. It could be slightly more concise, but it is appropriately sized for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (10 parameters, no output schema, no annotations), the description is highly complete. It explains all parameters, specifies return type ('The analysis result as text'), and provides comprehensive examples covering various use cases. The default system prompt is also elaborated, which adds valuable context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate entirely. It does so excellently by providing detailed explanations for all 10 parameters, including their purpose, defaults, and constraints. For example, it explains that 'query' should include context for better results and provides examples. This enables correct parameter usage without relying on the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Analyze an image using OpenRouter's vision capabilities.' It specifies the action (analyze), resource (image), and technology (OpenRouter's vision), leaving no ambiguity. With no sibling tools, differentiation is not needed, but the purpose is specific and actionable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear usage guidance through detailed parameter descriptions and multiple examples covering file paths, URLs, contextual queries, and custom system prompts. However, it does not explicitly state when not to use this tool or mention alternatives, though none exist. 'Clear context, no exclusions' accurately reflects this.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

TDQS

A4.1/5.0
Disambiguation5/5

With only one tool, there is no possibility of confusion between tools. The tool's purpose is clear and unique.

Naming Consistency5/5

A single tool means no inconsistency in naming patterns. The name 'image_analysis' is descriptive and follows a common noun_noun convention.

Tool Count3/5

A single tool for a vision server feels slightly thin. While the one tool is comprehensive, the server scope seems narrow; typically 3-15 tools are expected for a well-scoped server.

Completeness2/5

The server only offers image analysis. Missing other common vision operations like model listing, batch processing, or generation. The surface is incomplete for a vision-focused server.

Maintenance

ActivityInactive
ResponsivenessUnresponsive

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/Nazruden/mcp-openvision'

If you have feedback or need assistance with the MCP directory API, please join our Discord server