eyes-mcp
eyes-mcp
テキストのみのLLMに目を与えます。コーディングエージェントのためのローカルビジョン。
1コマンド・APIキー不要・あなたのマシンから何も出て行かない
DeepSeek、GLM、Qwen-Coder、Llama… 優れたモデルだが、すべて目が見えない。
❌ 目がない場合
テキストのみのモデルを実行しているエージェント(Claude Code / Codex / Cursor 経由)にスクリーンショットを貼り付けます:
> Here's the error in my UI, fix it [screenshot.png]
I'm sorry — I cannot see images. Please describe the error in text.Related MCP server: OpenSight MCP
✅ 目がある場合
代わりに、エージェントはローカルのVLM + OCRを呼び出し、スクリーンショットを自分で読み取ります:
> Here's the error in my UI, fix it [screenshot.png]
I see a React hydration error in `CartDrawer.tsx:142`. The OCR shows:
"Hydration failed because the server rendered HTML didn't match the client." …クイックスタート
git clone https://github.com/JamesbbBriz/eyes-mcp
cd eyes-mcp && ./scripts/install.shこれだけです。インストーラは:
希望するモデルを尋ねます。RAMとGPUから計算された推奨モデル付きです(
EYES_PRESETまたは--yesで質問をスキップできます)。依存関係をインストールし、モデルをダウンロードします(約0.3〜3.5GB、再開可能)。
Claude Code / Codex / Cursor の設定を読み取り、各モデルをモダリティデータベースと照合して、テキストのみのモデルを実行しているエージェントを検出します。
必要な場所にのみ eyes-mcp を登録します。マルチモーダルエージェントは自動的にスキップされます。
# options:
EYES_PRESET=fast ./scripts/install.sh # Qwen3.5-0.8B, natively multimodal
HF_ENDPOINT=https://hf-mirror.com ./install.sh # mainland-CN mirror
./install.sh --yes # accept all recommendations, no prompts
./install.sh --dry-run # preview without changing anythingモダリティチェックだけが必要ですか? python3 scripts/detect_modality.py
必要条件: Python ≥3.11、llama.cpp(brew install llama.cpp)、約1GBのRAM。
手動登録
自動インストールをスキップした場合、またはインストーラが認識しないエージェントは?手動で追加してください。
Claude Code(~/.claude.json → mcpServers):
"eyes-mcp": {
"command": "uv",
"args": ["--directory", "/ABS/PATH/eyes-mcp", "run", "eyes-mcp"],
"env": { "EYES_PRESET": "lfm-450m" }
}Codex(~/.codex/config.toml):
[mcp_servers.eyes-mcp]
command = "uv"
args = ["--directory", "/ABS/PATH/eyes-mcp", "run", "eyes-mcp"]
env = { EYES_PRESET = "lfm-450m" }Cursor(.cursor/mcp.json): Claude Code と同じ形式です。
エージェントを再起動して、次のように尋ねてください: 「このスクリーンショットには何が写っていますか?」
ツール
ツール | エンジン | 用途 |
| llama.cpp 経由のVLM | 説明、UI理解、ビジュアルQ&A |
| RapidOCR (onnx) | 高密度テキスト: ターミナル、ドキュメント、表;高速かつ高精度 |
モデルプリセット
プリセット | モデル | ダウンロード | RAM | ライセンス | 備考 |
| SmolVLM2-256M | ~0.3GB | ~1GB | Apache-2.0 | 最小限の実用VLM |
| LFM2.5-VL-450M | ~0.4GB | ~1.2GB | テスト済み;最速の起動 | |
| Qwen3.5-0.8B | ~0.7GB | ~1.8GB | Apache-2.0 | ネイティブでマルチモーダル(画像+動画) |
| GLM-OCR | ~1.4GB | ~3.5GB | MIT | 高密度テキスト/ドキュメントの王者(月間300万以上ダウンロード) |
| Qwen3.5-2B | ~2GB | ~3.5GB | Apache-2.0 | 最高の品質/サイズバランス |
| Qwen3.5-4B | ~3GB | ~6GB | Apache-2.0 | 最大ティア(GPU推奨) |
隠しエクストラ(同じく1コマンド): smol500(SmolVLM2-500M)、paddle(PaddleOCR-VL-1.6)、qwen3-2b(Qwen3-VL-2B)。
他のGGUFでも動作します。環境変数で指定して、プリセットを完全にスキップできます:
EYES_MODEL_DIR=~/models/my-vlm VLM_MODEL_FILE=model-Q4.gguf VLM_MMPROJ_FILE=mmproj.ggufプリセットとして同梱されていない良い候補: LFM2.5-VL-1.6B/3B、InternVL3.5-2B/4B、MiniCPM-V-4.6、DeepSeek-OCR、dots.ocr、gemma-3n-E2B、moondream2。llama.cpp が mmproj ファイル付きでサポートしているものなら何でも動作します。
いつでも切り替え可能: EYES_PRESET を設定して ./scripts/download_models.sh を再度実行してください。どれがいいか分からない? python3 scripts/choose_model.py が RAM/GPU を表示し、推奨をマークします。
仕組み
Claude Code / Codex / Cursor
│ MCP stdio
▼
eyes-mcp (stateless, mcp SDK 2.x)
├─ analyze_image → llama.cpp llama-server (local VLM) "understand"
└─ ocr_image → RapidOCR (onnx, ~20MB) "extract text"ライフサイクルはエージェントに追従: VLMサーバーはMCPの開始時に起動し、エージェントの終了時にシャットダウンするため、孤立プロセスや常駐デーモンの管理が不要です。
フローティングポート: VLMは固定ポートにバインドしないため(「8080は既に使用中」とはおさらば)、他のローカルサービスと共存できます。
外部VLMの再利用: すでに
VLM_BASE_URLでVLMを実行している場合、eyes-mcp は独自に起動する代わりにそれを使用します。
なぜ?
現在、最も安価で優れたコーディングモデル(DeepSeek-V4-Flash、GLM-5.x、Qwen-Coder)はテキストのみです。すべてのハーネスはスクリーンショットを貼り付けられることを前提としていますが、これらのモデルはどれも黙って失敗します。eyes-mcp は欠けているサイドカーです。小型のローカルVLMとOCRを、エージェントがすでに理解しているライフサイクルで包み込みます。
ロードマップ
遅延VLM起動(MCP開始時ではなく、最初のツール呼び出し時に起動)
screenshot_analyze(ファイル不要で画面を取得)PDFページ → ビジョン
npx eyes-mcpワンライナーインストーラモデルごとのプロンプトテンプレート(llama.cpp OCRモデルには特定のプロンプトが必要)
よくある質問
エージェントのモデルは重要ですか? これが役立つためには、モデルがテキストのみでなければならないという点だけです。マルチモーダルモデル(GPT、Claude、GLM-V)はすでに画像を見られるため、気にする必要はありません。
GPUは必要ですか? いいえ。CPUでも問題なく動作し、llama.cpp はApple MetalまたはCUDAがあれば自動的に使用します。
モデルはどこに保存されますか? ~/.eyes-mcp/models/<preset>/。リセットするには削除してください。
ライセンス
MIT。モデルの重みはそれぞれのライセンスに従います(プリセット表を参照)。インストール時にダウンロードされ、ここで再配布されることはありません。
This server cannot be deployed
Maintenance
Related MCP Connectors
OCR, transcription, file extraction, and image generation for AI agents via MCP.
Grabbit gives AI agents eyes on the web through a hosted MCP server. Send a public URL and get a pixel-perfect hosted image back, without maintaining Chromium, Playwright, or a browser fleet. Capture a full page, exact viewport, or single CSS selector as PNG, JPEG, or WebP. Grabbit handles cookie and consent banners, waits for JavaScript-heavy pages, blocks private and internal URLs, supports safe retries with idempotency keys, and delivers async results through signed webhooks. Completed captures include a CDN URL. Connect with OAuth 2.1 or an API key. Grabbit works with Claude, Cursor, Codex, and any MCP client. Live captures cost $0.002 each. The $50 annual plan includes 25,000 prepaid credits that never reset or expire. Free test keys return placeholder images, so you can wire up the integration before paying. Home: https://grabbit.live Docs: https://grabbit.live/screenshot-api Built by BrainGrid.
- mcpOAuthcom.screenshotink
Screenshot, diff, audit and sitemap-capture any web page — 5 MCP tools for AI agents.
Scrape, crawl and search the web for AI agents via MCP.
Related MCP Servers
- AlicenseAqualityBmaintenanceBridges vision models to text-only coding models using Florence-2, enabling non-vision LLMs to describe images, extract text, and analyze screenshots via MCP tools.6MIT
- AlicenseNot gradedqualityBmaintenanceMulti-backend AI vision for MCP agents. Analyze images, screenshots, and documents using local Ollama models or cloud APIs like OpenAI, Google Gemini, and OpenRouter.MIT
- AlicenseNot gradedqualityAmaintenanceLocal vision-capable MCP server that lets AI agents describe screenshots, UI, charts, and photos via vision and OCR tools, with support for multiple providers and automatic fallback.6MIT
- AlicenseNot gradedqualityCmaintenanceAdds image recognition and UI grounding capabilities to text-only LLMs through MCP tools, supporting local and cloud vision backends.40 npmMIT