Skip to main content
Glama

eyes-mcp

텍스트 전용 LLM에 눈을 달아 주세요. 코딩 에이전트를 위한 로컬 비전.

한 줄이면 끝납니다 · API 키 불필요 · 모든 처리는 내 컴퓨터 안에서

License: MIT MCP llama.cpp 简体中文

DeepSeek, GLM, Qwen-Coder, Llama… 훌륭한 모델이지만 모두 눈이 멀었습니다.


❌ 눈이 없을 때

에이전트(Claude Code / Codex / Cursor로 실행되는 텍스트 전용 모델)에 스크린샷을 붙여넣으면:

> Here's the error in my UI, fix it  [screenshot.png]

I'm sorry — I cannot see images. Please describe the error in text.

Related MCP server: OpenSight MCP

✅ 눈이 있을 때

에이전트가 대신 로컬 VLM + OCR을 호출해 스크린샷을 직접 읽습니다:

> Here's the error in my UI, fix it  [screenshot.png]

I see a React hydration error in `CartDrawer.tsx:142`. The OCR shows:
"Hydration failed because the server rendered HTML didn't match the client." …

빠른 시작

git clone https://github.com/JamesbbBriz/eyes-mcp
cd eyes-mcp && ./scripts/install.sh

이게 전부입니다. 설치 프로그램은:

  1. RAM과 GPU를 기반으로 계산한 권장 모델을 포함해 어떤 모델을 사용할지 묻습니다(EYES_PRESET 또는 --yes로 질문을 건너뛸 수 있습니다).

  2. 의존성을 설치하고 모델을 다운로드합니다(~0.3~3.5GB, 중단 후 재개 가능).

  3. Claude Code / Codex / Cursor 구성을 읽고 각 모델을 모달리티 데이터베이스와 대조하여, 어떤 에이전트가 텍스트 전용 모델을 사용하는지 감지합니다.

  4. 필요한 곳에만 eyes-mcp를 등록합니다. 멀티모달 에이전트는 자동으로 건너뜁니다.

# options:
EYES_PRESET=fast ./scripts/install.sh            # Qwen3.5-0.8B, natively multimodal
HF_ENDPOINT=https://hf-mirror.com ./install.sh  # mainland-CN mirror
./install.sh --yes                              # accept all recommendations, no prompts
./install.sh --dry-run                          # preview without changing anything

모달리티 확인만 원하세요? python3 scripts/detect_modality.py

요구 사항: Python ≥3.11, llama.cpp (brew install llama.cpp), ~1GB RAM.

수동 등록

자동 설치를 건너뛰었거나 설치 프로그램이 모르는 에이전트가 있나요? 직접 추가하세요.

Claude Code (~/.claude.json → mcpServers):

"eyes-mcp": {
  "command": "uv",
  "args": ["--directory", "/ABS/PATH/eyes-mcp", "run", "eyes-mcp"],
  "env": { "EYES_PRESET": "lfm-450m" }
}

Codex (~/.codex/config.toml):

[mcp_servers.eyes-mcp]
command = "uv"
args = ["--directory", "/ABS/PATH/eyes-mcp", "run", "eyes-mcp"]
env = { EYES_PRESET = "lfm-450m" }

Cursor (.cursor/mcp.json): Claude Code와 동일한 형식입니다.

에이전트를 재시작한 후 이렇게 물어보세요: "이 스크린샷에 뭐가 있나요?"

도구

도구

엔진

용도

analyze_image(path, question?)

llama.cpp 기반 VLM

설명, UI 이해, 시각적 Q&A

ocr_image(path)

RapidOCR (onnx)

터미널·문서·표 등 텍스트가 많은 콘텐츠; 빠르고 정확함

모델 프리셋

프리셋

모델

다운로드

RAM

라이선스

참고

nano

SmolVLM2-256M

~0.3GB

~1GB

Apache-2.0

가장 작은 실용적인 VLM

lfm-450m (기본값)

LFM2.5-VL-450M

~0.4GB

~1.2GB

Liquid

테스트 완료; 시작이 가장 빠름

fast

Qwen3.5-0.8B

~0.7GB

~1.8GB

Apache-2.0

기본적으로 멀티모달(이미지 + 비디오)

ocr

GLM-OCR

~1.4GB

~3.5GB

MIT

텍스트가 많은 콘텐츠/문서에 최강 (월 300만+ 다운로드)

strong

Qwen3.5-2B

~2GB

~3.5GB

Apache-2.0

품질/용량 밸런스 최고

xstrong

Qwen3.5-4B

~3GB

~6GB

Apache-2.0

최상위 등급 (GPU 권장)

숨겨진 추가 모델(여전히 한 줄 명령): smol500 (SmolVLM2-500M), paddle (PaddleOCR-VL-1.6), qwen3-2b (Qwen3-VL-2B).

다른 GGUF도 모두 작동합니다. 환경 변수로 지정하면 프리셋을 완전히 건너뜁니다:

EYES_MODEL_DIR=~/models/my-vlm  VLM_MODEL_FILE=model-Q4.gguf  VLM_MMPROJ_FILE=mmproj.gguf

프리셋으로 제공되지 않는 좋은 선택지: LFM2.5-VL-1.6B/3B, InternVL3.5-2B/4B, MiniCPM-V-4.6, DeepSeek-OCR, dots.ocr, gemma-3n-E2B, moondream2. llama.cpp에서 mmproj 파일로 지원하는 모든 모델이 작동합니다.

언제든 전환하세요: EYES_PRESET을 설정하고 ./scripts/download_models.sh을 다시 실행하세요. 어떤 모델인지 모르겠다면? python3 scripts/choose_model.py가 RAM/GPU를 표시하고 권장 사항을 표시합니다.

작동 방식

Claude Code / Codex / Cursor
      │ MCP stdio
      ▼
eyes-mcp  (stateless, mcp SDK 2.x)
   ├─ analyze_image → llama.cpp llama-server (local VLM)  "understand"
   └─ ocr_image     → RapidOCR (onnx, ~20MB)             "extract text"
  • 수명 주기가 에이전트를 따릅니다: MCP가 시작될 때 VLM 서버가 실행되고 에이전트가 종료되면 함께 종료됩니다. 따라서 고아 프로세스나 관리할 데몬이 생기지 않습니다.

  • 유동 포트: VLM은 고정 포트를 바인딩하지 않으므로(안녕, "8080 already in use"), 다른 로컬 서비스와 공존할 수 있습니다.

  • 외부 VLM 재사용: 이미 VLM_BASE_URL에서 VLM을 실행 중이라면 eyes-mcp는 자체 서버를 띄우는 대신 그 VLM을 사용합니다.

왜 필요한가

요즘 가장 저렴하고 성능 좋은 코딩 모델(DeepSeek-V4-Flash, GLM-5.x, Qwen-Coder)은 텍스트 전용입니다. 모든 하네스가 스크린샷을 붙여넣을 수 있다고 가정하지만, 이 모델들은 모두 조용히 실패합니다. eyes-mcp는 바로 그 빠진 사이드카입니다. 작은 로컬 VLM과 OCR을 에이전트가 이미 이해하는 수명 주기로 감쌌습니다.

로드맵

  • 지연 VLM 시작(첫 도구 호출 시 생성, MCP 시작 시 아님)

  • screenshot_analyze (화면 캡처, 파일 불필요)

  • PDF 페이지 → 비전

  • npx eyes-mcp 한 줄 설치기

  • 모델별 프롬프트 템플릿(llama.cpp OCR 모델은 특정 프롬프트 필요)

FAQ

에이전트 모델이 중요한가요? 이 도구가 유용하려면 모델이 텍스트 전용이어야 한다는 점에서만 그렇습니다. 멀티모달 모델(GPT, Claude, GLM-V)은 이미 이미지를 볼 수 있으므로 신경 쓸 필요가 없습니다.

GPU가 필요한가요? 아닙니다. CPU에서도 잘 실행되며, llama.cpp는 Apple Metal이나 CUDA가 있으면 자동으로 사용합니다.

모델은 어디에 저장되나요? ~/.eyes-mcp/models/<preset>/에 저장됩니다. 삭제하면 초기화됩니다.

라이선스

MIT. 모델 가중치는 각자의 라이선스를 따릅니다(프리셋 표 참조). 설치 시 다운로드되며, 이 저장소에서 재배포하지 않습니다.


Related MCP Connectors

  • OCR, transcription, file extraction, and image generation for AI agents via MCP.

  • Grabbit gives AI agents eyes on the web through a hosted MCP server. Send a public URL and get a pixel-perfect hosted image back, without maintaining Chromium, Playwright, or a browser fleet. Capture a full page, exact viewport, or single CSS selector as PNG, JPEG, or WebP. Grabbit handles cookie and consent banners, waits for JavaScript-heavy pages, blocks private and internal URLs, supports safe retries with idempotency keys, and delivers async results through signed webhooks. Completed captures include a CDN URL. Connect with OAuth 2.1 or an API key. Grabbit works with Claude, Cursor, Codex, and any MCP client. Live captures cost $0.002 each. The $50 annual plan includes 25,000 prepaid credits that never reset or expire. Free test keys return placeholder images, so you can wire up the integration before paying. Home: https://grabbit.live Docs: https://grabbit.live/screenshot-api Built by BrainGrid.

  • Screenshot, diff, audit and sitemap-capture any web page — 5 MCP tools for AI agents.

  • Scrape, crawl and search the web for AI agents via MCP.

Related MCP Servers