claude-vision-mcp
README.md
# Claude Vision MCP
A [Model Context Protocol](https://modelcontextprotocol.io) server that gives Claude Code the ability to **see** — capturing the screen, individual windows, screen regions, and live web pages, plus comparing screenshots for visual-regression checks.
Built so an AI coding agent can verify its own visual output: render a UI change, screenshot it, and look at the result instead of guessing.
## Features
Six tools, exposed over MCP via [FastMCP](https://github.com/jlowin/fastmcp):
- **`capture_screen`** — screenshot the full screen or a specific display, with token-aware downscaling.
- **`capture_window_tool`** — screenshot a single window by (partial, case-insensitive) title or app name.
- **`capture_screen_region`** — screenshot an arbitrary rectangular region by coordinates.
- **`capture_webpage_tool`** — screenshot any URL (including `localhost` dev servers) via a headless Playwright browser; supports full-page capture and waiting on a CSS selector.
- **`compare_screenshots`** — diff two images and report the percentage and region of changed pixels (visual regression).
- **`list_windows`** — list all open windows as `Application | Window Title` to find a capture target.
## Requirements
- Python ≥ 3.10
- macOS (window capture uses AppleScript; screen capture needs **Screen Recording** permission)
- Dependencies: `mcp[cli]`, `Pillow`, `playwright`, `numpy`
## Install
```bash
git clone https://github.com/wonderstone843/claude-vision-mcp.git
cd claude-vision-mcp
pip install -e .
playwright install chromium # only needed for capture_webpage_tool
```
Grant Screen Recording permission to your terminal in **System Settings → Privacy & Security → Screen Recording**.
## Use with Claude Code
Register the server (stdio):
```bash
claude mcp add claude-vision -- claude-vision-mcp
```
Or add it to your MCP config manually:
```json
{
"mcpServers": {
"claude-vision": {
"command": "claude-vision-mcp"
}
}
}
```
Then ask Claude to, e.g., *"screenshot localhost:3000 and check the hero section renders,"* or *"capture the Blender window."*
## Project layout
```
claude_vision_mcp/
server.py # FastMCP server + the 6 tool definitions
capture.py # full-screen / window / region capture
windows.py # AppleScript window enumeration
browser.py # Playwright headless webpage capture
compare.py # pixel-diff comparison
```
## License
MIT — see [LICENSE](LICENSE).
Author: Joshua Penn
TDQS
A4.1/5.0
Scored across 6 tools
Disambiguation5/5
Each tool has a clearly distinct purpose: capturing different screen sources (full, region, webpage, window), comparing images, or listing windows. No overlapping functionality.
Naming Consistency4/5
Most tools follow a 'capture_' prefix pattern, but capture_webpage_tool and capture_window_tool have an inconsistent '_tool' suffix, and compare_screenshots and list_windows use different verb forms.
Tool Count5/5
6 tools is well-scoped for a vision capture server, covering core capture methods plus comparison and window listing without excess.
Completeness4/5
Covers all major capture scenarios (screen, region, window, webpage) and adds comparison and window listing. Minor gap: no tool to list displays or get screen dimensions, but all essential workflows are present.
Maintenance
ActivityInactive
ResponsivenessNo issues