Skip to main content
Glama
lfzds4399-cpu

claude-screen-mcp

README.md
# claude-screen-mcp

An MCP server that exposes read-only screen state. Capture, OCR, and change
detection only — no input control.

[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)
[![Node](https://img.shields.io/badge/node-%3E%3D20-brightgreen.svg)](package.json)
[![MCP](https://img.shields.io/badge/MCP-1.29-blue.svg)](https://modelcontextprotocol.io)
[![CI](https://github.com/lfzds4399-cpu/claude-screen-mcp/actions/workflows/ci.yml/badge.svg)](https://github.com/lfzds4399-cpu/claude-screen-mcp/actions/workflows/ci.yml)

## Quick start

```bash
git clone https://github.com/lfzds4399-cpu/claude-screen-mcp
cd claude-screen-mcp
npm install
npm run build

claude mcp add screen -- node "$(pwd)/dist/index.js"
```

Restart the MCP host after registration.

## Tools

| Tool | Purpose |
|---|---|
| `screenshot` | Capture a full display and resize the image result. |
| `screenshot_region` | Capture a rectangular region. |
| `list_displays` | Enumerate connected displays. |
| `list_windows` | List visible top-level windows with optional title filter. |
| `read_screen_text` | Run OCR on the full display or a region. |
| `find_text_on_screen` | Search OCR text and return matching bounding boxes. |
| `screenshot_if_changed` | Capture only when perceptual-hash distance exceeds a threshold. |
| `get_screen_diff` | Return hash-distance diagnostics without an image. |
| `wait_for_change` | Poll until the screen changes or a timeout elapses. |
| `record_screen` | Sample a short interval and return deduplicated keyframes. |

## Platform support

Windows 10+ is the primary target and is exercised on every release. macOS 11+
and Linux (X11 and Wayland) pass CI and smoke tests but are not regularly
exercised. Window enumeration on macOS and Linux requires platform tooling;
multi-monitor display enumeration is supported on Windows.

## Security and privacy

All processing is local. No screenshot, OCR text, or telemetry leaves the
machine; the only network call is the initial Tesseract language data download.

OCR output is untrusted input. Text rendered on screen may attempt to influence
the model. Treat output as user-supplied data and avoid auto-executing commands
derived from it. Scope `read_screen_text` to a window when full-desktop capture
is not required.

## Configuration

| Variable | Default | Purpose |
|---|---|---|
| `SCREEN_MCP_LOG_LEVEL` | `info` | `debug`, `info`, `warn`, or `error`. |
| `SCREEN_MCP_OCR_LANGS` | `eng+chi_sim` | Tesseract language list (allowlist enforced). |

The first OCR call downloads language data (~10MB per language); subsequent calls reuse the local cache.

## Development

```bash
npm install
npm run build
npm test
node tests/e2e-wire.mjs
```

## Roadmap

- `screenshot_window(title)` for direct single-window capture.
- Improved multi-display enumeration on macOS and Linux.

## License — MIT, see [LICENSE](LICENSE).