browser-ocr-mcp
# browser-ocr-mcp
> MCP server for local OCR via Tesseract.js. Extract text from images and browser screenshots **without sending image data to the LLM**.
## Why
When browsing with Playwright MCP, you often encounter pages where text is embedded in images. Normally you'd send the screenshot to the LLM — slow, bandwidth-heavy. This server runs Tesseract.js locally so the image never leaves your machine.
## Tools
### `ocr_image`
Extract text from an image file or URL. Zero browser setup required.
**Input:**
- `path` (string, optional) — Absolute path to a local image file
- `url` (string, optional) — HTTP URL to download the image from
**Output:**
```json
{
"text": "The quick brown fox jumps over the lazy dog",
"confidence": 85,
"wordCount": 9
}
```
**Typical workflow with Playwright MCP:**
```
1. browser_take_screenshot({ filename: "page.png" })
2. ocr_image({ path: "/home/user/page.png" })
3. → text returned, zero image data sent to LLM
```
### `browser_ocr`
Take a screenshot of the current browser page via CDP and OCR it — all in one call. Requires a shared Chromium instance.
**Input:**
- `fullPage` (boolean, optional) — Capture full scrollable page. Default: `false`
**Setup (one-time):**
```bash
# Terminal 1: Launch Chromium with debugging port
chromium --remote-debugging-port=9222
# Terminal 2: Start Playwright MCP connected to it
npx @playwright/mcp --cdp-endpoint http://localhost:9222
# Terminal 3: Start OCR MCP server
node server.js --cdp-endpoint http://localhost:9222
```
## Install
```bash
git clone https://github.com/Ismapik/browser-ocr-mcp.git
cd browser-ocr-mcp
npm install
```
The first time you run OCR, Tesseract.js downloads English language data (~12 MB). Subsequent runs use the cached data.
## MCP Client Config
Add to your MCP client configuration (e.g., `mcp.json` or Claude Desktop config):
```json
{
"mcpServers": {
"browser-ocr": {
"command": "node",
"args": ["/path/to/browser-ocr-mcp/server.js"],
"env": {
"CDP_ENDPOINT": "http://localhost:9222"
}
}
}
}
```
The `CDP_ENDPOINT` env var is only needed if you use the `browser_ocr` tool. The `ocr_image` tool works without it.
## Test
```bash
npm test
```
Runs a test suite that:
1. Starts the MCP server
2. Calls `ocr_image` on a sample image
3. Verifies text extraction, confidence, and word count
4. Tests error handling for missing/invalid inputs
## Requirements
- Node.js ≥ 18
- Chromium (only for `browser_ocr` tool)
- `playwright-core` (optional, only for `browser_ocr`)
## License
MIT
TDQS
Scored across 2 tools
Both tools perform OCR but on clearly distinct inputs: one captures the current browser page via screenshot, the other processes an image file or URL. No ambiguity in their purposes.
Both tools use a consistent snake_case naming pattern with a source prefix (browser_, ocr_) followed by the action (ocr, image). The pattern is straightforward and predictable.
With only 2 tools, the server is lean but covers the core OCR use cases. While slightly below the typical 3-15 range, it feels appropriately scoped for a focused OCR utility.
The two tools cover the primary OCR scenarios: extracting text from the live browser page and from external images. Minor gaps like language configuration or PDF support exist but are not essential for basic usage.