markitdown-ocr-mcp
# markitdown-ocr-mcp
MCP server exposing OCR-enabled PDF → Markdown conversion to Claude Code.
MarkItDown + the official `markitdown-ocr` LLM-vision plugin, backed by a local
oMLX server running GLM-OCR (or any OpenAI-compatible vision endpoint).
## Setup
```bash
uv sync --group dev
uv tool install --editable .
claude mcp add markitdown-ocr-mcp --scope user -- ~/.local/bin/markitdown-ocr-mcp
```
oMLX must be running (`omlx start`) with a vision model loaded.
## Tools
- `inspect_pdf(path)` — per-page classification (text / scanned / mixed / blank), no OCR
- `ocr_pdf(path, pages?, out_path?, dpi?)` — hybrid conversion to Markdown: text-layer pages via MarkItDown (exact text), scanned/mixed pages rendered and OCR'd directly by the vision model; `pages` ("1-5,9") extracts only a page subset; `out_path` writes to file; `dpi` overrides OCR_DPI
- `omlx_models()` — oMLX health check + available models + resolved OCR model
## Testing
```bash
uv run pytest # unit + integration, no oMLX needed
uv run python scripts/smoke_omlx.py # live OCR smoke test (needs oMLX)
uv run python scripts/probe_mcp.py # end-to-end probe of the installed binary
```
## Config (env vars)
| Var | Default |
|---|---|
| `OMLX_URL` | `http://127.0.0.1:8080/v1` |
| `OMLX_API_KEY` | auto-read from `~/.omlx/settings.json` |
| `OMLX_OCR_MODEL` | auto-discovered from `/v1/models` — prefers GLM-OCR (small prefill footprint) |
| `OMLX_OCR_PROMPT` | `OCR:` |
| `OCR_DPI` | `150` (render DPI for direct OCR) |
| `OCR_MAX_LONG_SIDE` | `1300` (pixel cap keeping oMLX's prefill inside its memory guard) |
TDQS
Scored across 3 tools
Each tool has a clearly distinct role: inspect_pdf classifies pages, ocr_pdf performs conversion, and omlx_models checks OCR service health. There is no functional overlap or ambiguity between them.
All names use lowercase snake_case and follow a pattern of verb_noun (inspect_pdf, ocr_pdf). omlx_models breaks the verb_noun pattern but is still readable and consistent in style, making it a minor deviation.
With only three tools, the server is tightly scoped to the PDF-to-Markdown OCR workflow. Each tool serves a necessary purpose without redundancy or bloat.
The server supports the full workflow: inspecting PDFs to identify OCR needs, converting with optional page selection, and verifying OCR model availability. No critical missing operations are apparent for its stated purpose.