Skip to main content
Glama
README.md
# markitdown-ocr-mcp

MCP server exposing OCR-enabled PDF → Markdown conversion to Claude Code.

MarkItDown + the official `markitdown-ocr` LLM-vision plugin, backed by a local
oMLX server running GLM-OCR (or any OpenAI-compatible vision endpoint).

## Setup

```bash
uv sync --group dev
uv tool install --editable .
claude mcp add markitdown-ocr-mcp --scope user -- ~/.local/bin/markitdown-ocr-mcp
```

oMLX must be running (`omlx start`) with a vision model loaded.

## Tools

- `inspect_pdf(path)` — per-page classification (text / scanned / mixed / blank), no OCR
- `ocr_pdf(path, pages?, out_path?, dpi?)` — hybrid conversion to Markdown: text-layer pages via MarkItDown (exact text), scanned/mixed pages rendered and OCR'd directly by the vision model; `pages` ("1-5,9") extracts only a page subset; `out_path` writes to file; `dpi` overrides OCR_DPI
- `omlx_models()` — oMLX health check + available models + resolved OCR model

## Testing

```bash
uv run pytest                        # unit + integration, no oMLX needed
uv run python scripts/smoke_omlx.py  # live OCR smoke test (needs oMLX)
uv run python scripts/probe_mcp.py   # end-to-end probe of the installed binary
```

## Config (env vars)

| Var | Default |
|---|---|
| `OMLX_URL` | `http://127.0.0.1:8080/v1` |
| `OMLX_API_KEY` | auto-read from `~/.omlx/settings.json` |
| `OMLX_OCR_MODEL` | auto-discovered from `/v1/models` — prefers GLM-OCR (small prefill footprint) |
| `OMLX_OCR_PROMPT` | `OCR:` |
| `OCR_DPI` | `150` (render DPI for direct OCR) |
| `OCR_MAX_LONG_SIDE` | `1300` (pixel cap keeping oMLX's prefill inside its memory guard) |

TDQS

A4.4/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a clearly distinct role: inspect_pdf classifies pages, ocr_pdf performs conversion, and omlx_models checks OCR service health. There is no functional overlap or ambiguity between them.

Naming Consistency4/5

All names use lowercase snake_case and follow a pattern of verb_noun (inspect_pdf, ocr_pdf). omlx_models breaks the verb_noun pattern but is still readable and consistent in style, making it a minor deviation.

Tool Count5/5

With only three tools, the server is tightly scoped to the PDF-to-Markdown OCR workflow. Each tool serves a necessary purpose without redundancy or bloat.

Completeness5/5

The server supports the full workflow: inspecting PDFs to identify OCR needs, converting with optional page selection, and verifying OCR model availability. No critical missing operations are apparent for its stated purpose.

Maintenance

ActivityMaintained
ResponsivenessNo issues