Skip to main content
Glama
README.md
<div align="center">

# πŸ“„ MCP PDF

<img src="https://img.shields.io/badge/MCP-PDF%20Tools-red?style=for-the-badge&logo=adobe-acrobat-reader" alt="MCP PDF">

**A FastMCP server for PDF processing**

*52 tools for text extraction, OCR, tables, forms, XFA, annotations, markdown↔PDF, and more*

[![Python 3.11+](https://img.shields.io/badge/python-3.11+-blue.svg?style=flat-square)](https://www.python.org/downloads/)
[![FastMCP](https://img.shields.io/badge/FastMCP-2.0+-green.svg?style=flat-square)](https://github.com/jlowin/fastmcp)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg?style=flat-square)](https://opensource.org/licenses/MIT)
[![PyPI](https://img.shields.io/pypi/v/mcp-pdf?style=flat-square)](https://pypi.org/project/mcp-pdf/)

**Works great with [MCP Office Tools](https://git.supported.systems/MCP/mcp-office-tools)**

</div>

---

## What It Does

MCP PDF extracts content from PDFs using multiple libraries with automatic fallbacks. If one method fails, it tries another.

**Core capabilities:**
- **Text extraction** via PyMuPDF, pdfplumber, or pypdf (auto-fallback)
- **Table extraction** via Camelot, pdfplumber, or Tabula (auto-fallback)
- **OCR** for scanned documents via Tesseract
- **Form handling** - extract, fill, and create PDF forms
- **Document assembly** - merge, split, reorder pages
- **Annotations** - sticky notes, highlights, stamps
- **Vector graphics** - extract to SVG for schematics and technical drawings
- **Format conversion** - PDF ↔ Markdown (PDFβ†’MD via PyMuPDF, MDβ†’PDF via pandoc)
- **XFA forms** - Schema extraction for dynamic Adobe LiveCycle forms that no open-source library can render

---

## Quick Start

```bash
# Run from PyPI (one-shot, no permanent install)
uvx mcp-pdf

# Add to Claude Code β€” note the `--` separator before uvx
claude mcp add pdf-tools -- uvx mcp-pdf

# Include the markdown_to_pdf tool (requires pandoc on host)
claude mcp add pdf-tools -- uvx --from "mcp-pdf[markdown]" mcp-pdf
```

> `uvx` caches tool installs aggressively. After upgrading to a new release, force a refresh with `uvx --refresh mcp-pdf` (or `uvx --refresh --from "mcp-pdf[markdown]" mcp-pdf` if you're using extras).

<details>
<summary><b>Development Installation</b></summary>

```bash
git clone https://github.com/rsp2k/mcp-pdf
cd mcp-pdf
uv sync

# System dependencies (Ubuntu/Debian)
sudo apt-get install tesseract-ocr tesseract-ocr-eng poppler-utils ghostscript

# For markdown_to_pdf β€” pick one PDF-engine route:
sudo apt-get install pandoc tectonic                                          # recommended (small)
# or:  sudo apt-get install pandoc texlive-xetex texlive-latex-extra          # full TeX
# or:  sudo apt-get install pandoc && pip install weasyprint                  # skip TeX

# Verify
uv run python examples/verify_installation.py
```

</details>

---

## Tools

### Content Extraction

| Tool | What it does |
|------|-------------|
| `extract_text` | Pull text from PDF pages with automatic chunking for large files |
| `extract_tables` | Extract tables to JSON, CSV, or Markdown |
| `extract_images` | Extract embedded images |
| `extract_links` | Get all hyperlinks with page filtering |
| `ocr_pdf` | OCR scanned documents using Tesseract |
| `extract_vector_graphics` | Export vector graphics to SVG (schematics, charts, drawings) |

### Format Conversion

| Tool | What it does |
|------|-------------|
| `pdf_to_markdown` | Convert PDF to markdown preserving structure; extracts images and SVG vectors to disk |
| `markdown_to_pdf` | Convert `.md` files (or inline text) to PDF via pandoc with auto-detected engine |

**`markdown_to_pdf` requires:** `pip install mcp-pdf[markdown]` plus the pandoc binary and at least one PDF engine (`xelatex`, `pdflatex`, `tectonic`, `weasyprint`, or `wkhtmltopdf`) on PATH. The tool auto-detects what's available and uses the highest-quality one. Pass `pdf_engine=` to override or `extra_args=` for raw pandoc options.

### Document Analysis

| Tool | What it does |
|------|-------------|
| `extract_metadata` | Get title, author, creation date, page count, etc. |
| `get_document_structure` | Extract table of contents and bookmarks |
| `analyze_layout` | Detect columns, headers, footers |
| `is_scanned_pdf` | Check if PDF needs OCR |
| `compare_pdfs` | Diff two PDFs by text, structure, or metadata |
| `analyze_pdf_health` | Check for corruption, optimization opportunities |
| `analyze_pdf_security` | Report encryption, permissions, signatures |

### Forms

| Tool | What it does |
|------|-------------|
| `extract_form_data` | Get form field names and values (AcroForm) |
| `fill_form_pdf` | Fill form fields from JSON |
| `create_form_pdf` | Create new forms with text fields, checkboxes, dropdowns |
| `add_form_fields` | Add fields to existing PDFs |
| `add_date_field` | Add a date field with format validation |
| `add_radio_group` | Add a radio button group with mutual exclusion |
| `add_textarea_field` | Add a multi-line text area with word limits |
| `add_field_validation` | Add validation rules to existing form fields |
| `validate_form_data` | Check data against field rules and constraints before filling |

Field types are reported in a **portable vocabulary** (`text/checkbox/radio/dropdown/date/signature` plus `button/unknown`) shared between the AcroForm and XFA tools, so callers don't have to learn two models. `extract_form_data` also returns `field_type_raw` with the unmerged AcroForm widget type, since `listbox` and `combobox` both report as `dropdown` but differ on free-text entry and multi-select.

### XFA Forms (Dynamic Adobe LiveCycle)

Real-estate forms, mortgage forms, government forms β€” many are **dynamic XFA**, where the layout + fields live in an XFA program that only Adobe's runtime can render. Every open-source PDF library (PyMuPDF, pdfium, MuPDF, pikepdf) only sees the static "Open in Adobe Reader" placeholder page. These tools recover the form *schema* instead.

| Tool | What it does |
|------|-------------|
| `is_xfa_pdf` | Detect XFA and classify as dynamic / static. Use for branching before extract_form_data or convert_to_images |
| `extract_xfa_fields` | Parse the XFA template for field names, captions, UI types. Splits into shared (cross-form canonical), positional (opaque codes), other, and plumbing (producer internals, dropped) |

**Check `detection_failed` before trusting `is_xfa`.** Detection has three outcomes: `True` with an `xfa_type`, `False` for a readable non-XFA PDF, and `None` with `detection_failed=True` when the file could not be read. A truncated or corrupt form package lands in the third case and gets told so, rather than being reported as a file that simply has no XFA.

`extract_xfa_fields` defaults to the **zipForm producer profile** (Lone Wolf / zipForm Plus, the most common XFA producer in the wild), matched case-insensitively. Pass `profile="generic"` plus `extra_plumbing_patterns` / `extra_positional_patterns` for other producers; an unrecognized profile name is an error rather than a silent fallback. Plumbing patterns are substring searches while positional patterns are anchored at the start.

Every response carries `success`, `is_xfa` and `xfa_type`, including failures. Fields land in four categories, not three: `shared`, `positional`, `other` (the default bucket, which holds the majority on non-zipForm producers) and `plumbing_fields_dropped`. `canonical_collisions` flags distinct XFA names that canonicalize to the same key, which matters because that key is the cross-form join. The `original` XFA name is on every field as the round-trip key for filling; `canonical_name` appears only on shared fields. `canonical_separator` chooses `_` (snake, default), `.` (dotted) or `-` (kebab). `include_design_time_bbox=True` opts into best-effort geometry, page-relative with a top-left origin, and not authoritative for dynamic XFA since subforms reflow at render time.

### Permit Forms (Coordinate-Based)

For scanned PDFs or forms without interactive fields. Draws text at (x, y) coordinates.

| Tool | What it does |
|------|-------------|
| `fill_permit_form` | Fill any PDF by drawing at coordinates (works with scanned forms) |
| `get_field_schema` | Get field definitions for validation or UI generation |
| `validate_permit_form_data` | Check data against field schema before filling |
| `preview_field_positions` | Generate PDF showing field boundaries (debugging) |
| `insert_attachment_pages` | Insert image/text pages with "See page X" references |

**Requires:** `pip install mcp-pdf[forms]` (adds reportlab dependency)

### Document Assembly

| Tool | What it does |
|------|-------------|
| `merge_pdfs` | Combine multiple PDFs with bookmark preservation |
| `merge_pdfs_advanced` | Merge with page numbering, generated TOC, and per-file page ranges |
| `split_pdf` | Split into separate documents |
| `split_pdf_by_pages` | Split by page ranges |
| `split_pdf_by_bookmarks` | Split at chapter/section boundaries |
| `reorder_pdf_pages` | Rearrange pages in custom order |
| `rotate_pages` | Rotate specific pages by 90, 180 or 270 degrees |

### Structure Detection

Chapter-aware analysis, for long documents where you want to work a section at a time instead of paging through blindly.

| Tool | What it does |
|------|-------------|
| `detect_structure` | Find headings via bookmarks, font-size heuristics and numbering patterns |
| `split_pdf_by_structure` | Auto-split into per-chapter directories with markdown and images |
| `batch_extract` | Process several page ranges in one call, replacing dozens of individual calls |

`detect_structure` writes the full structure to JSON and returns a compact summary plus the path (roughly 1k tokens against ~20k inline). Pass `inline=True` when you want the whole thing in the response.

### Content Analysis

| Tool | What it does |
|------|-------------|
| `classify_content` | Identify document type (invoice, contract, report) and structure |
| `summarize_content` | Generate a summary and key insights |
| `extract_charts` | Extract and analyze charts, diagrams and visual elements |
| `detect_watermarks` | Detect and analyze watermarks |

### PDF Utilities

| Tool | What it does |
|------|-------------|
| `convert_to_images` | Render pages to PNG or JPEG at a chosen DPI |
| `optimize_pdf` | Reduce file size and improve load performance |
| `repair_pdf` | Attempt recovery of a corrupted or damaged PDF |

### Server Introspection

Not PDF tools, so they sit outside the count above. Useful when you want to know what a given install can actually do.

| Tool | What it does |
|------|-------------|
| `list_capabilities` | Enumerate what this server build can do |
| `server_info` | Report version, registered mixins and configuration |

### Annotations

| Tool | What it does |
|------|-------------|
| `add_sticky_notes` | Add comment annotations |
| `add_highlights` | Highlight text regions |
| `add_stamps` | Add Approved/Draft/Confidential stamps |
| `extract_all_annotations` | Export annotations to JSON |

---

## How Fallbacks Work

The server tries multiple libraries for each operation:

**Text extraction:**
1. PyMuPDF (fastest)
2. pdfplumber (better for complex layouts)
3. pypdf (most compatible)

**Table extraction:**
1. Camelot (best accuracy, requires Ghostscript)
2. pdfplumber (no dependencies)
3. Tabula (requires Java)

If a PDF fails with one library, the next is tried automatically. An engine that raises *and* one that succeeds but finds no text both fall through, so a page PyMuPDF silently returns nothing for still gets tried by pdfplumber and pypdf.

The response tells you what happened: `method_used` names the engine that produced the text, and `methods_attempted` lists the ones skipped and why. When all three come back empty you get `method_used: "none"` plus an `extraction_warning`, rather than the emptiness being attributed to whichever engine happened to run last. Naming an engine explicitly (`method="pdfplumber"`) disables the cascade, and a failure then raises instead of falling through.

---

## Token Management

Large PDFs can overflow MCP response limits. The server handles this:

- **Automatic chunking** splits large documents into page groups
- **Table row limits** prevent huge tables from blowing up responses
- **Summary mode** returns structure without full content

```python
# Get first 10 pages
result = await extract_text("huge.pdf", pages="1-10")

# Limit table rows
tables = await extract_tables("data.pdf", max_rows_per_table=50)

# Structure only
tables = await extract_tables("data.pdf", summary_only=True)
```

---

## URL Processing

PDFs can be fetched directly from HTTPS URLs:

```python
result = await extract_text("https://example.com/report.pdf")
```

Files are cached locally for subsequent operations.

---

## System Dependencies

Some features require system packages:

| Feature | Dependency |
|---------|-----------|
| OCR | `tesseract-ocr` |
| Camelot tables | `ghostscript` |
| Tabula tables | `default-jre-headless` |
| PDF to images | `poppler-utils` |
| `markdown_to_pdf` | `pandoc` + one of: `tectonic`, `texlive-xetex` (+ `texlive-latex-extra`), `weasyprint`, `wkhtmltopdf` |

### Picking a PDF engine for `markdown_to_pdf`

Pandoc takes markdown β†’ HTML or LaTeX β†’ PDF. The LaTeX path produces the most polished output but needs a TeX install. Trade-offs:

| Engine | Disk size | Notes |
|--------|----------|-------|
| **`tectonic`** | ~30 MB | **Recommended for new installs.** Single static binary. Downloads LaTeX packages on demand β€” no upfront mass-install. |
| `xelatex` + `texlive-latex-extra` | ~500 MB | Best output once installed. Use if you already run TeX. The `-extra` package matters: pandoc's default template needs `lastpage`, `xcolor`, `framed`, `fancyhdr`, etc. β€” all of which live there, **not** in `texlive-xetex`. |
| `xelatex` alone (just `texlive-xetex`) | ~200 MB | **Often breaks.** Expect `! LaTeX Error: File 'X.sty' not found` on real docs. |
| `weasyprint` | ~40 MB | Pure-Python (`pip install weasyprint`) + cairo/pango system libs. HTML/CSS path β€” no LaTeX. Good for simple docs; weaker on math, footnotes, citations. |
| `wkhtmltopdf` | ~40 MB | Older HTML-to-PDF tool. Adequate but less actively maintained. |

**Ubuntu/Debian:**
```bash
sudo apt-get install tesseract-ocr tesseract-ocr-eng poppler-utils ghostscript default-jre-headless

# For markdown_to_pdf β€” pick one engine route:

# Option A β€” tectonic (smallest, downloads packages on demand)
sudo apt-get install pandoc
# tectonic isn't in apt β€” install via cargo or download static binary:
#   https://tectonic-typesetting.github.io/en-US/install.html

# Option B β€” full TeX (best quality, large download)
sudo apt-get install pandoc texlive-xetex texlive-latex-extra texlive-fonts-extra

# Option C β€” weasyprint (skip TeX entirely)
sudo apt-get install pandoc
pip install weasyprint
```

**Arch Linux:**
```bash
sudo pacman -S tesseract tesseract-data-eng poppler ghostscript jre-openjdk-headless

# For markdown_to_pdf β€” pick one engine route:

# Option A β€” tectonic (recommended for new installs, in official repo)
sudo pacman -S pandoc tectonic

# Option B β€” full TeX (best output, ~500 MB)
sudo pacman -S pandoc texlive-xetex texlive-latexextra texlive-fontsextra

# Option C β€” weasyprint (skip TeX)
sudo pacman -S pandoc
pip install weasyprint   # or: uv pip install weasyprint

# Option D β€” wkhtmltopdf (from AUR)
yay -S wkhtmltopdf-static
```

**macOS (Homebrew):**
```bash
brew install tesseract poppler ghostscript

# For markdown_to_pdf β€” pick one engine route:

# Option A β€” tectonic (recommended)
brew install pandoc tectonic

# Option B β€” full TeX (mactex-no-gui includes the latex-extra equivalent)
brew install pandoc
brew install --cask mactex-no-gui

# Option C β€” weasyprint
brew install pandoc weasyprint
```

## Optional Extras

The base install stays lean. Heavy or niche dependencies are gated behind extras:

| Extra | Adds | When to install |
|-------|------|----------------|
| `mcp-pdf[forms]` | `reportlab` | Form creation tools (`create_form_pdf`, permit forms) |
| `mcp-pdf[tables]` | `camelot-py`, `tabula-py` | Higher-accuracy table extraction (also needs Java + Ghostscript) |
| `mcp-pdf[markdown]` | `pypandoc` | `markdown_to_pdf` tool (also needs pandoc binary) |
| `mcp-pdf[all]` | All of the above | Want everything |

---

## Configuration

All optional. Defaults are tuned for local stdio use (Claude Desktop, Claude Code), which is the common case.

| Variable | Default | Purpose |
|----------|---------|---------|
| `PDF_TEMP_DIR` | `/tmp/mcp-pdf-processing` | Working directory for intermediate files |
| `MCP_PDF_MAX_SIZE` | *no limit* | Max input PDF size in **MB**. Set `0` or leave unset to disable |
| `ALLOWED_DOMAINS` | *all* | Comma-separated host allowlist for fetching PDFs over HTTPS |
| `DEBUG` | `false` | Set `true` for verbose logging |
| `TESSDATA_PREFIX` | *system* | Tesseract language data location (read by Tesseract, not by this server) |

### Output-path restriction

`MCP_PDF_ALLOWED_PATHS` takes a colon-separated list of directories that tools may write into. **It is not enforced in stdio mode**, which is the default, on the reasoning that a local server writing to the user's own filesystem does not need to be fenced off from it. Enforcement switches on when either of these is set:

| Variable | Effect |
|----------|--------|
| `MCP_TRANSPORT=http` | Treat as network-exposed; enforce `MCP_PDF_ALLOWED_PATHS` |
| `MCP_PUBLIC_MODE` | Any non-empty value does the same |

If you expose this server over HTTP, set both `MCP_TRANSPORT=http` and `MCP_PDF_ALLOWED_PATHS`. Setting the allowlist alone has no effect in stdio mode. And note the framing in [CLAUDE.md](CLAUDE.md): application-level path checks are a speed bump, not a boundary. Real isolation comes from running as an unprivileged user in a container with the filesystem it actually needs and nothing more.

### XFA limits

Bound the XFA parser against hostile or merely enormous templates. expat caps entity amplification relative to input size, so bounding the input is what bounds the expansion.

| Variable | Default | Purpose |
|----------|---------|---------|
| `MCP_PDF_MAX_XFA_TEMPLATE_BYTES` | `16777216` (16 MB) | Reject XFA templates larger than this |
| `MCP_PDF_MAX_XFA_DEPTH` | `100` | Max `<subform>` nesting before giving up |
| `MCP_PDF_MAX_XFA_INLINE_FIELDS` | `5000` | Fields serialized into one response before truncating |

---

## Development

```bash
# Run tests
uv run pytest

# With coverage
uv run pytest --cov=mcp_pdf

# Format
uv run black src/ tests/

# Lint
uv run ruff check src/ tests/
```

---

## Versioning

CalVer, `YYYY.MM.DD`, with a PEP 440 post-release segment for same-day fixes (`2026.09.21.1`). This package is a thin layer over PyMuPDF, pdfplumber, pypdf, Camelot, Tabula, Tesseract, pandoc and several system binaries, and most surprising behavior traces back to one of those. When a PDF misbehaves, the useful question is "when was this last tested against those?", which a date answers.

Releases before `2026.09.21` used semver. PEP 440 compares the release tuple numerically, so `(2026, 9, 21)` sorts after `(2, 3, 1)` and `pip install --upgrade` works correctly across the switch. See [CHANGELOG.md](CHANGELOG.md).

## License

MIT

</div>