Skip to main content
Glama

📄 MCP PDF

A FastMCP server for PDF processing

52 tools for text extraction, OCR, tables, forms, XFA, annotations, markdown↔PDF, and more

Python 3.11+ FastMCP License: MIT PyPI

Works great with MCP Office Tools


What It Does

MCP PDF extracts content from PDFs using multiple libraries with automatic fallbacks. If one method fails, it tries another.

Core capabilities:

  • Text extraction via PyMuPDF, pdfplumber, or pypdf (auto-fallback)

  • Table extraction via Camelot, pdfplumber, or Tabula (auto-fallback)

  • OCR for scanned documents via Tesseract

  • Form handling - extract, fill, and create PDF forms

  • Document assembly - merge, split, reorder pages

  • Annotations - sticky notes, highlights, stamps

  • Vector graphics - extract to SVG for schematics and technical drawings

  • Format conversion - PDF ↔ Markdown (PDF→MD via PyMuPDF, MD→PDF via pandoc)

  • XFA forms - Schema extraction for dynamic Adobe LiveCycle forms that no open-source library can render


Related MCP server: PDF MCP Flow

Quick Start

# Run from PyPI (one-shot, no permanent install)
uvx mcp-pdf

# Add to Claude Code — note the `--` separator before uvx
claude mcp add pdf-tools -- uvx mcp-pdf

# Include the markdown_to_pdf tool (requires pandoc on host)
claude mcp add pdf-tools -- uvx --from "mcp-pdf[markdown]" mcp-pdf

uvx caches tool installs aggressively. After upgrading to a new release, force a refresh with uvx --refresh mcp-pdf (or uvx --refresh --from "mcp-pdf[markdown]" mcp-pdf if you're using extras).

git clone https://github.com/rsp2k/mcp-pdf
cd mcp-pdf
uv sync

# System dependencies (Ubuntu/Debian)
sudo apt-get install tesseract-ocr tesseract-ocr-eng poppler-utils ghostscript

# For markdown_to_pdf — pick one PDF-engine route:
sudo apt-get install pandoc tectonic                                          # recommended (small)
# or:  sudo apt-get install pandoc texlive-xetex texlive-latex-extra          # full TeX
# or:  sudo apt-get install pandoc && pip install weasyprint                  # skip TeX

# Verify
uv run python examples/verify_installation.py

Tools

Content Extraction

Tool

What it does

extract_text

Pull text from PDF pages with automatic chunking for large files

extract_tables

Extract tables to JSON, CSV, or Markdown

extract_images

Extract embedded images

extract_links

Get all hyperlinks with page filtering

ocr_pdf

OCR scanned documents using Tesseract

extract_vector_graphics

Export vector graphics to SVG (schematics, charts, drawings)

Format Conversion

Tool

What it does

pdf_to_markdown

Convert PDF to markdown preserving structure; extracts images and SVG vectors to disk

markdown_to_pdf

Convert .md files (or inline text) to PDF via pandoc with auto-detected engine

markdown_to_pdf requires: pip install mcp-pdf[markdown] plus the pandoc binary and at least one PDF engine (xelatex, pdflatex, tectonic, weasyprint, or wkhtmltopdf) on PATH. The tool auto-detects what's available and uses the highest-quality one. Pass pdf_engine= to override or extra_args= for raw pandoc options.

Document Analysis

Tool

What it does

extract_metadata

Get title, author, creation date, page count, etc.

get_document_structure

Extract table of contents and bookmarks

analyze_layout

Detect columns, headers, footers

is_scanned_pdf

Check if PDF needs OCR

compare_pdfs

Diff two PDFs by text, structure, or metadata

analyze_pdf_health

Check for corruption, optimization opportunities

analyze_pdf_security

Report encryption, permissions, signatures

Forms

Tool

What it does

extract_form_data

Get form field names and values (AcroForm)

fill_form_pdf

Fill form fields from JSON

create_form_pdf

Create new forms with text fields, checkboxes, dropdowns

add_form_fields

Add fields to existing PDFs

add_date_field

Add a date field with format validation

add_radio_group

Add a radio button group with mutual exclusion

add_textarea_field

Add a multi-line text area with word limits

add_field_validation

Add validation rules to existing form fields

validate_form_data

Check data against field rules and constraints before filling

Field types are reported in a portable vocabulary (text/checkbox/radio/dropdown/date/signature plus button/unknown) shared between the AcroForm and XFA tools, so callers don't have to learn two models. extract_form_data also returns field_type_raw with the unmerged AcroForm widget type, since listbox and combobox both report as dropdown but differ on free-text entry and multi-select.

XFA Forms (Dynamic Adobe LiveCycle)

Real-estate forms, mortgage forms, government forms — many are dynamic XFA, where the layout + fields live in an XFA program that only Adobe's runtime can render. Every open-source PDF library (PyMuPDF, pdfium, MuPDF, pikepdf) only sees the static "Open in Adobe Reader" placeholder page. These tools recover the form schema instead.

Tool

What it does

is_xfa_pdf

Detect XFA and classify as dynamic / static. Use for branching before extract_form_data or convert_to_images

extract_xfa_fields

Parse the XFA template for field names, captions, UI types. Splits into shared (cross-form canonical), positional (opaque codes), other, and plumbing (producer internals, dropped)

Check detection_failed before trusting is_xfa. Detection has three outcomes: True with an xfa_type, False for a readable non-XFA PDF, and None with detection_failed=True when the file could not be read. A truncated or corrupt form package lands in the third case and gets told so, rather than being reported as a file that simply has no XFA.

extract_xfa_fields defaults to the zipForm producer profile (Lone Wolf / zipForm Plus, the most common XFA producer in the wild), matched case-insensitively. Pass profile="generic" plus extra_plumbing_patterns / extra_positional_patterns for other producers; an unrecognized profile name is an error rather than a silent fallback. Plumbing patterns are substring searches while positional patterns are anchored at the start.

Every response carries success, is_xfa and xfa_type, including failures. Fields land in four categories, not three: shared, positional, other (the default bucket, which holds the majority on non-zipForm producers) and plumbing_fields_dropped. canonical_collisions flags distinct XFA names that canonicalize to the same key, which matters because that key is the cross-form join. The original XFA name is on every field as the round-trip key for filling; canonical_name appears only on shared fields. canonical_separator chooses _ (snake, default), . (dotted) or - (kebab). include_design_time_bbox=True opts into best-effort geometry, page-relative with a top-left origin, and not authoritative for dynamic XFA since subforms reflow at render time.

Permit Forms (Coordinate-Based)

For scanned PDFs or forms without interactive fields. Draws text at (x, y) coordinates.

Tool

What it does

fill_permit_form

Fill any PDF by drawing at coordinates (works with scanned forms)

get_field_schema

Get field definitions for validation or UI generation

validate_permit_form_data

Check data against field schema before filling

preview_field_positions

Generate PDF showing field boundaries (debugging)

insert_attachment_pages

Insert image/text pages with "See page X" references

Requires: pip install mcp-pdf[forms] (adds reportlab dependency)

Document Assembly

Tool

What it does

merge_pdfs

Combine multiple PDFs with bookmark preservation

merge_pdfs_advanced

Merge with page numbering, generated TOC, and per-file page ranges

split_pdf

Split into separate documents

split_pdf_by_pages

Split by page ranges

split_pdf_by_bookmarks

Split at chapter/section boundaries

reorder_pdf_pages

Rearrange pages in custom order

rotate_pages

Rotate specific pages by 90, 180 or 270 degrees

Structure Detection

Chapter-aware analysis, for long documents where you want to work a section at a time instead of paging through blindly.

Tool

What it does

detect_structure

Find headings via bookmarks, font-size heuristics and numbering patterns

split_pdf_by_structure

Auto-split into per-chapter directories with markdown and images

batch_extract

Process several page ranges in one call, replacing dozens of individual calls

detect_structure writes the full structure to JSON and returns a compact summary plus the path (roughly 1k tokens against ~20k inline). Pass inline=True when you want the whole thing in the response.

Content Analysis

Tool

What it does

classify_content

Identify document type (invoice, contract, report) and structure

summarize_content

Generate a summary and key insights

extract_charts

Extract and analyze charts, diagrams and visual elements

detect_watermarks

Detect and analyze watermarks

PDF Utilities

Tool

What it does

convert_to_images

Render pages to PNG or JPEG at a chosen DPI

optimize_pdf

Reduce file size and improve load performance

repair_pdf

Attempt recovery of a corrupted or damaged PDF

Server Introspection

Not PDF tools, so they sit outside the count above. Useful when you want to know what a given install can actually do.

Tool

What it does

list_capabilities

Enumerate what this server build can do

server_info

Report version, registered mixins and configuration

Annotations

Tool

What it does

add_sticky_notes

Add comment annotations

add_highlights

Highlight text regions

add_stamps

Add Approved/Draft/Confidential stamps

extract_all_annotations

Export annotations to JSON


How Fallbacks Work

The server tries multiple libraries for each operation:

Text extraction:

  1. PyMuPDF (fastest)

  2. pdfplumber (better for complex layouts)

  3. pypdf (most compatible)

Table extraction:

  1. Camelot (best accuracy, requires Ghostscript)

  2. pdfplumber (no dependencies)

  3. Tabula (requires Java)

If a PDF fails with one library, the next is tried automatically. An engine that raises and one that succeeds but finds no text both fall through, so a page PyMuPDF silently returns nothing for still gets tried by pdfplumber and pypdf.

The response tells you what happened: method_used names the engine that produced the text, and methods_attempted lists the ones skipped and why. When all three come back empty you get method_used: "none" plus an extraction_warning, rather than the emptiness being attributed to whichever engine happened to run last. Naming an engine explicitly (method="pdfplumber") disables the cascade, and a failure then raises instead of falling through.


Token Management

Large PDFs can overflow MCP response limits. The server handles this:

  • Automatic chunking splits large documents into page groups

  • Table row limits prevent huge tables from blowing up responses

  • Summary mode returns structure without full content

# Get first 10 pages
result = await extract_text("huge.pdf", pages="1-10")

# Limit table rows
tables = await extract_tables("data.pdf", max_rows_per_table=50)

# Structure only
tables = await extract_tables("data.pdf", summary_only=True)

URL Processing

PDFs can be fetched directly from HTTPS URLs:

result = await extract_text("https://example.com/report.pdf")

Files are cached locally for subsequent operations.


System Dependencies

Some features require system packages:

Feature

Dependency

OCR

tesseract-ocr

Camelot tables

ghostscript

Tabula tables

default-jre-headless

PDF to images

poppler-utils

markdown_to_pdf

pandoc + one of: tectonic, texlive-xetex (+ texlive-latex-extra), weasyprint, wkhtmltopdf

Picking a PDF engine for markdown_to_pdf

Pandoc takes markdown → HTML or LaTeX → PDF. The LaTeX path produces the most polished output but needs a TeX install. Trade-offs:

Engine

Disk size

Notes

tectonic

~30 MB

Recommended for new installs. Single static binary. Downloads LaTeX packages on demand — no upfront mass-install.

xelatex + texlive-latex-extra

~500 MB

Best output once installed. Use if you already run TeX. The -extra package matters: pandoc's default template needs lastpage, xcolor, framed, fancyhdr, etc. — all of which live there, not in texlive-xetex.

xelatex alone (just texlive-xetex)

~200 MB

Often breaks. Expect ! LaTeX Error: File 'X.sty' not found on real docs.

weasyprint

~40 MB

Pure-Python (pip install weasyprint) + cairo/pango system libs. HTML/CSS path — no LaTeX. Good for simple docs; weaker on math, footnotes, citations.

wkhtmltopdf

~40 MB

Older HTML-to-PDF tool. Adequate but less actively maintained.

Ubuntu/Debian:

sudo apt-get install tesseract-ocr tesseract-ocr-eng poppler-utils ghostscript default-jre-headless

# For markdown_to_pdf — pick one engine route:

# Option A — tectonic (smallest, downloads packages on demand)
sudo apt-get install pandoc
# tectonic isn't in apt — install via cargo or download static binary:
#   https://tectonic-typesetting.github.io/en-US/install.html

# Option B — full TeX (best quality, large download)
sudo apt-get install pandoc texlive-xetex texlive-latex-extra texlive-fonts-extra

# Option C — weasyprint (skip TeX entirely)
sudo apt-get install pandoc
pip install weasyprint

Arch Linux:

sudo pacman -S tesseract tesseract-data-eng poppler ghostscript jre-openjdk-headless

# For markdown_to_pdf — pick one engine route:

# Option A — tectonic (recommended for new installs, in official repo)
sudo pacman -S pandoc tectonic

# Option B — full TeX (best output, ~500 MB)
sudo pacman -S pandoc texlive-xetex texlive-latexextra texlive-fontsextra

# Option C — weasyprint (skip TeX)
sudo pacman -S pandoc
pip install weasyprint   # or: uv pip install weasyprint

# Option D — wkhtmltopdf (from AUR)
yay -S wkhtmltopdf-static

macOS (Homebrew):

brew install tesseract poppler ghostscript

# For markdown_to_pdf — pick one engine route:

# Option A — tectonic (recommended)
brew install pandoc tectonic

# Option B — full TeX (mactex-no-gui includes the latex-extra equivalent)
brew install pandoc
brew install --cask mactex-no-gui

# Option C — weasyprint
brew install pandoc weasyprint

Optional Extras

The base install stays lean. Heavy or niche dependencies are gated behind extras:

Extra

Adds

When to install

mcp-pdf[forms]

reportlab

Form creation tools (create_form_pdf, permit forms)

mcp-pdf[tables]

camelot-py, tabula-py

Higher-accuracy table extraction (also needs Java + Ghostscript)

mcp-pdf[markdown]

pypandoc

markdown_to_pdf tool (also needs pandoc binary)

mcp-pdf[all]

All of the above

Want everything


Configuration

All optional. Defaults are tuned for local stdio use (Claude Desktop, Claude Code), which is the common case.

Variable

Default

Purpose

PDF_TEMP_DIR

/tmp/mcp-pdf-processing

Working directory for intermediate files

MCP_PDF_MAX_SIZE

no limit

Max input PDF size in MB. Set 0 or leave unset to disable

ALLOWED_DOMAINS

all

Comma-separated host allowlist for fetching PDFs over HTTPS

DEBUG

false

Set true for verbose logging

TESSDATA_PREFIX

system

Tesseract language data location (read by Tesseract, not by this server)

Output-path restriction

MCP_PDF_ALLOWED_PATHS takes a colon-separated list of directories that tools may write into. It is not enforced in stdio mode, which is the default, on the reasoning that a local server writing to the user's own filesystem does not need to be fenced off from it. Enforcement switches on when either of these is set:

Variable

Effect

MCP_TRANSPORT=http

Treat as network-exposed; enforce MCP_PDF_ALLOWED_PATHS

MCP_PUBLIC_MODE

Any non-empty value does the same

If you expose this server over HTTP, set both MCP_TRANSPORT=http and MCP_PDF_ALLOWED_PATHS. Setting the allowlist alone has no effect in stdio mode. And note the framing in CLAUDE.md: application-level path checks are a speed bump, not a boundary. Real isolation comes from running as an unprivileged user in a container with the filesystem it actually needs and nothing more.

XFA limits

Bound the XFA parser against hostile or merely enormous templates. expat caps entity amplification relative to input size, so bounding the input is what bounds the expansion.

Variable

Default

Purpose

MCP_PDF_MAX_XFA_TEMPLATE_BYTES

16777216 (16 MB)

Reject XFA templates larger than this

MCP_PDF_MAX_XFA_DEPTH

100

Max <subform> nesting before giving up

MCP_PDF_MAX_XFA_INLINE_FIELDS

5000

Fields serialized into one response before truncating


Development

# Run tests
uv run pytest

# With coverage
uv run pytest --cov=mcp_pdf

# Format
uv run black src/ tests/

# Lint
uv run ruff check src/ tests/

Versioning

CalVer, YYYY.MM.DD, with a PEP 440 post-release segment for same-day fixes (2026.09.21.1). This package is a thin layer over PyMuPDF, pdfplumber, pypdf, Camelot, Tabula, Tesseract, pandoc and several system binaries, and most surprising behavior traces back to one of those. When a PDF misbehaves, the useful question is "when was this last tested against those?", which a date answers.

Releases before 2026.09.21 used semver. PEP 440 compares the release tuple numerically, so (2026, 9, 21) sorts after (2, 3, 1) and pip install --upgrade works correctly across the switch. See CHANGELOG.md.

License

MIT

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI-driven PDF document processing including PDF to Markdown conversion, intelligent text and table extraction, image extraction, format conversion between PDF/Word/Markdown, batch processing, and fuzzy search - optimized for LLM context and RAG workflows.
    2
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Provides intelligent OCR and PDF processing capabilities that automatically detect whether PDFs contain digital text or scanned images and apply appropriate extraction methods. Supports text extraction, OCR processing, structure analysis, and batch operations.
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables AI applications to read and process PDF files with intelligent file search, text extraction, image processing, and optional OCR support for scanned documents.
    MIT