MCP PDF
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@MCP PDFsummarize the quarterly report PDF for me"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
📄 MCP PDF
A FastMCP server for PDF processing
52 tools for text extraction, OCR, tables, forms, XFA, annotations, markdown↔PDF, and more
Works great with MCP Office Tools
What It Does
MCP PDF extracts content from PDFs using multiple libraries with automatic fallbacks. If one method fails, it tries another.
Core capabilities:
Text extraction via PyMuPDF, pdfplumber, or pypdf (auto-fallback)
Table extraction via Camelot, pdfplumber, or Tabula (auto-fallback)
OCR for scanned documents via Tesseract
Form handling - extract, fill, and create PDF forms
Document assembly - merge, split, reorder pages
Annotations - sticky notes, highlights, stamps
Vector graphics - extract to SVG for schematics and technical drawings
Format conversion - PDF ↔ Markdown (PDF→MD via PyMuPDF, MD→PDF via pandoc)
XFA forms - Schema extraction for dynamic Adobe LiveCycle forms that no open-source library can render
Related MCP server: PDF MCP Flow
Quick Start
# Run from PyPI (one-shot, no permanent install)
uvx mcp-pdf
# Add to Claude Code — note the `--` separator before uvx
claude mcp add pdf-tools -- uvx mcp-pdf
# Include the markdown_to_pdf tool (requires pandoc on host)
claude mcp add pdf-tools -- uvx --from "mcp-pdf[markdown]" mcp-pdf
uvxcaches tool installs aggressively. After upgrading to a new release, force a refresh withuvx --refresh mcp-pdf(oruvx --refresh --from "mcp-pdf[markdown]" mcp-pdfif you're using extras).
git clone https://github.com/rsp2k/mcp-pdf
cd mcp-pdf
uv sync
# System dependencies (Ubuntu/Debian)
sudo apt-get install tesseract-ocr tesseract-ocr-eng poppler-utils ghostscript
# For markdown_to_pdf — pick one PDF-engine route:
sudo apt-get install pandoc tectonic # recommended (small)
# or: sudo apt-get install pandoc texlive-xetex texlive-latex-extra # full TeX
# or: sudo apt-get install pandoc && pip install weasyprint # skip TeX
# Verify
uv run python examples/verify_installation.pyTools
Content Extraction
Tool | What it does |
| Pull text from PDF pages with automatic chunking for large files |
| Extract tables to JSON, CSV, or Markdown |
| Extract embedded images |
| Get all hyperlinks with page filtering |
| OCR scanned documents using Tesseract |
| Export vector graphics to SVG (schematics, charts, drawings) |
Format Conversion
Tool | What it does |
| Convert PDF to markdown preserving structure; extracts images and SVG vectors to disk |
| Convert |
markdown_to_pdf requires: pip install mcp-pdf[markdown] plus the pandoc binary and at least one PDF engine (xelatex, pdflatex, tectonic, weasyprint, or wkhtmltopdf) on PATH. The tool auto-detects what's available and uses the highest-quality one. Pass pdf_engine= to override or extra_args= for raw pandoc options.
Document Analysis
Tool | What it does |
| Get title, author, creation date, page count, etc. |
| Extract table of contents and bookmarks |
| Detect columns, headers, footers |
| Check if PDF needs OCR |
| Diff two PDFs by text, structure, or metadata |
| Check for corruption, optimization opportunities |
| Report encryption, permissions, signatures |
Forms
Tool | What it does |
| Get form field names and values (AcroForm) |
| Fill form fields from JSON |
| Create new forms with text fields, checkboxes, dropdowns |
| Add fields to existing PDFs |
| Add a date field with format validation |
| Add a radio button group with mutual exclusion |
| Add a multi-line text area with word limits |
| Add validation rules to existing form fields |
| Check data against field rules and constraints before filling |
Field types are reported in a portable vocabulary (text/checkbox/radio/dropdown/date/signature plus button/unknown) shared between the AcroForm and XFA tools, so callers don't have to learn two models. extract_form_data also returns field_type_raw with the unmerged AcroForm widget type, since listbox and combobox both report as dropdown but differ on free-text entry and multi-select.
XFA Forms (Dynamic Adobe LiveCycle)
Real-estate forms, mortgage forms, government forms — many are dynamic XFA, where the layout + fields live in an XFA program that only Adobe's runtime can render. Every open-source PDF library (PyMuPDF, pdfium, MuPDF, pikepdf) only sees the static "Open in Adobe Reader" placeholder page. These tools recover the form schema instead.
Tool | What it does |
| Detect XFA and classify as dynamic / static. Use for branching before extract_form_data or convert_to_images |
| Parse the XFA template for field names, captions, UI types. Splits into shared (cross-form canonical), positional (opaque codes), other, and plumbing (producer internals, dropped) |
Check detection_failed before trusting is_xfa. Detection has three outcomes: True with an xfa_type, False for a readable non-XFA PDF, and None with detection_failed=True when the file could not be read. A truncated or corrupt form package lands in the third case and gets told so, rather than being reported as a file that simply has no XFA.
extract_xfa_fields defaults to the zipForm producer profile (Lone Wolf / zipForm Plus, the most common XFA producer in the wild), matched case-insensitively. Pass profile="generic" plus extra_plumbing_patterns / extra_positional_patterns for other producers; an unrecognized profile name is an error rather than a silent fallback. Plumbing patterns are substring searches while positional patterns are anchored at the start.
Every response carries success, is_xfa and xfa_type, including failures. Fields land in four categories, not three: shared, positional, other (the default bucket, which holds the majority on non-zipForm producers) and plumbing_fields_dropped. canonical_collisions flags distinct XFA names that canonicalize to the same key, which matters because that key is the cross-form join. The original XFA name is on every field as the round-trip key for filling; canonical_name appears only on shared fields. canonical_separator chooses _ (snake, default), . (dotted) or - (kebab). include_design_time_bbox=True opts into best-effort geometry, page-relative with a top-left origin, and not authoritative for dynamic XFA since subforms reflow at render time.
Permit Forms (Coordinate-Based)
For scanned PDFs or forms without interactive fields. Draws text at (x, y) coordinates.
Tool | What it does |
| Fill any PDF by drawing at coordinates (works with scanned forms) |
| Get field definitions for validation or UI generation |
| Check data against field schema before filling |
| Generate PDF showing field boundaries (debugging) |
| Insert image/text pages with "See page X" references |
Requires: pip install mcp-pdf[forms] (adds reportlab dependency)
Document Assembly
Tool | What it does |
| Combine multiple PDFs with bookmark preservation |
| Merge with page numbering, generated TOC, and per-file page ranges |
| Split into separate documents |
| Split by page ranges |
| Split at chapter/section boundaries |
| Rearrange pages in custom order |
| Rotate specific pages by 90, 180 or 270 degrees |
Structure Detection
Chapter-aware analysis, for long documents where you want to work a section at a time instead of paging through blindly.
Tool | What it does |
| Find headings via bookmarks, font-size heuristics and numbering patterns |
| Auto-split into per-chapter directories with markdown and images |
| Process several page ranges in one call, replacing dozens of individual calls |
detect_structure writes the full structure to JSON and returns a compact summary plus the path (roughly 1k tokens against ~20k inline). Pass inline=True when you want the whole thing in the response.
Content Analysis
Tool | What it does |
| Identify document type (invoice, contract, report) and structure |
| Generate a summary and key insights |
| Extract and analyze charts, diagrams and visual elements |
| Detect and analyze watermarks |
PDF Utilities
Tool | What it does |
| Render pages to PNG or JPEG at a chosen DPI |
| Reduce file size and improve load performance |
| Attempt recovery of a corrupted or damaged PDF |
Server Introspection
Not PDF tools, so they sit outside the count above. Useful when you want to know what a given install can actually do.
Tool | What it does |
| Enumerate what this server build can do |
| Report version, registered mixins and configuration |
Annotations
Tool | What it does |
| Add comment annotations |
| Highlight text regions |
| Add Approved/Draft/Confidential stamps |
| Export annotations to JSON |
How Fallbacks Work
The server tries multiple libraries for each operation:
Text extraction:
PyMuPDF (fastest)
pdfplumber (better for complex layouts)
pypdf (most compatible)
Table extraction:
Camelot (best accuracy, requires Ghostscript)
pdfplumber (no dependencies)
Tabula (requires Java)
If a PDF fails with one library, the next is tried automatically. An engine that raises and one that succeeds but finds no text both fall through, so a page PyMuPDF silently returns nothing for still gets tried by pdfplumber and pypdf.
The response tells you what happened: method_used names the engine that produced the text, and methods_attempted lists the ones skipped and why. When all three come back empty you get method_used: "none" plus an extraction_warning, rather than the emptiness being attributed to whichever engine happened to run last. Naming an engine explicitly (method="pdfplumber") disables the cascade, and a failure then raises instead of falling through.
Token Management
Large PDFs can overflow MCP response limits. The server handles this:
Automatic chunking splits large documents into page groups
Table row limits prevent huge tables from blowing up responses
Summary mode returns structure without full content
# Get first 10 pages
result = await extract_text("huge.pdf", pages="1-10")
# Limit table rows
tables = await extract_tables("data.pdf", max_rows_per_table=50)
# Structure only
tables = await extract_tables("data.pdf", summary_only=True)URL Processing
PDFs can be fetched directly from HTTPS URLs:
result = await extract_text("https://example.com/report.pdf")Files are cached locally for subsequent operations.
System Dependencies
Some features require system packages:
Feature | Dependency |
OCR |
|
Camelot tables |
|
Tabula tables |
|
PDF to images |
|
|
|
Picking a PDF engine for markdown_to_pdf
Pandoc takes markdown → HTML or LaTeX → PDF. The LaTeX path produces the most polished output but needs a TeX install. Trade-offs:
Engine | Disk size | Notes |
| ~30 MB | Recommended for new installs. Single static binary. Downloads LaTeX packages on demand — no upfront mass-install. |
| ~500 MB | Best output once installed. Use if you already run TeX. The |
| ~200 MB | Often breaks. Expect |
| ~40 MB | Pure-Python ( |
| ~40 MB | Older HTML-to-PDF tool. Adequate but less actively maintained. |
Ubuntu/Debian:
sudo apt-get install tesseract-ocr tesseract-ocr-eng poppler-utils ghostscript default-jre-headless
# For markdown_to_pdf — pick one engine route:
# Option A — tectonic (smallest, downloads packages on demand)
sudo apt-get install pandoc
# tectonic isn't in apt — install via cargo or download static binary:
# https://tectonic-typesetting.github.io/en-US/install.html
# Option B — full TeX (best quality, large download)
sudo apt-get install pandoc texlive-xetex texlive-latex-extra texlive-fonts-extra
# Option C — weasyprint (skip TeX entirely)
sudo apt-get install pandoc
pip install weasyprintArch Linux:
sudo pacman -S tesseract tesseract-data-eng poppler ghostscript jre-openjdk-headless
# For markdown_to_pdf — pick one engine route:
# Option A — tectonic (recommended for new installs, in official repo)
sudo pacman -S pandoc tectonic
# Option B — full TeX (best output, ~500 MB)
sudo pacman -S pandoc texlive-xetex texlive-latexextra texlive-fontsextra
# Option C — weasyprint (skip TeX)
sudo pacman -S pandoc
pip install weasyprint # or: uv pip install weasyprint
# Option D — wkhtmltopdf (from AUR)
yay -S wkhtmltopdf-staticmacOS (Homebrew):
brew install tesseract poppler ghostscript
# For markdown_to_pdf — pick one engine route:
# Option A — tectonic (recommended)
brew install pandoc tectonic
# Option B — full TeX (mactex-no-gui includes the latex-extra equivalent)
brew install pandoc
brew install --cask mactex-no-gui
# Option C — weasyprint
brew install pandoc weasyprintOptional Extras
The base install stays lean. Heavy or niche dependencies are gated behind extras:
Extra | Adds | When to install |
|
| Form creation tools ( |
|
| Higher-accuracy table extraction (also needs Java + Ghostscript) |
|
|
|
| All of the above | Want everything |
Configuration
All optional. Defaults are tuned for local stdio use (Claude Desktop, Claude Code), which is the common case.
Variable | Default | Purpose |
|
| Working directory for intermediate files |
| no limit | Max input PDF size in MB. Set |
| all | Comma-separated host allowlist for fetching PDFs over HTTPS |
|
| Set |
| system | Tesseract language data location (read by Tesseract, not by this server) |
Output-path restriction
MCP_PDF_ALLOWED_PATHS takes a colon-separated list of directories that tools may write into. It is not enforced in stdio mode, which is the default, on the reasoning that a local server writing to the user's own filesystem does not need to be fenced off from it. Enforcement switches on when either of these is set:
Variable | Effect |
| Treat as network-exposed; enforce |
| Any non-empty value does the same |
If you expose this server over HTTP, set both MCP_TRANSPORT=http and MCP_PDF_ALLOWED_PATHS. Setting the allowlist alone has no effect in stdio mode. And note the framing in CLAUDE.md: application-level path checks are a speed bump, not a boundary. Real isolation comes from running as an unprivileged user in a container with the filesystem it actually needs and nothing more.
XFA limits
Bound the XFA parser against hostile or merely enormous templates. expat caps entity amplification relative to input size, so bounding the input is what bounds the expansion.
Variable | Default | Purpose |
|
| Reject XFA templates larger than this |
|
| Max |
|
| Fields serialized into one response before truncating |
Development
# Run tests
uv run pytest
# With coverage
uv run pytest --cov=mcp_pdf
# Format
uv run black src/ tests/
# Lint
uv run ruff check src/ tests/Versioning
CalVer, YYYY.MM.DD, with a PEP 440 post-release segment for same-day fixes (2026.09.21.1). This package is a thin layer over PyMuPDF, pdfplumber, pypdf, Camelot, Tabula, Tesseract, pandoc and several system binaries, and most surprising behavior traces back to one of those. When a PDF misbehaves, the useful question is "when was this last tested against those?", which a date answers.
Releases before 2026.09.21 used semver. PEP 440 compares the release tuple numerically, so (2026, 9, 21) sorts after (2, 3, 1) and pip install --upgrade works correctly across the switch. See CHANGELOG.md.
License
MIT
This server cannot be deployed
Maintenance
Related MCP Connectors
Turn any PDF into structured JSON via AI + OCR: invoices, bank statements, contracts.
Extract text, tables and metadata from every PDF linked in a dataset, CSV or Google Sheet.
PDF tools + invoice extraction, bank statement parsing, GST reconciliation & GSTIN validation.
Extract tables, text and formulas from PDFs, including scanned pages and broken text layers.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceEnables comprehensive PDF processing including text extraction, image extraction, and OCR capabilities for reading text within images across multiple languages.12MIT
- AlicenseNot gradedqualityDmaintenanceEnables AI-driven PDF document processing including PDF to Markdown conversion, intelligent text and table extraction, image extraction, format conversion between PDF/Word/Markdown, batch processing, and fuzzy search - optimized for LLM context and RAG workflows.2MIT
- AlicenseNot gradedqualityDmaintenanceProvides intelligent OCR and PDF processing capabilities that automatically detect whether PDFs contain digital text or scanned images and apply appropriate extraction methods. Supports text extraction, OCR processing, structure analysis, and batch operations.MIT
- AlicenseNot gradedqualityDmaintenanceEnables AI applications to read and process PDF files with intelligent file search, text extraction, image processing, and optional OCR support for scanned documents.MIT