Skip to main content
Glama

pdf-extract-mcp

FastMCP server for PDF text extraction and metadata — part of the Palimpsest intelligence toolkit.

Why

LLMs can't read PDFs natively. This server fills the gap — extract text from PDF documents, pull document metadata (title, author, producer, creation date), and extract specific page ranges. All results are returned as structured JSON for easy downstream processing.

Related MCP server: pdf-mcp

Architecture

Pluggable PDF backends with graceful degradation:

  • base install (pip install pdf-extract-mcp) — pypdf-based text extraction and metadata (~5 MB)

The server tries backends in priority order: pypdf → none.

Tools

Tool

Description

extract_text(path, pages="")

Extract text from a PDF — all pages or a specific range (e.g. '1-5')

extract_metadata(path)

Full metadata: file info, SHA-256, PDF metadata, encryption status

extract_pages(path, page_range)

Convenience wrapper for extracting specific page ranges

Installation

git clone https://github.com/palimpsest-labs/pdf-extract-mcp
cd pdf-extract-mcp
python3 -m venv .venv
source .venv/bin/activate

pip install -e .

Usage

from pdf_extract_mcp.server import extract_text, extract_metadata, extract_pages

# Get everything at once
result = extract_text("~/documents/report.pdf")

# Or separate calls
meta = extract_metadata("~/documents/report.pdf")
pages = extract_pages("~/documents/report.pdf", "1-3,5")

Security

Paths must be absolute and resolve to a location under the user's home directory. Path traversal (..) and paths starting with / are rejected. Symlinks are resolved before checking home-directory containment.

License

MIT

F
license - not found
Not graded
quality - not tested
C
maintenance

Maintenance

Maintainers
Response time
Release cycle
Releases (12mo)
Commit activity

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    An MCP server that provides comprehensive PDF processing capabilities including text extraction, image extraction, table detection, annotation extraction, metadata retrieval, page rendering, and document structure analysis.
  • A
    license
    A
    quality
    C
    maintenance
    An MCP server for reading, rendering, and searching PDF files, specifically optimized for LLMs to extract text, tables, and technical diagrams. It enables metadata retrieval, multi-format text extraction, and page-to-image rendering using PyMuPDF.
    5
    66
    MIT
  • F
    license
    Not graded
    quality
    D
    maintenance
    MCP server for extracting text from PDF files, supporting local files and URLs.

View all related MCP servers

Related MCP Connectors

  • MCP server for the PDFGate API. Generate PDFs, manage documents and handle e-signatures.

  • OCR.space MCP — wraps the OCR.space API (ocr.space) for image/PDF → text OCR.

  • MCP server for the FFmpeg Micro video transcoding API — create, monitor, download transcodes.

View all MCP Connectors

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/palimpsest-labs/pdf-extract-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server