Skip to main content
Glama
README.md
# pdf-extract-mcp

FastMCP server for PDF text extraction and metadata — part of the [Palimpsest](https://github.com/palimpsest-labs/palimpsest) intelligence toolkit.

## Why

LLMs can't read PDFs natively. This server fills the gap — extract text from PDF documents, pull document metadata (title, author, producer, creation date), and extract specific page ranges. All results are returned as structured JSON for easy downstream processing.

## Architecture

Pluggable PDF backends with graceful degradation:

* **base install** (`pip install pdf-extract-mcp`) — pypdf-based text extraction and metadata (~5 MB)

The server tries backends in priority order: pypdf → none.

## Tools

| Tool | Description |
|---|---|
| `extract_text(path, pages="")` | Extract text from a PDF — all pages or a specific range (e.g. '1-5') |
| `extract_metadata(path)` | Full metadata: file info, SHA-256, PDF metadata, encryption status |
| `extract_pages(path, page_range)` | Convenience wrapper for extracting specific page ranges |

## Installation

```bash
git clone https://github.com/palimpsest-labs/pdf-extract-mcp
cd pdf-extract-mcp
python3 -m venv .venv
source .venv/bin/activate

pip install -e .
```

## Usage

```python
from pdf_extract_mcp.server import extract_text, extract_metadata, extract_pages

# Get everything at once
result = extract_text("~/documents/report.pdf")

# Or separate calls
meta = extract_metadata("~/documents/report.pdf")
pages = extract_pages("~/documents/report.pdf", "1-3,5")
```

## Security

Paths must be absolute and resolve to a location under the user's home directory. Path traversal (`..`) and paths starting with `/` are rejected. Symlinks are resolved before checking home-directory containment.

## License

MIT