pdf-extract-mcp
README.md
# pdf-extract-mcp
FastMCP server for PDF text extraction and metadata — part of the [Palimpsest](https://github.com/palimpsest-labs/palimpsest) intelligence toolkit.
## Why
LLMs can't read PDFs natively. This server fills the gap — extract text from PDF documents, pull document metadata (title, author, producer, creation date), and extract specific page ranges. All results are returned as structured JSON for easy downstream processing.
## Architecture
Pluggable PDF backends with graceful degradation:
* **base install** (`pip install pdf-extract-mcp`) — pypdf-based text extraction and metadata (~5 MB)
The server tries backends in priority order: pypdf → none.
## Tools
| Tool | Description |
|---|---|
| `extract_text(path, pages="")` | Extract text from a PDF — all pages or a specific range (e.g. '1-5') |
| `extract_metadata(path)` | Full metadata: file info, SHA-256, PDF metadata, encryption status |
| `extract_pages(path, page_range)` | Convenience wrapper for extracting specific page ranges |
## Installation
```bash
git clone https://github.com/palimpsest-labs/pdf-extract-mcp
cd pdf-extract-mcp
python3 -m venv .venv
source .venv/bin/activate
pip install -e .
```
## Usage
```python
from pdf_extract_mcp.server import extract_text, extract_metadata, extract_pages
# Get everything at once
result = extract_text("~/documents/report.pdf")
# Or separate calls
meta = extract_metadata("~/documents/report.pdf")
pages = extract_pages("~/documents/report.pdf", "1-3,5")
```
## Security
Paths must be absolute and resolve to a location under the user's home directory. Path traversal (`..`) and paths starting with `/` are rejected. Symlinks are resolved before checking home-directory containment.
## License
MIT
This server cannot be deployed
Maintenance
ActivitySlowing
ResponsivenessNo issues