@shuji-bonji/pdf-spec-mcp
by shuji-bonji
README.md
# PDF SPEC MCP Server
[](https://github.com/shuji-bonji/pdf-spec-mcp/actions/workflows/ci.yml)
[](https://www.npmjs.com/package/@shuji-bonji/pdf-spec-mcp)
[日本語版 README はこちら](README.ja.md)
An MCP (Model Context Protocol) server that provides structured access to ISO 32000 (PDF) specification documents. Enables LLMs to navigate, search, and analyze PDF specifications through well-defined tools.
> [!IMPORTANT]
> **This is a specification *reference*, not a rule engine.**
> It retrieves and structures the text of ISO 32000 — clauses, tables, definitions, and
> `shall`/`should`/`may` requirements. It does **not** examine a PDF file, and it cannot tell
> you whether a document conforms to anything. Conformance verdicts come from
> [pdf-verify-mcp](https://github.com/shuji-bonji/pdf-verify-mcp)
> (`validate_conformance` / `evaluate_policy`).
>
> The distinction matters because three different things get conflated:
> **declaration** — a label the file wrote about itself ("I am PDF/A" in the metadata). Writing it is not evidence /
> **conformance** — whether the file actually meets the standard. There is no way to prove it in full; you can only find where it breaks the rules /
> **validation** — what a validator (veraPDF and the like) reports against the checks it implements. A pass means "this inspection did not fail", not "the file conforms to the standard".
> Reading a `shall` here tells you what the standard requires — not whether your file meets it.
>
> **A search that returns nothing means "cannot answer", not "no such requirement."**
> ISO 19005 (PDF/A) and ETSI PAdES are outside this corpus; see `list_specs` → `coverage.gaps`.
### What each PDF family server does — and does not do
| Server | Does | **Does not** |
|---|---|---|
| **pdf-spec-mcp** (this) | Search, retrieve and extract requirements from 17 PDF-related documents | **Is not a rule engine.** Does not define business rules, inspect PDF files, or validate schemas. ISO 19005 (PDF/A) is not part of the corpus |
| [pdf-reader-mcp](https://github.com/shuji-bonji/pdf-reader-mcp) | Extract text / tables / structure tree / fonts / annotations / images / signature *fields* | **Does not verify cryptography.** Does not read the incremental-update history, does not map object IDs to coordinates, does not OCR |
| [pdf-writer-mcp](https://github.com/shuji-bonji/pdf-writer-mcp) | Create, page operations, tagging, forms, annotations, metadata, attachments, PDF/A-3b scaffolding | **Does not sign.** Does not make the file meet the standard — it can write a *label*, not conformance |
| [pdf-verify-mcp](https://github.com/shuji-bonji/pdf-verify-mcp) | Conformance validation (delegated to veraPDF), cryptographic signature verification, tamper detection, policy verdicts | **Does not prove the file meets the standard** (it can only find where it breaks the rules). Does not vouch for the signer's identity. Does not judge whether the content is true |
> [!IMPORTANT]
> **PDF specification files are NOT included in this package.**
> You must obtain the PDF specification documents separately and place them in a local directory.
>
> **Download from:** [PDF Association — Sponsored Standards](https://pdfa.org/sponsored-standards/)
>
> See "[Setup](#setup)" for details.
## Features
- **Multi-spec support** — Auto-discovers and manages up to 17 PDF-related documents (ISO 32000-2, PDF/UA, Tagged PDF guides, etc.)
- **Structured content extraction** — Headings, paragraphs, lists, tables, and notes from any section
- **Full-text search** — Keyword search with section-aware context snippets
- **Requirements extraction** — Extracts normative language (shall / must / may) per ISO conventions
- **Definitions lookup** — Term definitions from Section 3 (Definitions)
- **Table extraction** — Multi-page table detection with header merging
- **Version comparison** — Diff PDF 1.7 vs PDF 2.0 section structures
- **Bounded-concurrency processing** — Parallel page processing for large documents
- **On-disk index cache** — The search index and the full requirements scan are built once per PDF and reused by every later process (ISO 32000-2: ~6 s → ~0.2 s)
## Architecture
```mermaid
graph LR
subgraph Client["MCP Client"]
LLM["LLM<br/>(Claude, etc.)"]
end
subgraph Server["PDF Spec MCP Server"]
direction TB
MCP["MCP Server<br/>index.ts"]
subgraph Tools["Tools Layer"]
direction LR
T1["list_specs"]
T2["get_structure"]
T3["get_section"]
T4["search_spec"]
T5["get_requirements"]
T6["get_definitions"]
T7["get_tables"]
T8["compare_versions"]
end
subgraph Services["Services Layer"]
direction LR
REG["Registry<br/>Auto-discovery"]
LOADER["Loader<br/>LRU Cache"]
SVC["PDFService<br/>Orchestration"]
CMP["CompareService<br/>Version Diff"]
end
subgraph Extractors["Extractors"]
direction LR
OUTLINE["OutlineResolver<br/>TOC & Section Index"]
CONTENT["ContentExtractor<br/>Structured Extraction"]
SEARCH["SearchIndex<br/>Full-text Search"]
REQ["RequirementExtractor"]
DEF["DefinitionExtractor"]
end
subgraph Utils["Utils"]
direction LR
CACHE["LRU Cache"]
CONC["Concurrency"]
VALID["Validation"]
end
end
subgraph PDFs["PDF Spec Files (obtained separately)"]
direction LR
PDF1["ISO 32000-2<br/>(PDF 2.0)"]
PDF2["ISO 32000-1<br/>(PDF 1.7)"]
PDF3["TS 32001–32005<br/>PDF/UA, etc."]
end
LLM <-->|"stdio / JSON-RPC"| MCP
MCP --> Tools
Tools --> Services
Services --> Extractors
Services --> Utils
LOADER --> PDFs
REG -->|"Filename pattern<br/>auto-discovery"| PDFs
style Client fill:#e8f4f8,stroke:#2196F3
style PDFs fill:#fff3e0,stroke:#FF9800
style Tools fill:#e8f5e9,stroke:#4CAF50
style Services fill:#f3e5f5,stroke:#9C27B0
style Extractors fill:#fce4ec,stroke:#E91E63
style Utils fill:#f5f5f5,stroke:#9E9E9E
```
### Layer Overview
| Layer | Responsibility |
| -------------- | ---------------------------------------------------------------------------------- |
| **Tools** | MCP tool schema definitions & handlers (input validation) |
| **Services** | Business logic (PDF registry, loader, orchestration) |
| **Extractors** | Information extraction from PDFs (TOC, content, search, requirements, definitions) |
| **Utils** | Shared utilities (cache, concurrency, validation) |
## Setup
### 1. Obtain PDF Specification Files
> [!WARNING]
> PDF specifications are **copyrighted documents** and are not included in this package.
> Download them from the sources below and place them in a local directory.
| Document | Source |
| ---------------------------- | ----------------------------------------------------------------------------------------------- |
| ISO 32000-2 (PDF 2.0) | [PDF Association](https://pdfa.org/resource/iso-32000-pdf/) |
| ISO 32000-1 (PDF 1.7) | [Adobe (free)](https://opensource.adobe.com/dc-acrobat-sdk-docs/pdfstandards/PDF32000_2008.pdf) |
| TS 32001–32005, PDF/UA, etc. | [PDF Association — Sponsored Standards](https://pdfa.org/sponsored-standards/) |
All 17 files below are supported. You do not need all of them — place only the specs you need (at minimum, ISO 32000-2 is recommended).
```
pdf-specs/
│
│ ── Standards ─────────────────────────────
├── ISO_32000-2_sponsored_EC3.pdf # iso32000-2 : PDF 2.0 EC3 (recommended; falls back to -ec2.pdf)
├── ISO_32000-2-2020_sponsored.pdf # iso32000-2-2020 : PDF 2.0 original
├── PDF32000_2008.pdf # pdf17 : PDF 1.7 (for version comparison)
├── pdfreference1.7old.pdf # pdf17old : Adobe PDF Reference 1.7
│
│ ── Technical Specifications (TS) ─────────
├── ISO_TS_32001-2022_sponsored_EC3.pdf # ts32001 : Hash extensions (SHA-3)
├── ISO_TS_32002-2022_sponsored_EC3.pdf # ts32002 : Digital signature extensions (ECC/PAdES)
├── ISO_TS_32003-2023_sponsored.pdf # ts32003 : AES-GCM encryption
├── ISO-TS-32004-2024_sponsored.pdf # ts32004 : Integrity protection
├── ISO-TS-32005-2023-sponsored.pdf # ts32005 : Namespace mapping
│
│ ── PDF/UA (Accessibility) ────────────────
├── ISO-14289-1-2014-sponsored.pdf # pdfua1 : PDF/UA-1
├── ISO-14289-2-2024-sponsored.pdf # pdfua2 : PDF/UA-2
│
│ ── Guides ────────────────────────────────
├── Tagged-PDF-Best-Practice-Guide.pdf # tagged-bpg : Tagged PDF Best Practice
├── Well-Tagged-PDF-WTPDF-1.0.pdf # wtpdf : Well-Tagged PDF
├── PDF-Declarations.pdf # declarations: PDF Declarations
│
│ ── Application Notes ─────────────────────
├── PDF20_AN001-BPC.pdf # an001 : Black Point Compensation
├── PDF20_AN002-AF.pdf # an002 : Associated Files
└── PDF20_AN003-ObjectMetadataLocations.pdf # an003 : Object Metadata
```
### 2. Install
This package ships a CLI binary (`pdf-spec-mcp`) intended to be launched by an MCP client.
**You do not need to install it manually** — just point your MCP client to `npx @shuji-bonji/pdf-spec-mcp@latest` as shown in the next step.
If you want to run it directly from the shell (e.g. for debugging):
```bash
PDF_SPEC_DIR=/path/to/pdf-specs npx -y @shuji-bonji/pdf-spec-mcp@latest
```
Or install it globally (optional):
```bash
npm install -g @shuji-bonji/pdf-spec-mcp
PDF_SPEC_DIR=/path/to/pdf-specs pdf-spec-mcp
```
### 3. Configure MCP Client
#### Environment Variable
| Variable | Description | Default |
| -------------------- | --------------------------------------------------------------------- | ------------------------------------------ |
| `PDF_SPEC_DIR` | Directory containing PDF specification files | (required) |
| `PDF_SPEC_CACHE_DIR` | Where the on-disk index cache lives (see [Index cache](#index-cache)) | `${XDG_CACHE_HOME:-~/.cache}/pdf-spec-mcp` |
| `PDF_SPEC_CACHE` | Set to `off` to neither read nor write the index cache | on |
#### Claude Desktop
Add to `claude_desktop_config.json`:
```json
{
"mcpServers": {
"pdf-spec": {
"command": "npx",
"args": ["-y", "@shuji-bonji/pdf-spec-mcp@latest"],
"env": {
"PDF_SPEC_DIR": "/path/to/pdf-specs"
}
}
}
}
```
> [!IMPORTANT]
> **Use `@latest` (or pin a version).** `npx -y <pkg>` without a version keeps running whatever
> it cached the first time — `-y` only skips the install prompt, it does not check for updates.
> A bare specifier will happily run a months-old release. `@latest` makes npx check the registry
> on each start; pin `@0.4.0` instead if you want reproducibility.
> To clear a stale cache: `rm -rf ~/.npm/_npx`.
#### Cursor / VS Code
Add to `.cursor/mcp.json` or VS Code MCP settings:
```json
{
"mcpServers": {
"pdf-spec": {
"command": "npx",
"args": ["-y", "@shuji-bonji/pdf-spec-mcp@latest"],
"env": {
"PDF_SPEC_DIR": "/path/to/pdf-specs"
}
}
}
}
```
## Index cache
Two operations walk every page of a specification: the first `search_spec` on a spec builds
its full-text index (ISO 32000-2, 1023 pages: about 6 s on a laptop), and `get_requirements`
without a `section` scans every section (about 11 s). Everything else opens only the pages it
needs and answers in well under a second.
Since 0.5.0 those two results are written to disk after the first build and read back by every
later process — an MCP client that starts one server per session no longer pays the build each
time. The second process answers the same `search_spec` in about 0.2 s and the full
requirements scan in about 0.02 s, from byte-for-byte the same index.
- **Location:** `${PDF_SPEC_CACHE_DIR:-${XDG_CACHE_HOME:-~/.cache}/pdf-spec-mcp}/v1/<version>/<spec>.<kind>.<sha256[0:16]>.json`.
The whole 17-spec corpus is about 18 MB per package version.
- **Key:** package version, `pdfjs-dist` version, spec id, and the SHA-256 of the PDF. A
replaced PDF, an upgraded server, or an upgraded pdfjs all miss and rebuild. Entries of
older versions are left in place (another install may still use them); `--clear-cache`
removes everything.
- **Failure is a miss, never an error:** an unreadable, truncated, or foreign file is rebuilt;
an unwritable directory is reported once on stderr and the server carries on without a cache.
- **It is derived from *your* copy of the PDFs and stays on your machine.** It is not part of
the package and must not be redistributed — the specifications are copyrighted.
Nothing about searching changes: the same in-memory structure is searched by the same code.
Only where it comes from (built vs. read) does.
### Pre-building the cache
The cache fills lazily, one spec at a time as tools touch it. To warm every spec up front — after
installing, after upgrading, or from cron — run the CLI (it uses the same code path as the tools,
processes specs sequentially, and exits):
```bash
PDF_SPEC_DIR=/path/to/pdf-specs npx -y @shuji-bonji/pdf-spec-mcp@latest --build-cache
# --spec=iso32000-2,pdf17 only these specs
# --force rebuild even when a valid entry exists
npx -y @shuji-bonji/pdf-spec-mcp@latest --cache-info # directory, key, entries
npx -y @shuji-bonji/pdf-spec-mcp@latest --clear-cache # remove the directory
```
A full build of the 17-spec corpus takes about a minute on a laptop.
## Available Tools
All tools accept an optional `spec` parameter to target a specific specification (default: `iso32000-2`).
| Tool | Description |
| ------------------ | ----------------------------------------------------------------- |
| `list_specs` | List all discovered PDF specifications with metadata |
| `get_structure` | Get section hierarchy (table of contents) with configurable depth |
| `get_section` | Get structured content of a specific section |
| `search_spec` | Full-text keyword search across a specification |
| `get_requirements` | Extract normative requirements (shall/must/may) |
| `get_definitions` | Lookup term definitions |
| `get_tables` | Extract table structures from a section |
| `compare_versions` | Compare PDF 1.7 and PDF 2.0 section structures |
### `list_specs` — Discover Specifications
List all available specification documents. Use the returned IDs as the `spec` parameter in other tools.
```jsonc
// List all specs
{ }
// Filter by category
{ "category": "ts" } // Technical specs only
{ "category": "pdfua" } // PDF/UA only
{ "category": "guide" } // Guide documents only
```
### `get_structure` — Table of Contents
Get the section hierarchy (TOC tree) of a specification.
```jsonc
// PDF 2.0 top-level sections only
{ "max_depth": 1 }
// Expand to 2 levels
{ "max_depth": 2 }
// TS 32002 (Digital Signatures) full structure
{ "spec": "ts32002" }
// PDF/UA-2 structure
{ "spec": "pdfua2", "max_depth": 2 }
```
### `get_section` — Section Content
Get structured content (headings, paragraphs, lists, tables, notes) of a specific section.
A parent section returns its entire subtree (its preamble followed by all subsections, in document order). Top-level clauses can be very large — prefer the most specific section number.
```jsonc
// PDF 2.0 Section 7.3.4.2 (Literal Strings)
{ "section": "7.3.4.2" }
// PDF 2.0 Annex A
{ "section": "Annex A" }
// TS 32002 Section 5
{ "spec": "ts32002", "section": "5" }
// PDF/UA-2 Section 8 (Tagged PDF)
{ "spec": "pdfua2", "section": "8" }
```
### `search_spec` — Full-text Search
Search across a specification with section-aware context snippets. The first call on a spec
builds its index (a few seconds); the index is then cached on disk (see [Index cache](#index-cache)).
```jsonc
// Search PDF 2.0 for "digital signature"
{ "query": "digital signature" }
// Limit results
{ "query": "font", "max_results": 5 }
// Search within TS 32002
{ "spec": "ts32002", "query": "CMS" }
```
### `get_requirements` — Normative Requirements
Extract normative requirements (shall / must / may) per ISO conventions.
```jsonc
// All requirements in section 12.8
{ "section": "12.8" }
// Only "shall" requirements
{ "section": "12.8", "level": "shall" }
// Only "shall not" requirements
{ "section": "7.3", "level": "shall not" }
// PDF/UA-2 requirements
{ "spec": "pdfua2", "section": "8", "level": "shall" }
```
### `get_definitions` — Term Definitions
Look up term definitions from Section 3 (Definitions).
```jsonc
// Search for "font" definitions
{ "term": "font" }
// List all definitions
{ }
// PDF/UA definitions
{ "spec": "pdfua2", "term": "artifact" }
```
### `get_tables` — Table Extraction
Extract table structures (headers, rows, captions) from a section. Multi-page tables are automatically merged.
```jsonc
// All tables in section 7.3.4.2 (Table 3 — Escape sequences)
{ "section": "7.3.4.2" }
// Specific table only (0-based index)
{ "section": "7.3.4.2", "table_index": 0 }
// TS spec tables
{ "spec": "ts32002", "section": "5" }
```
### `compare_versions` — Version Comparison
Compare section structures between PDF 1.7 (ISO 32000-1) and PDF 2.0 (ISO 32000-2). Uses title-based automatic matching to detect matched, added, and removed sections.
> [!NOTE]
> This tool requires both PDF 1.7 (`PDF32000_2008.pdf`) and PDF 2.0 files in `PDF_SPEC_DIR`.
```jsonc
// Diff section 12.8 (Digital Signatures)
{ "section": "12.8" }
// Compare all top-level sections
{ }
```
## Supported Specifications
The server auto-discovers PDF files in `PDF_SPEC_DIR` by filename pattern matching:
| Category | Spec IDs | Documents |
| ------------------ | ---------------------------------------------------- | ------------------------------------------------------- |
| **Standard** | `iso32000-2`, `iso32000-2-2020`, `pdf17`, `pdf17old` | ISO 32000-2 (PDF 2.0), ISO 32000-1 (PDF 1.7) |
| **Technical Spec** | `ts32001` – `ts32005` | Hash, Digital Signatures, AES-GCM, Integrity, Namespace |
| **PDF/UA** | `pdfua1`, `pdfua2` | Accessibility (ISO 14289-1, 14289-2) |
| **Guide** | `tagged-bpg`, `wtpdf`, `declarations` | Tagged PDF, Well-Tagged PDF, Declarations |
| **App Note** | `an001` – `an003` | BPC, Associated Files, Object Metadata |
## Directory Structure
```
src/
├── index.ts # Entry point: MCP server on stdio, or the cache CLI
├── cli.ts # --build-cache / --clear-cache / --cache-info
├── config.ts # Configuration & spec patterns
├── errors.ts # Error hierarchy (PDFSpecError → sub-classes)
├── services/
│ ├── pdf-registry.ts # Auto-discovery of PDF files
│ ├── pdf-loader.ts # PDF loading with LRU cache
│ ├── pdf-service.ts # Orchestration layer
│ ├── index-store.ts # On-disk cache for the search / requirements indexes
│ ├── compare-service.ts # Version comparison
│ ├── outline-resolver.ts # Section index builder
│ ├── content-extractor.ts # Structured content extraction
│ ├── search-index.ts # Full-text search index
│ ├── requirement-extractor.ts
│ └── definition-extractor.ts
├── tools/
│ ├── definitions.ts # MCP tool schemas
│ └── handlers.ts # Tool implementations
├── types/
│ └── index.ts # Shared type definitions
└── utils/
├── concurrency.ts # mapConcurrent (bounded Promise.all)
├── text.ts # Text normalization
├── cache.ts # LRU cache
├── file-hash.ts # SHA-256 of a PDF (index cache key)
├── validation.ts # Input validation
└── logger.ts # Structured logger
```
## Development
```bash
git clone https://github.com/shuji-bonji/pdf-spec-mcp.git
cd pdf-spec-mcp
npm install
npm run build
# Unit tests
npm run test
# E2E tests (requires PDF files in ./pdf-spec/)
npm run test:e2e
# Lint & format
npm run lint
npm run format:check
```
## License
[MIT](LICENSE)
TDQS
A4.2/5.0
Scored across 8 tools
Disambiguation5/5
Each tool targets a distinct aspect of the PDF specification: version comparison, definitions, requirements, sections, structure, tables, listing specs, and searching. No overlap in purpose.
Naming Consistency5/5
All tool names follow a consistent verb_noun pattern in snake_case (e.g., get_section, search_spec). There is no mixing of conventions.
Tool Count5/5
With 8 tools, the set is well-scoped for exploring and comparing PDF specifications. Each tool serves a clear function without being excessive or insufficient.
Completeness5/5
The tool surface covers common operations for a reference document: listing, searching, retrieving structure, content, tables, definitions, requirements, and version comparison. No obvious gaps for its read-only domain.
Maintenance
ActivityActive
ResponsivenessSlow