Skip to main content
Glama
shuji-bonji

@shuji-bonji/pdf-spec-mcp

by shuji-bonji
README.md
# PDF SPEC MCP Server

[![CI](https://github.com/shuji-bonji/pdf-spec-mcp/actions/workflows/ci.yml/badge.svg)](https://github.com/shuji-bonji/pdf-spec-mcp/actions/workflows/ci.yml)
[![npm version](https://img.shields.io/npm/v/@shuji-bonji/pdf-spec-mcp)](https://www.npmjs.com/package/@shuji-bonji/pdf-spec-mcp)

[日本語版 README はこちら](README.ja.md)

An MCP (Model Context Protocol) server that provides structured access to ISO 32000 (PDF) specification documents. Enables LLMs to navigate, search, and analyze PDF specifications through well-defined tools.

> [!IMPORTANT]
> **This is a specification *reference*, not a rule engine.**
> It retrieves and structures the text of ISO 32000 — clauses, tables, definitions, and
> `shall`/`should`/`may` requirements. It does **not** examine a PDF file, and it cannot tell
> you whether a document conforms to anything. Conformance verdicts come from
> [pdf-verify-mcp](https://github.com/shuji-bonji/pdf-verify-mcp)
> (`validate_conformance` / `evaluate_policy`).
>
> The distinction matters because three different things get conflated:
> **declaration** — a label the file wrote about itself ("I am PDF/A" in the metadata). Writing it is not evidence /
> **conformance** — whether the file actually meets the standard. There is no way to prove it in full; you can only find where it breaks the rules /
> **validation** — what a validator (veraPDF and the like) reports against the checks it implements. A pass means "this inspection did not fail", not "the file conforms to the standard".
> Reading a `shall` here tells you what the standard requires — not whether your file meets it.
>
> **A search that returns nothing means "cannot answer", not "no such requirement."**
> ISO 19005 (PDF/A) and ETSI PAdES are outside this corpus; see `list_specs` → `coverage.gaps`.

### What each PDF family server does — and does not do

| Server | Does | **Does not** |
|---|---|---|
| **pdf-spec-mcp** (this) | Search, retrieve and extract requirements from 17 PDF-related documents | **Is not a rule engine.** Does not define business rules, inspect PDF files, or validate schemas. ISO 19005 (PDF/A) is not part of the corpus |
| [pdf-reader-mcp](https://github.com/shuji-bonji/pdf-reader-mcp) | Extract text / tables / structure tree / fonts / annotations / images / signature *fields* | **Does not verify cryptography.** Does not read the incremental-update history, does not map object IDs to coordinates, does not OCR |
| [pdf-writer-mcp](https://github.com/shuji-bonji/pdf-writer-mcp) | Create, page operations, tagging, forms, annotations, metadata, attachments, PDF/A-3b scaffolding | **Does not sign.** Does not make the file meet the standard — it can write a *label*, not conformance |
| [pdf-verify-mcp](https://github.com/shuji-bonji/pdf-verify-mcp) | Conformance validation (delegated to veraPDF), cryptographic signature verification, tamper detection, policy verdicts | **Does not prove the file meets the standard** (it can only find where it breaks the rules). Does not vouch for the signer's identity. Does not judge whether the content is true |

> [!IMPORTANT]
> **PDF specification files are NOT included in this package.**
> You must obtain the PDF specification documents separately and place them in a local directory.
>
> **Download from:** [PDF Association — Sponsored Standards](https://pdfa.org/sponsored-standards/)
>
> See "[Setup](#setup)" for details.

## Features

- **Multi-spec support** — Auto-discovers and manages up to 17 PDF-related documents (ISO 32000-2, PDF/UA, Tagged PDF guides, etc.)
- **Structured content extraction** — Headings, paragraphs, lists, tables, and notes from any section
- **Full-text search** — Keyword search with section-aware context snippets
- **Requirements extraction** — Extracts normative language (shall / must / may) per ISO conventions
- **Definitions lookup** — Term definitions from Section 3 (Definitions)
- **Table extraction** — Multi-page table detection with header merging
- **Version comparison** — Diff PDF 1.7 vs PDF 2.0 section structures
- **Bounded-concurrency processing** — Parallel page processing for large documents
- **On-disk index cache** — The search index and the full requirements scan are built once per PDF and reused by every later process (ISO 32000-2: ~6 s → ~0.2 s)

## Architecture

```mermaid
graph LR
    subgraph Client["MCP Client"]
        LLM["LLM<br/>(Claude, etc.)"]
    end

    subgraph Server["PDF Spec MCP Server"]
        direction TB
        MCP["MCP Server<br/>index.ts"]

        subgraph Tools["Tools Layer"]
            direction LR
            T1["list_specs"]
            T2["get_structure"]
            T3["get_section"]
            T4["search_spec"]
            T5["get_requirements"]
            T6["get_definitions"]
            T7["get_tables"]
            T8["compare_versions"]
        end

        subgraph Services["Services Layer"]
            direction LR
            REG["Registry<br/>Auto-discovery"]
            LOADER["Loader<br/>LRU Cache"]
            SVC["PDFService<br/>Orchestration"]
            CMP["CompareService<br/>Version Diff"]
        end

        subgraph Extractors["Extractors"]
            direction LR
            OUTLINE["OutlineResolver<br/>TOC & Section Index"]
            CONTENT["ContentExtractor<br/>Structured Extraction"]
            SEARCH["SearchIndex<br/>Full-text Search"]
            REQ["RequirementExtractor"]
            DEF["DefinitionExtractor"]
        end

        subgraph Utils["Utils"]
            direction LR
            CACHE["LRU Cache"]
            CONC["Concurrency"]
            VALID["Validation"]
        end
    end

    subgraph PDFs["PDF Spec Files (obtained separately)"]
        direction LR
        PDF1["ISO 32000-2<br/>(PDF 2.0)"]
        PDF2["ISO 32000-1<br/>(PDF 1.7)"]
        PDF3["TS 32001–32005<br/>PDF/UA, etc."]
    end

    LLM <-->|"stdio / JSON-RPC"| MCP
    MCP --> Tools
    Tools --> Services
    Services --> Extractors
    Services --> Utils
    LOADER --> PDFs
    REG -->|"Filename pattern<br/>auto-discovery"| PDFs

    style Client fill:#e8f4f8,stroke:#2196F3
    style PDFs fill:#fff3e0,stroke:#FF9800
    style Tools fill:#e8f5e9,stroke:#4CAF50
    style Services fill:#f3e5f5,stroke:#9C27B0
    style Extractors fill:#fce4ec,stroke:#E91E63
    style Utils fill:#f5f5f5,stroke:#9E9E9E
```

### Layer Overview

| Layer          | Responsibility                                                                     |
| -------------- | ---------------------------------------------------------------------------------- |
| **Tools**      | MCP tool schema definitions & handlers (input validation)                          |
| **Services**   | Business logic (PDF registry, loader, orchestration)                               |
| **Extractors** | Information extraction from PDFs (TOC, content, search, requirements, definitions) |
| **Utils**      | Shared utilities (cache, concurrency, validation)                                  |

## Setup

### 1. Obtain PDF Specification Files

> [!WARNING]
> PDF specifications are **copyrighted documents** and are not included in this package.
> Download them from the sources below and place them in a local directory.

| Document                     | Source                                                                                          |
| ---------------------------- | ----------------------------------------------------------------------------------------------- |
| ISO 32000-2 (PDF 2.0)        | [PDF Association](https://pdfa.org/resource/iso-32000-pdf/)                                     |
| ISO 32000-1 (PDF 1.7)        | [Adobe (free)](https://opensource.adobe.com/dc-acrobat-sdk-docs/pdfstandards/PDF32000_2008.pdf) |
| TS 32001–32005, PDF/UA, etc. | [PDF Association — Sponsored Standards](https://pdfa.org/sponsored-standards/)                  |

All 17 files below are supported. You do not need all of them — place only the specs you need (at minimum, ISO 32000-2 is recommended).

```
pdf-specs/
│
│ ── Standards ─────────────────────────────
├── ISO_32000-2_sponsored_EC3.pdf          # iso32000-2  : PDF 2.0 EC3 (recommended; falls back to -ec2.pdf)
├── ISO_32000-2-2020_sponsored.pdf         # iso32000-2-2020 : PDF 2.0 original
├── PDF32000_2008.pdf                      # pdf17       : PDF 1.7 (for version comparison)
├── pdfreference1.7old.pdf                 # pdf17old    : Adobe PDF Reference 1.7
│
│ ── Technical Specifications (TS) ─────────
├── ISO_TS_32001-2022_sponsored_EC3.pdf    # ts32001     : Hash extensions (SHA-3)
├── ISO_TS_32002-2022_sponsored_EC3.pdf    # ts32002     : Digital signature extensions (ECC/PAdES)
├── ISO_TS_32003-2023_sponsored.pdf        # ts32003     : AES-GCM encryption
├── ISO-TS-32004-2024_sponsored.pdf        # ts32004     : Integrity protection
├── ISO-TS-32005-2023-sponsored.pdf        # ts32005     : Namespace mapping
│
│ ── PDF/UA (Accessibility) ────────────────
├── ISO-14289-1-2014-sponsored.pdf         # pdfua1      : PDF/UA-1
├── ISO-14289-2-2024-sponsored.pdf         # pdfua2      : PDF/UA-2
│
│ ── Guides ────────────────────────────────
├── Tagged-PDF-Best-Practice-Guide.pdf     # tagged-bpg  : Tagged PDF Best Practice
├── Well-Tagged-PDF-WTPDF-1.0.pdf          # wtpdf       : Well-Tagged PDF
├── PDF-Declarations.pdf                   # declarations: PDF Declarations
│
│ ── Application Notes ─────────────────────
├── PDF20_AN001-BPC.pdf                    # an001       : Black Point Compensation
├── PDF20_AN002-AF.pdf                     # an002       : Associated Files
└── PDF20_AN003-ObjectMetadataLocations.pdf # an003      : Object Metadata
```

### 2. Install

This package ships a CLI binary (`pdf-spec-mcp`) intended to be launched by an MCP client.
**You do not need to install it manually** — just point your MCP client to `npx @shuji-bonji/pdf-spec-mcp@latest` as shown in the next step.

If you want to run it directly from the shell (e.g. for debugging):

```bash
PDF_SPEC_DIR=/path/to/pdf-specs npx -y @shuji-bonji/pdf-spec-mcp@latest
```

Or install it globally (optional):

```bash
npm install -g @shuji-bonji/pdf-spec-mcp
PDF_SPEC_DIR=/path/to/pdf-specs pdf-spec-mcp
```

### 3. Configure MCP Client

#### Environment Variable

| Variable             | Description                                                           | Default                                    |
| -------------------- | --------------------------------------------------------------------- | ------------------------------------------ |
| `PDF_SPEC_DIR`       | Directory containing PDF specification files                          | (required)                                 |
| `PDF_SPEC_CACHE_DIR` | Where the on-disk index cache lives (see [Index cache](#index-cache)) | `${XDG_CACHE_HOME:-~/.cache}/pdf-spec-mcp` |
| `PDF_SPEC_CACHE`     | Set to `off` to neither read nor write the index cache                | on                                         |

#### Claude Desktop

Add to `claude_desktop_config.json`:

```json
{
  "mcpServers": {
    "pdf-spec": {
      "command": "npx",
      "args": ["-y", "@shuji-bonji/pdf-spec-mcp@latest"],
      "env": {
        "PDF_SPEC_DIR": "/path/to/pdf-specs"
      }
    }
  }
}
```

> [!IMPORTANT]
> **Use `@latest` (or pin a version).** `npx -y <pkg>` without a version keeps running whatever
> it cached the first time — `-y` only skips the install prompt, it does not check for updates.
> A bare specifier will happily run a months-old release. `@latest` makes npx check the registry
> on each start; pin `@0.4.0` instead if you want reproducibility.
> To clear a stale cache: `rm -rf ~/.npm/_npx`.

#### Cursor / VS Code

Add to `.cursor/mcp.json` or VS Code MCP settings:

```json
{
  "mcpServers": {
    "pdf-spec": {
      "command": "npx",
      "args": ["-y", "@shuji-bonji/pdf-spec-mcp@latest"],
      "env": {
        "PDF_SPEC_DIR": "/path/to/pdf-specs"
      }
    }
  }
}
```

## Index cache

Two operations walk every page of a specification: the first `search_spec` on a spec builds
its full-text index (ISO 32000-2, 1023 pages: about 6 s on a laptop), and `get_requirements`
without a `section` scans every section (about 11 s). Everything else opens only the pages it
needs and answers in well under a second.

Since 0.5.0 those two results are written to disk after the first build and read back by every
later process — an MCP client that starts one server per session no longer pays the build each
time. The second process answers the same `search_spec` in about 0.2 s and the full
requirements scan in about 0.02 s, from byte-for-byte the same index.

- **Location:** `${PDF_SPEC_CACHE_DIR:-${XDG_CACHE_HOME:-~/.cache}/pdf-spec-mcp}/v1/<version>/<spec>.<kind>.<sha256[0:16]>.json`.
  The whole 17-spec corpus is about 18 MB per package version.
- **Key:** package version, `pdfjs-dist` version, spec id, and the SHA-256 of the PDF. A
  replaced PDF, an upgraded server, or an upgraded pdfjs all miss and rebuild. Entries of
  older versions are left in place (another install may still use them); `--clear-cache`
  removes everything.
- **Failure is a miss, never an error:** an unreadable, truncated, or foreign file is rebuilt;
  an unwritable directory is reported once on stderr and the server carries on without a cache.
- **It is derived from *your* copy of the PDFs and stays on your machine.** It is not part of
  the package and must not be redistributed — the specifications are copyrighted.

Nothing about searching changes: the same in-memory structure is searched by the same code.
Only where it comes from (built vs. read) does.

### Pre-building the cache

The cache fills lazily, one spec at a time as tools touch it. To warm every spec up front — after
installing, after upgrading, or from cron — run the CLI (it uses the same code path as the tools,
processes specs sequentially, and exits):

```bash
PDF_SPEC_DIR=/path/to/pdf-specs npx -y @shuji-bonji/pdf-spec-mcp@latest --build-cache
#   --spec=iso32000-2,pdf17   only these specs
#   --force                   rebuild even when a valid entry exists
npx -y @shuji-bonji/pdf-spec-mcp@latest --cache-info     # directory, key, entries
npx -y @shuji-bonji/pdf-spec-mcp@latest --clear-cache    # remove the directory
```

A full build of the 17-spec corpus takes about a minute on a laptop.

## Available Tools

All tools accept an optional `spec` parameter to target a specific specification (default: `iso32000-2`).

| Tool               | Description                                                       |
| ------------------ | ----------------------------------------------------------------- |
| `list_specs`       | List all discovered PDF specifications with metadata              |
| `get_structure`    | Get section hierarchy (table of contents) with configurable depth |
| `get_section`      | Get structured content of a specific section                      |
| `search_spec`      | Full-text keyword search across a specification                   |
| `get_requirements` | Extract normative requirements (shall/must/may)                   |
| `get_definitions`  | Lookup term definitions                                           |
| `get_tables`       | Extract table structures from a section                           |
| `compare_versions` | Compare PDF 1.7 and PDF 2.0 section structures                    |

### `list_specs` — Discover Specifications

List all available specification documents. Use the returned IDs as the `spec` parameter in other tools.

```jsonc
// List all specs
{ }

// Filter by category
{ "category": "ts" }        // Technical specs only
{ "category": "pdfua" }     // PDF/UA only
{ "category": "guide" }     // Guide documents only
```

### `get_structure` — Table of Contents

Get the section hierarchy (TOC tree) of a specification.

```jsonc
// PDF 2.0 top-level sections only
{ "max_depth": 1 }

// Expand to 2 levels
{ "max_depth": 2 }

// TS 32002 (Digital Signatures) full structure
{ "spec": "ts32002" }

// PDF/UA-2 structure
{ "spec": "pdfua2", "max_depth": 2 }
```

### `get_section` — Section Content

Get structured content (headings, paragraphs, lists, tables, notes) of a specific section.

A parent section returns its entire subtree (its preamble followed by all subsections, in document order). Top-level clauses can be very large — prefer the most specific section number.

```jsonc
// PDF 2.0 Section 7.3.4.2 (Literal Strings)
{ "section": "7.3.4.2" }

// PDF 2.0 Annex A
{ "section": "Annex A" }

// TS 32002 Section 5
{ "spec": "ts32002", "section": "5" }

// PDF/UA-2 Section 8 (Tagged PDF)
{ "spec": "pdfua2", "section": "8" }
```

### `search_spec` — Full-text Search

Search across a specification with section-aware context snippets. The first call on a spec
builds its index (a few seconds); the index is then cached on disk (see [Index cache](#index-cache)).

```jsonc
// Search PDF 2.0 for "digital signature"
{ "query": "digital signature" }

// Limit results
{ "query": "font", "max_results": 5 }

// Search within TS 32002
{ "spec": "ts32002", "query": "CMS" }
```

### `get_requirements` — Normative Requirements

Extract normative requirements (shall / must / may) per ISO conventions.

```jsonc
// All requirements in section 12.8
{ "section": "12.8" }

// Only "shall" requirements
{ "section": "12.8", "level": "shall" }

// Only "shall not" requirements
{ "section": "7.3", "level": "shall not" }

// PDF/UA-2 requirements
{ "spec": "pdfua2", "section": "8", "level": "shall" }
```

### `get_definitions` — Term Definitions

Look up term definitions from Section 3 (Definitions).

```jsonc
// Search for "font" definitions
{ "term": "font" }

// List all definitions
{ }

// PDF/UA definitions
{ "spec": "pdfua2", "term": "artifact" }
```

### `get_tables` — Table Extraction

Extract table structures (headers, rows, captions) from a section. Multi-page tables are automatically merged.

```jsonc
// All tables in section 7.3.4.2 (Table 3 — Escape sequences)
{ "section": "7.3.4.2" }

// Specific table only (0-based index)
{ "section": "7.3.4.2", "table_index": 0 }

// TS spec tables
{ "spec": "ts32002", "section": "5" }
```

### `compare_versions` — Version Comparison

Compare section structures between PDF 1.7 (ISO 32000-1) and PDF 2.0 (ISO 32000-2). Uses title-based automatic matching to detect matched, added, and removed sections.

> [!NOTE]
> This tool requires both PDF 1.7 (`PDF32000_2008.pdf`) and PDF 2.0 files in `PDF_SPEC_DIR`.

```jsonc
// Diff section 12.8 (Digital Signatures)
{ "section": "12.8" }

// Compare all top-level sections
{ }
```

## Supported Specifications

The server auto-discovers PDF files in `PDF_SPEC_DIR` by filename pattern matching:

| Category           | Spec IDs                                             | Documents                                               |
| ------------------ | ---------------------------------------------------- | ------------------------------------------------------- |
| **Standard**       | `iso32000-2`, `iso32000-2-2020`, `pdf17`, `pdf17old` | ISO 32000-2 (PDF 2.0), ISO 32000-1 (PDF 1.7)            |
| **Technical Spec** | `ts32001` – `ts32005`                                | Hash, Digital Signatures, AES-GCM, Integrity, Namespace |
| **PDF/UA**         | `pdfua1`, `pdfua2`                                   | Accessibility (ISO 14289-1, 14289-2)                    |
| **Guide**          | `tagged-bpg`, `wtpdf`, `declarations`                | Tagged PDF, Well-Tagged PDF, Declarations               |
| **App Note**       | `an001` – `an003`                                    | BPC, Associated Files, Object Metadata                  |

## Directory Structure

```
src/
├── index.ts              # Entry point: MCP server on stdio, or the cache CLI
├── cli.ts                # --build-cache / --clear-cache / --cache-info
├── config.ts             # Configuration & spec patterns
├── errors.ts             # Error hierarchy (PDFSpecError → sub-classes)
├── services/
│   ├── pdf-registry.ts       # Auto-discovery of PDF files
│   ├── pdf-loader.ts         # PDF loading with LRU cache
│   ├── pdf-service.ts        # Orchestration layer
│   ├── index-store.ts        # On-disk cache for the search / requirements indexes
│   ├── compare-service.ts    # Version comparison
│   ├── outline-resolver.ts   # Section index builder
│   ├── content-extractor.ts  # Structured content extraction
│   ├── search-index.ts       # Full-text search index
│   ├── requirement-extractor.ts
│   └── definition-extractor.ts
├── tools/
│   ├── definitions.ts    # MCP tool schemas
│   └── handlers.ts       # Tool implementations
├── types/
│   └── index.ts          # Shared type definitions
└── utils/
    ├── concurrency.ts    # mapConcurrent (bounded Promise.all)
    ├── text.ts           # Text normalization
    ├── cache.ts          # LRU cache
    ├── file-hash.ts      # SHA-256 of a PDF (index cache key)
    ├── validation.ts     # Input validation
    └── logger.ts         # Structured logger
```

## Development

```bash
git clone https://github.com/shuji-bonji/pdf-spec-mcp.git
cd pdf-spec-mcp
npm install
npm run build

# Unit tests
npm run test

# E2E tests (requires PDF files in ./pdf-spec/)
npm run test:e2e

# Lint & format
npm run lint
npm run format:check
```

## License

[MIT](LICENSE)

TDQS

A4.2/5.0

Scored across 8 tools

Disambiguation5/5

Each tool targets a distinct aspect of the PDF specification: version comparison, definitions, requirements, sections, structure, tables, listing specs, and searching. No overlap in purpose.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern in snake_case (e.g., get_section, search_spec). There is no mixing of conventions.

Tool Count5/5

With 8 tools, the set is well-scoped for exploring and comparing PDF specifications. Each tool serves a clear function without being excessive or insufficient.

Completeness5/5

The tool surface covers common operations for a reference document: listing, searching, retrieving structure, content, tables, definitions, requirements, and version comparison. No obvious gaps for its read-only domain.

Maintenance

ActivityActive
ResponsivenessSlow