@shuji-bonji/pdf-reader-mcp
This is a specialized MCP server for reading, inspecting, and analyzing the internal structure of PDF documents across 18 tools.
Basic Operations
get_page_count— Retrieve total page countget_metadata— Extract full metadata (title, author, PDF version, creation date, tagged/encrypted/signature flags, etc.)read_text— Extract text with Y-coordinate reading order; supports multi-column reordering and whitespace compaction for Japanese formssearch_text— Case-insensitive full-text search with surrounding context, page filtering, and configurable result limitsread_images— Extract embedded images as base64-encoded data with metadata (dimensions, color space)read_url— Fetch and process remote PDFs from HTTP/HTTPS URLs (max 50MB, 30s timeout)summarize— Quick overview combining metadata, text presence, image count, and a first-page text preview
Structure Inspection
inspect_structure— Examine PDF internal object structure: catalog entries, page tree, object statistics, and encryption statusinspect_tags— Analyze the Tagged PDF structure tree: hierarchy, roles, nesting depth, and element distributioninspect_fonts— List all fonts with type, encoding, embedded/subset status, and pages usedinspect_annotations— Categorize all annotations by subtype (Link, Widget, Highlight, etc.) with per-page breakdowninspect_signatures— Examine digital signature field structure (structural only, no cryptographic verification)extract_structured_text— Extract text from Tagged PDFs in logical content order with structure labels (H1, P, Table…) and optional bounding box coordinatesextract_tables— Extract tables from Tagged PDFs as Markdown or structured JSON, handling cross-page tableslocate_objects— Map PDF object numbers to their page and bounding rectangle coordinates
Validation & Analysis
validate_tagged(deprecated) — Validate Tagged PDF / PDF/UA structure requirementsvalidate_metadata(deprecated) — Validate PDF metadata completeness against PDF/A and PDF/UA best practicescompare_structure— Structural diff between two PDFs: page count, version, encryption, tags, object counts, fonts, page dimensions, and catalog entries
PDF Reader MCP Server
English | 日本語
An MCP (Model Context Protocol) server specialized in deciphering PDF internal structures.
While typical PDF MCP servers are thin wrappers for text extraction, this project focuses on reading and analyzing the internal structure of PDF documents. Pair it with pdf-spec-mcp for specification-aware structural analysis and validation.
PDF family
Server | Role |
PDF specification knowledge (ISO 32000, PDF/A, PDF/UA) | |
pdf-reader-mcp (this) | Read and inspect PDF internal structure — what is in a PDF |
Authenticity verification — whether it is genuine: cryptographic signature verification, tamper detection, PAdES level, PDF/A validation, encrypted-PDF decryption |
pdf-reader-mcp inspects signature structure (inspect_signatures); for cryptographic signature verification, trust/revocation evaluation, and PDF/A conformance validation, use pdf-verify-mcp.
Features
18 tools organized into three tiers:
Tier 1: Basic Operations
Tool | Description |
| Lightweight page count retrieval |
| Full metadata extraction (title, author, PDF version...) |
| Text extraction with Y-coordinate reading order (opt-in |
| Full-text search with surrounding context. Searches the same text |
| Image extraction as base64 with metadata |
| Fetch and process remote PDFs from URLs |
| Quick overview report (metadata + text + image count) |
Tier 2: Structure Inspection
Tool | Description |
| Object tree and catalog dictionary analysis |
| Tagged PDF structure tree visualization |
| Font inventory (embedded/subset/type detection) |
| Annotation listing (categorized by subtype) |
| Digital signature field structure analysis |
| Tagged PDF text in logical content order (ISO 32000-2 §14.8.2.5), each piece labelled with its structure type ( |
| Tagged PDF |
| Object number → page and rectangle, in the coordinate form pdf-writer-mcp |
Tier 3: Validation & Analysis
Tool | Description |
| Deprecated — PDF/UA pass/fail belongs to pdf-verify-mcp |
| Deprecated — same migration path as above. Kept until the next major |
| Structural diff between two PDFs (properties + fonts) |
Related MCP server: PDF Reader MCP Server
Installation
npx (recommended)
npx @shuji-bonji/pdf-reader-mcp@latestClaude Desktop
Add to your claude_desktop_config.json:
{
"mcpServers": {
"pdf-reader-mcp": {
"command": "npx",
"args": ["-y", "@shuji-bonji/pdf-reader-mcp@latest"]
}
}
}Use
@latest. The-yflag innpx -y <pkg>only skips the install prompt — it does not check for updates. Without@latest, npx keeps running whichever version it cached the first time, so new releases never reach you. If you suspect you are on a stale version, runrm -rf ~/.npm/_npxand restart your client.
Claude Code
claude mcp add pdf-reader-mcp -- npx -y @shuji-bonji/pdf-reader-mcp@latestFrom Source
git clone https://github.com/shuji-bonji/pdf-reader-mcp.git
cd pdf-reader-mcp
npm install
npm run buildUsage Examples
Get Page Count
get_page_count({ file_path: "/path/to/document.pdf" })
→ 42Search Text
search_text({
file_path: "/path/to/spec.pdf",
query: "digital signature",
pages: "1-20",
max_results: 10
})
→ Found 5 matches (page 3, 7, 12, 15, 18)Summarize
summarize({ file_path: "/path/to/document.pdf" })
→ | Pages | 42 |
| PDF Version | 2.0 |
| Tagged | Yes |
| Signatures | No |
| Images | 15 |Validate Tagged Structure (PDF/UA)
validate_tagged({ file_path: "/path/to/document.pdf" })
→ ✅ [TAG-001] Document is marked as tagged
✅ [TAG-002] Structure tree root exists
⚠️ [TAG-004] Heading hierarchy has gaps: H1, H3
❌ [TAG-005] Document has 3 image(s) but no Figure tagsValidate Metadata
validate_metadata({ file_path: "/path/to/document.pdf" })
→ ✅ [META-001] Title: "Annual Report 2025"
⚠️ [META-002] Author is missing
✅ [META-006] PDF version: 2.0Compare Structure
compare_structure({
file_path_1: "/path/to/v1.pdf",
file_path_2: "/path/to/v2.pdf"
})
→ | Page Count | 10 | 12 | ❌ |
| PDF Version | 1.7 | 2.0 | ❌ |
| Tagged | true | true | ✅ |Extract Structured Text (Tagged PDF, logical order)
extract_structured_text({ file_path: "/path/to/report.pdf", pages: "1-2" })
→ # Structured Text
- **Tagged**: Yes / **Language**: en-US / **Elements**: 7
## Logical Content Order
- **Document** (pages 1–2)
- **H1** (page 1) — Quarterly Report
- **P** (pages 1–2) — This paragraph begins on page one and continues on page two.
- **Table** (page 1)
| Item | Amount |
|---|---|
| Sales | 100 |
- **Figure** (page 2) — *alt:* A bar chart of salesThis answers "what is the text of the H1?" — which read_text (flat,
coordinate order) cannot. Order is a depth-first traversal of the structure
tree (ISO 32000-2 §14.8.2.5). /ActualText replaces the glyphs (§14.9.4),
/Alt stays out of the body text (§14.9.3), list labels are reported
separately, and an element spanning a page break stays ONE element. Use
roles: ["H1", "H2"] to pull an outline. Untagged PDFs return
isTagged: false with a reason — nothing is guessed from coordinates.
include_bbox: where the element is drawn
extract_structured_text({ file_path: "/doc.pdf", roles: ["P"], include_bbox: true })
→ - **P** (page 1) — Measured paragraph
- *bbox* p1 `(50.0, 297.5, 161.4, 308.6)` — text-extent
- **Figure** (page 1)
- *bbox* p1 `(50.0, 150.0, 110.0, 190.0)` — layout-attribute-bboxRectangles are in PDF default user space (origin bottom-left, pt, normalised) —
exactly what pdf-writer-mcp
add_annotation takes, so "annotate this paragraph" needs no coordinate
conversion in between. /Rotate and a shifted /CropBox do not move them.
basis says how strong the claim is, and the two are different in kind:
| What it is |
| The |
| Measured from the element's text: baseline origin plus the font's ascent/descent. The line box, not the glyph outlines. Images and vector art contribute nothing |
An element spanning pages gets ONE RECTANGLE PER PAGE — merging them would put
content on a page it is not on. An element with no rectangle says why in
boxNote rather than returning a zero-sized one.
A declaration is reported as-is, and cross-checked. Files state nonsense:
the cover Figure of both Well-Tagged PDF 1.0 and the Tagged PDF Best Practice
Guide declares /BBox [-32768 -32768 32767 32767] — int16 sentinels where a
rectangle should be — and PDF32000_2008 has 131 of its 545 declarations reaching
past the page edge. Since this output is meant to go straight into
add_annotation, a declaration is checked against the page box (§7.7.3.3) and
against the element's own text; either contradiction is reported in boxNote,
with the rectangle still returned unaltered.
Measured against independent ground truth: on Well-Tagged PDF (WTPDF) 1.0, the
166 Link structure elements were compared with the 173 Link annotation
/Rect values the producer placed for the same links — median IoU 0.972,
none disjoint.
Extract Tables (Tagged PDF)
extract_tables({ file_path: "/path/to/kaisei-tsutatsu.pdf", pages: "1" })
→ # Extracted Tables
- **Tagged**: Yes / **Pages Scanned**: 1 / **Tables Found**: 1
## Table 1 — Page 1
| 改正後 | 改正前 |
| --- | --- |
| …第2条第 16 項《定義》… | …第2条第 15 項《定義》… |A table that continues across a page break is reported as ONE table —
pages is an array (e.g. ## Table 3 — Pages 5–7), and a table touching
the requested pages range is returned whole. Cell text honours
/ActualText replacements (as do read_text and search_text since #18).
Untagged PDFs return an empty result with a
note recommending the column-aware fallback below.
Read Untagged Multi-Column PDF
read_text({ file_path: "/path/to/older-shinkyu.pdf", split_columns: 2 })
→ // Plain Y-sort would interleave columns:
// "改正後セル1 改正前セル1\n 改正後セル2 改正前セル2..."
//
// With split_columns: 2 the left column is emitted first, then the right:
// "改正後セル1\n改正後セル2\n…\n\n改正前セル1\n改正前セル2\n…"Use split_columns: 2 | 3 for untagged multi-column PDFs. For Tagged
PDFs with proper <Table> markup, extract_tables (above) is preferred.
Compact Whitespace (Japanese Forms)
read_text({ file_path: "/path/to/form.pdf", compact_whitespace: true })
→ // Original PDF uses U+3000 fullwidth space as visual indentation:
// " ( ) 自 年 月 日 法 有 ( 年 月 日) 有 有"
//
// With compact_whitespace: true:
// "( ) 自 年 月 日 法 有 ( 年 月 日) 有 有"
//
// Empirically reduces character count by ~40% on form PDFs.compact_whitespace is orthogonal to split_columns — both can be combined.
Tech Stack
TypeScript + MCP TypeScript SDK
pdfjs-dist (Mozilla) — text/image extraction, tag tree, annotations
pdf-lib — low-level object structure analysis
Vitest — unit + E2E testing (171 tests)
Biome — linting + formatting
Zod — input validation
Testing
npm test # Run all tests (unit: 39 tests)
npm run test:e2e # E2E tests only (132 tests)
npm run test:watch # Watch modeArchitecture
pdf-reader-mcp/
├── src/
│ ├── index.ts # MCP Server entry point
│ ├── constants.ts # Shared constants
│ ├── types.ts # Type definitions
│ ├── tools/
│ │ ├── tier1/ # Basic tools (7)
│ │ ├── tier2/ # Structure inspection (6)
│ │ ├── tier3/ # Validation & analysis (3)
│ │ └── index.ts # Tool registration
│ ├── services/
│ │ ├── pdfjs-service.ts # pdfjs-dist wrapper (parallel page processing)
│ │ ├── pdflib-service.ts # pdf-lib wrapper
│ │ ├── validation-service.ts # Validation & comparison logic
│ │ └── url-fetcher.ts # URL fetching
│ ├── schemas/ # Zod validation schemas
│ └── utils/
│ ├── pdf-helpers.ts # PDF utilities (page range parsing, file I/O)
│ ├── batch-processor.ts # Batch processing for large PDFs
│ ├── formatter.ts # Output formatting
│ └── error-handler.ts # Error handling
└── tests/
├── tier1/ # Unit tests
└── e2e/ # E2E tests (9 suites, 132 tests)Error Contract (houki-hub family)
Since v0.6.0, this MCP returns structured errors that follow the houki-hub family error contract, sharing a unified code vocabulary across the family. Combined with houki-egov-mcp / houki-nta-mcp, an LLM or Skill layer can interpret errors with consistent logic.
docs/ERROR-CODES.md— error code vocabulary (houki-research-skill)docs/ERROR-HANDLING.md— handling policy / next_actions templates
Implementation is independent — no dependency on houki-abbreviations or other family packages. The reference implementation is houki-egov-mcp/src/errors.ts; pdf-reader-mcp's local definition is in src/errors.ts.
On error, every tool returns isError: true and the JSON-stringified LawServiceError in content[0].text:
{
"error": "The file does not appear to be a valid PDF.",
"code": "INVALID_PDF",
"hint": "ファイルが破損していないか確認してください。",
"next_actions": [
{
"action": "inspect_structure",
"reason": "PDF が壊れている可能性があります。Catalog / Pages 等の構造を確認してください"
}
],
"detail": { "cause": "Invalid PDF structure" }
}Codes used by pdf-reader-mcp
code | 用途 |
| パス・URL・ページ範囲などクライアント側引数の不正 |
| ファイル未存在 (ENOENT) |
| PDF として不正・破損 |
| 暗号化 PDF (現状未対応) |
| サポート外の PDF 機能 |
| 50MB 上限超過 (pdf-reader 固有) |
| URL fetch の HTTP エラー (4xx/5xx) |
| リモート取得タイムアウト |
| DNS / 接続失敗 |
| パーミッション拒否を含むその他バグ |
Migration note (v0.5.x → v0.6.0)
旧 v0.5.x までは content[0].text に Error: ...\n\nSuggestion: ... という人間可読文字列を入れていました。v0.6.0 では同じ場所に JSON 文字列 が入ります。LLM 側でテキスト解釈に依存していた場合は、JSON.parse(content[0].text) での解釈に切り替えてください。isError: true フラグで構造化エラーかどうかを判定できます。
Pairing with pdf-spec-mcp
pdf-spec-mcp provides PDF specification knowledge (ISO 32000-2, etc.). With both servers enabled, an LLM can perform specification-aware workflows:
summarize— get a PDF overviewinspect_tags— examine the tag structurepdf-spec-mcp
get_requirements— fetch PDF/UA requirementsvalidate_tagged— check conformancecompare_structure— diff before/after fixes
License
MIT
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- Flicense-qualityDmaintenanceA PDF processing server that extracts text via normal parsing or OCR, and retrieves images from PDF files through the MCP protocol with a built-in web debugger.Last updated36
- FlicenseAquality-maintenanceA Model Context Protocol server that extracts and processes content from PDF documents, providing text extraction, metadata retrieval, page-level processing, and PDF validation capabilities.Last updated41
- AlicenseAqualityCmaintenanceAn MCP server for reading, rendering, and searching PDF files, specifically optimized for LLMs to extract text, tables, and technical diagrams. It enables metadata retrieval, multi-format text extraction, and page-to-image rendering using PyMuPDF.Last updated552MIT
- Flicense-qualityDmaintenanceMCP server for extracting text from PDF files, supporting local files and URLs.Last updated
Related MCP Connectors
An MCP server for deep research or task groups
MCP server for ScanMalware.com URL scanning, malware detection, and analysis.
MCP server for network documentation, generated by doc2mcp.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/shuji-bonji/pdf-reader-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server