pdfcheck
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@pdfcheckrun pdf_qa on /tmp/report.pdf"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
pdfcheck-mcp
MCP server that validates PDF page-tree structure. Two tools, stdio transport, zero native dependencies.
Why
Most PDF checks stop at "does a stream object exist". That misses the real failure mode: a PDF can have 24 stream objects and still be unreadable because the page tree never points at them. We shipped 13 broken product PDFs this way. Every page dictionary was orphaned, /Contents pointed at font objects, and the stream-count QA passed anyway.
This server walks the object graph instead:
every
/Kidsentry resolves to a/Type /Pageobjectevery page
/Contentspoints to a stream objectno orphan
/Type /Pageobjects unreachable from/Kidsreachable page count matches the
/Countclaimblank-page check: inflated content streams contain at least one
Tj/TJtext op
It handles classic (uncompressed) xref tables, which covers output from typical PDF generators. Compressed xref streams are rejected with a clear error rather than a false pass.
Related MCP server: Visual Document Forensics MCP Server
Install
npm install -g @jayjex/pdfcheck-mcpOr run from GitHub without npm:
npx -y github:jayjex/pdfcheck-mcpClaude Desktop config
{
"mcpServers": {
"pdfcheck": {
"command": "npx",
"args": ["-y", "@jayjex/pdfcheck-mcp"]
}
}
}Tools
pdf_qa(path)
Full report for one PDF:
{
"file": "/tmp/report.pdf",
"ok": false,
"problems": [
"contents-not-stream=3 [{\"page\":7,\"reason\":\"obj 5 not stream (type Font)\"}]",
"orphan-pages=1 [13]",
"count-mismatch claimed=4 reachable=3"
],
"report": {
"size_claim": 4, "kids": 4, "kids_ok": 3,
"contents_ok": 0, "contents_bad": 3,
"zero_tj": [], "orphans": [13], "xref_entries": 14, "size_field": 14
}
}pdf_batch(dir)
Flat scan of a directory, one pdf_qa per *.pdf (max 200 files, 50 MB each). Returns totals, a per-file kids/contents/orphans summary, and the full problem list for every broken file.
Local development
git clone https://github.com/jayjex/pdfcheck-mcp
cd pdfcheck-mcp && npm install
node test/smoke.mjstest/fixtures/ holds two healthy PDFs and one deliberately broken (broken-wiring.pdf, a page-tree pointer-shift bug). The smoke test asserts 2 pass, 1 fails with the exact graph problems.
License
MIT
This server cannot be installed
Maintenance
Related MCP Connectors
Build and validate IndexNow payloads, check key files, and diff sitemaps into a submission list.
Parse WebVTT, SRT, or TTML for conformance, timing, overlaps, line length, and reading speed.
PDF URLs to per-page text, tables as rows, Markdown, metadata and OCR for scanned pages.
Validate HTML/CSS, audit SEO and JSON-LD, check links, and capture responsive screenshots.
Related MCP Servers
AlicenseBqualityDmaintenanceEnables low-level PDF object-tree inspection and debugging through natural language, with lazy loading for token efficiency.24MIT- FlicenseNot gradedqualityBmaintenanceEnables deterministic visual and structural analysis of PDF and DOCX documents, extracting measurable evidence such as blur, OCR confidence, and image anomalies for auditable forensic workflows.1-
- AlicenseNot gradedqualityAmaintenanceValidates Excel (.xlsx) workbooks against reviewed, lockable acceptance contracts, catching errors and outputting structured issues that agents can repair.1MIT
- FlicenseNot gradedqualityCmaintenanceInspects local PDF files, extracting text, counting pages, and performing OCR on scanned PDFs using standard libraries.-