Skip to main content
Glama

pdf-library-mcp

A local-first MCP server for mathematical PDFs, scanned notes, and technical documents.

Import a PDF once, turn it into searchable Markdown, cache it permanently, and let AI clients retrieve only the pages or sections they actually need.

Built for mathematics, physics, computer science, English and Greek material, including scanned and handwritten lecture notes.

PDF
 │
 ▼
Hash + inspect
 │
 ▼
Fast extraction
 │
 ├── good enough ───────────────────────────────┐
 │                                             │
 └── suspicious / scanned / math-heavy         │
                  │                            │
                  ▼                            │
             Marker OCR                        │
                  │                            │
                  ▼                            │
       optional MLX-VLM correction             │
                  │                            │
                  └──────────────┬─────────────┘
                                 ▼
                         canonical Markdown
                                 │
                         cache + index
                                 │
                                 ▼
                               MCP

The goal is simple:

Do the expensive document work once. Retrieve only what the agent needs afterward.

A 500-page textbook should not cost 500 pages of context every time you ask a question about it.


Why this exists

There are already excellent PDF extraction tools.

The missing piece is that they are mostly extractors, not libraries.

Project

What it gives

What is still missing

PyMuPDF4LLM

Very fast local Markdown extraction, layout and tables

Weak mathematical extraction, no persistent library or search

Marker

Strong LaTeX, reading order and OCR

Much slower, model-heavy, no persistent retrieval layer

MinerU

Strong formula-aware document parsing

Still primarily an extractor

Generic PDF MCP servers

Page retrieval and search

Usually not designed around mathematical documents, OCR quality or page-level correction

pdf-library-mcp sits above the extraction engines.

It remembers what has already been processed, indexes the result, identifies bad pages, upgrades only those pages, and exposes the library to an AI client through MCP.

It adds:

  • Content-addressed caching

  • Fast and high-quality extraction tiers

  • Per-page upgrades

  • Mathematics-aware quality checks

  • Greek-aware search

  • OCR repair

  • Visual OCR review

  • Local MLX-VLM OCR correction

  • Page-level provenance and correction history

  • Token-budgeted retrieval

  • Original-page image access when OCR cannot be trusted

The extraction engines remain replaceable.


Core idea

A PDF is identified by the SHA-256 hash of its bytes.

Import the exact same PDF twice and the second import performs:

no extraction
no OCR
no model inference

The existing cached document is returned instead.

Storage is page-based, which means one bad page can be reprocessed without touching the other 800 pages in the book.


Extraction pipeline

There are now three levels of document processing.

Related MCP server: OpenPapers MCP

1. Fast extraction

Engine: PyMuPDF4LLM

Designed to make a document searchable quickly.

It handles:

  • prose

  • headings

  • tables

  • basic layout

  • text-layer PDFs

It is extremely fast, but mathematical PDFs expose an important weakness:

Display equations can disappear completely.

The resulting Markdown may still look perfectly valid, which makes the failure difficult to detect by inspecting the extracted text alone.


2. High-quality extraction

Engine: Marker

Marker is used when a page deserves heavier processing.

It provides:

  • much stronger LaTeX extraction

  • OCR for scanned pages

  • better mathematical layout

  • better reading order

  • block-level bounding boxes

Import does not blindly run Marker over every page.

Instead:

fast extraction
      │
      ▼
 quality gate
      │
      ├── page looks good → keep fast result
      │
      └── page looks suspicious → candidate for reprocessing

A single page can then be upgraded with:

pdf-library reprocess real-analysis --pages 243

The rest of the document remains untouched.


3. Local vision OCR correction

For difficult scanned or handwritten material, Marker is not always enough.

The library sends scanned/OCR pages to a local Vision Language Model running through MLX-VLM after Marker. The image is authoritative and the proposed Markdown must pass deterministic validation before replacing the Marker text.

This is particularly useful for:

  • handwritten notes

  • Greek handwriting

  • mathematical notation

  • inverse functions

  • hyperbolic functions

  • superscripts and subscripts

  • OCR errors that are impossible to detect from text alone

The pipeline becomes:

Marker transcription
        +
original page image
        │
        ▼
     MLX-VLM
        │
        ▼
vision comparison
        │
        ▼
corrected Markdown + LaTeX

The vision model is instructed to transcribe, not solve or rewrite the mathematics.

For example, it is explicitly told to distinguish:

sinh  cosh  tanh  coth  sech  csch

from:

sin   cos   tan   cot   sec   csc

and preserve notation such as:

\sinh^{-1}(x)

rather than silently changing what appears on the page.


MLX-VLM on Apple Silicon

The recommended vision backend on Apple Silicon is MLX-VLM.

The MCP server talks to its local OpenAI-compatible API, so the vision backend is isolated from the rest of the library.

Install

Create a separate environment if desired:

python3 -m venv ~/.venvs/pdf-vision
source ~/.venvs/pdf-vision/bin/activate

pip install -U mlx-vlm

Start a local model:

python -m mlx_vlm.server \
  --model mlx-community/Qwen3-VL-8B-Instruct-4bit

For an Apple Silicon Mac with limited unified memory, a quantized 7B/8B vision model is a sensible starting point.


Configure the library

Add to config.toml:

[vision]
enabled = true

base_url = "http://127.0.0.1:8080/v1"
model = "mlx-community/Qwen3-VL-8B-Instruct-4bit"

timeout_seconds = 300
render_scale = 3.0

temperature = 0.0
max_tokens = 4096

apply_min_confidence = 0.80

The high-resolution renderer used for vision OCR is separate from the token-budgeted renderer used when sending images through MCP.

That distinction is intentional:

MCP image
→ optimized for context/token cost

Vision OCR image
→ optimized for transcription accuracy

OCR correction safety

Vision models are useful, but they are not allowed to silently rewrite the library.

Every vision correction records:

original Markdown
corrected Markdown
model
confidence
status
warnings
whether it was applied
timestamp

The original OCR is therefore preserved.

A correction can be generated without applying it:

correct_ocr_page(
    document="notes.pdf",
    page=4,
    apply=false
)

or for several pages:

correct_ocr_pages(
    document="notes.pdf",
    pages=[1, 2, 3, 4],
    apply=false
)

After inspection:

correct_ocr_pages(
    document="notes.pdf",
    pages=[1, 2, 3, 4],
    apply=true
)

A correction is only automatically accepted when it satisfies the configured confidence threshold.

Low-confidence pages remain available for review instead.


Why OCR needs visual verification

Some errors cannot be detected from the extracted text.

Consider an underbrace annotation.

A handwritten or typeset expression with:

f(x)
g'(x)

written underneath terms may be interpreted by OCR as a fraction.

The resulting output can be perfectly valid LaTeX while representing completely different mathematics.

Text-only validation cannot reliably detect that.

The original page remains the source of truth.


Visual OCR review

Scan-derived pages have a separate review workflow.

An AI client can request:

ocr_review_queue

and then inspect one page with:

review_ocr_page

The tool returns:

  • the original page image

  • the saved Markdown transcription

The client can then record:

approved
needs_correction
unreadable

using:

record_ocr_review

Manual visual-review verdicts are audit records.

They do not silently modify the transcription.

A new extraction that changes the page returns it to the review queue.


Mathematics-aware quality detection

One of the most damaging extraction failures is also one of the least obvious:

A display equation disappears and leaves behind a blank line.

The resulting Markdown contains no malformed LaTeX and no obvious error.

To detect this, the library inspects the fonts used by the original PDF page.

Pages containing TeX mathematical fonts such as:

CMEX
AMS symbol fonts
OpenType mathematical fonts

but producing no display-math block can be flagged as:

display_math_missing

This allows the library to identify pages that deserve high-quality reprocessing even when their extracted Markdown appears superficially valid.


Greek is a first-class language

Greek technical documents introduce problems that ordinary Unicode search does not solve.

SQLite's unicode61 tokenizer does not stem Greek words.

For example:

μερικά κλάσματα

should still find:

μερικών κλασμάτων

The index therefore performs:

  • accent folding

  • final-sigma normalization

  • lightweight Greek stemming

  • identical normalization of queries

LaTeX is removed from the search index while remaining untouched in the readable Markdown.


Search that survives OCR

OCR corruption is not the same problem as grammatical inflection.

The library therefore searches in stages.

Stage

Behaviour

Result label

Exact

All normalized/stemmed words present

exact

Partial

At least one query word present

partial

Approximate

Character-trigram similarity

approximate

Search stops at the first stage that produces useful results.

Approximate matches receive a much lower score than real lexical matches, so OCR-tolerant search can rescue a failed query without outranking correct results.


Greek mathematical notation repair

Greek mathematical material frequently uses localized trigonometric names:

ημ   → sine
συν  → cosine
εφ   → tangent

OCR creates two recurring problems.

Character confusion

Handwriting may turn:

συνx

into something resembling:

60vx

Incorrect mathematical interpretation

Even correctly recognized characters may become separate variables:

ημx

can be emitted as:

\eta \mu x

which means the product:

η · μ · x

rather than a function.

Because the set of Greek mathematical function names is small and closed, these cases can be repaired deterministically.

Repairs are applied inside mathematics only so ordinary Greek words are not modified.

Existing documents can be repaired without rerunning OCR:

pdf-library repair real-analysis

Handwritten Greek OCR repair

Several deterministic OCR failures are also handled.

Examples include:

  • words split across lines

  • duplicated word tails after line breaks

  • the Greek article η being recognized as Latin h

  • known mathematical-function substitutions

Other handwriting errors are deliberately not guessed.

Greek handwriting frequently confuses characters such as:

γ ↔ χ
η ↔ υ
σ ↔ δ

Those cases are better handled by:

  • approximate search

  • visual review

  • the optional MLX-VLM correction layer

rather than aggressive automatic replacement.


Lecture-note structure

Handwritten lecture notes often have no Markdown headings.

They still contain semantic structure such as:

Παράδειγμα
Λύση
Περίπτωση 2
Βήμα 3
Θεώρημα 2.5

The chunker recognizes these markers and can use them as both:

  • chunk headings

  • semantic chunk types

This makes queries such as:

search --type solution

possible even when the original notes have no formal document structure.

OCR-damaged near-matches are tolerated with safeguards against accidentally turning ordinary verbs into headings.


The image escape hatch

OCR will never be perfect.

For that reason the library can expose the original page to the AI client.

Images are:

  • opt-in

  • never returned by ordinary search

  • rendered according to a token budget

  • optionally cropped to a single detected block

For example:

get_page_image(page=4)

returns the page.

But:

get_page_image(page=4, block=2)

can return only one equation.

That is often both cheaper and sharper than sending the whole page.

Marker's block bounding boxes make this possible.


Performance

The expensive parts of the pipeline are third-party extraction and model inference.

Everything else is intentionally lightweight.

The important optimization is not shaving milliseconds from SQLite.

It is this:

Extraction should never happen twice unless the user explicitly asks for it.

A cached document can be reopened and searched without running:

  • PyMuPDF extraction

  • Marker

  • OCR

  • MLX-VLM

Reindexing also does not re-extract PDFs.

Text repair modifies cached text without rerunning OCR.


Installation

Requires Python 3.11.

python3.11 -m venv .venv

.venv/bin/pip install -e .

.venv/bin/pdf-library doctor

Marker support

The high-quality extraction tier is optional.

.venv/bin/pip install -e '.[marker]'

brew install llama.cpp

The first Marker run may download several gigabytes of models.

pdf-library doctor checks that the required dependencies are available.


Optional Greek spell repair

For the complete inflection-aware deterministic repair mode, install:

.venv/bin/pip install -e '.[spellcheck]'

Place:

Greek.aff
Greek.dic

at:

<library root>/dictionaries/Greek.*

or configure another path:

[repair]

greek_lexicon = "/absolute/path/to/greek-words.txt"
greek_hunspell = "/absolute/path/to/Greek"

Corrections are deliberately conservative.

A word is changed only when the configured repair system finds a sufficiently constrained candidate.

Mathematical spans are not modified by prose repair.


Command line

# Import one document
pdf-library import ~/books/real-analysis.pdf

# Import a directory
pdf-library import ~/books/

# List documents
pdf-library list

# Search
pdf-library search "dominated convergence"

# Read specific pages
pdf-library page real-analysis 243 244

# Retrieve a section
pdf-library section real-analysis "Dominated Convergence"

# Inspect quality problems
pdf-library report real-analysis --problems-only

# Upgrade one page with Marker
pdf-library reprocess real-analysis --pages 243

# Apply deterministic repair without OCR
pdf-library repair real-analysis

# Inspect detected page blocks
pdf-library blocks real-analysis 243

# Render one block
pdf-library image real-analysis 243 --block 2

# Rebuild indexes without re-extracting
pdf-library reindex --all

# Environment and dependency checks
pdf-library doctor

Proofreading

The repository contains:

tools/proofread.py <document>

It generates a self-contained HTML proofreader showing:

original page | extracted Markdown

with LaTeX rendered.

For extraction quality, side-by-side visual comparison remains the most trustworthy test.


MCP setup

Claude Code

claude mcp add pdf-library -- /absolute/path/to/.venv/bin/pdf-library-mcp

Claude Desktop

Add the server to claude_desktop_config.json:

{
  "mcpServers": {
    "pdf-library": {
      "command": "/absolute/path/to/.venv/bin/pdf-library-mcp",
      "env": {
        "PDF_LIBRARY_ROOT": "/Users/you/Documents/pdf-library"
      }
    }
  }
}

Restart Claude Desktop after changing the configuration or adding new MCP tools.


MCP tools

Tool

Purpose

import_pdf

Import a PDF in the background or return an existing cached document

search_library

Search headings, pages and snippets

get_pages

Return Markdown for specific pages

get_chunk

Retrieve one semantic chunk

get_section

Retrieve the chunks belonging to a section

list_documents

List library metadata

document_status

Import status and document quality information

get_page_image

Return an original page or cropped block

ocr_review_queue

List OCR pages awaiting visual verification

review_ocr_page

Return the original image together with its transcription

record_ocr_review

Store a visual-review verdict

correct_ocr_page

Compare one OCR page with the original using the local vision model

correct_ocr_pages

Run vision OCR correction across selected pages

reprocess

Re-extract selected pages using Marker

import_pdf and reprocess return background job IDs when appropriate.

Use:

job_status(job_id)

to follow the operation.


Example vision workflow

Suppose OCR of four handwritten mathematics pages looks suspicious.

First ask the local VLM to evaluate them without modifying anything:

correct_ocr_pages(
    document="ΜΑΘΗΜΑΤΙΚΑ Ι ΜΑΘΗΜΑ 10.pdf",
    pages=[1, 2, 3, 4],
    apply=false
)

Inspect the proposed corrections.

Then:

correct_ocr_pages(
    document="ΜΑΘΗΜΑΤΙΚΑ Ι ΜΑΘΗΜΑ 10.pdf",
    pages=[1, 2, 3, 4],
    apply=true
)

Accepted pages are:

  1. written back as canonical Markdown

  2. recorded in the correction audit history

  3. rebuilt into document.md

  4. reindexed for search

The previous transcription remains recorded.


Storage

Default layout:

~/Documents/pdf-library/
│
├── library.db
│
└── documents/
    └── <document_id>/
        ├── source.pdf
        ├── document.md
        ├── metadata.json
        └── pages/
            ├── 0001.md
            ├── 0002.md
            └── ...

SQLite stores:

  • documents

  • pages

  • chunks

  • FTS search data

  • jobs

  • OCR reviews

  • vision correction provenance

Page files are both the caching unit and retrieval unit.

That is what makes single-page replacement possible.


Scanned and handwritten documents

A page without a usable text layer is not silently treated as an empty page.

It is marked as scanned and becomes eligible for OCR/high-quality processing.

Handwritten Greek remains one of the hardest cases.

Marker often recovers mathematical structure surprisingly well, including expressions such as:

\frac{A_{2,m_2}}{(x-r_2)^{m_2}}

while surrounding prose can still contain character substitutions.

The local vision layer exists specifically for the cases where text-only post-processing has reached its limit.


Configuration

See:

config.example.toml

The library root can be overridden with:

PDF_LIBRARY_ROOT=/path/to/library

and the configuration file with:

PDF_LIBRARY_CONFIG=/path/to/config.toml

Nothing is hard-coded.

Changes to text normalization may invalidate existing search indexes.

Run:

pdf-library doctor

to detect problems and:

pdf-library reindex --all

to rebuild search data without reprocessing the PDFs.


Tests

Run:

.venv/bin/python -m pytest tests -q

Fixtures include compiled LaTeX documents so mathematical extraction can be compared against known source material.

The test suite covers core guarantees such as:

  • cached imports do not invoke extraction again

  • digital PDFs do not unnecessarily trigger OCR

  • display equations remain intact across chunks

  • Greek search works across inflection

  • approximate OCR search does not outrank exact search

  • rendered images respect their token budget

  • MCP tools respect their response budgets

  • vision OCR provenance is preserved

  • low-confidence vision output does not silently replace canonical text


Design principles

Local first

Documents, indexes, OCR and optional vision inference stay on the user's machine.

Extraction is expensive. Retrieval should not be.

Run extraction once.

Reuse the result indefinitely.

Page-level quality beats whole-document perfection

Do not spend hours rerunning an 800-page book because three pages are bad.

The original PDF remains the source of truth

Every automated extraction layer can be wrong.

The original page image is always available for verification.

Corrections require provenance

The system keeps the previous transcription and records how a replacement was produced.

Deterministic fixes before generative fixes

Known problems such as Greek normalization and closed mathematical notation sets are handled deterministically where possible.

Vision models are reserved for cases where visual understanding actually adds information.


Deliberately not included

This project intentionally avoids several pieces of infrastructure that are not currently necessary.

Vector database

SQLite FTS5 with language-aware normalization already handles the intended library queries.

Semantic/vector search can be added later behind the same retrieval interface.

Cloud OCR as a requirement

The project is designed to remain usable locally.

Vision correction can run entirely on Apple Silicon through MLX-VLM.

Automatic LLM rewriting of every page

A vision model is not part of the mandatory import path.

It is an optional correction layer for pages that justify the additional cost.

A third traditional extraction engine

PyMuPDF4LLM and Marker already cover the fast/high-quality extraction split.

Another extractor should only be added if benchmarks show a concrete advantage.

Postgres, Redis, workers or a web application

This is a local tool for one person's document library.

SQLite and the filesystem are enough.


Project direction

See:

docs/NEXT.md

for planned work and design decisions.

Potential future areas include:

  • automatic routing of suspicious OCR pages to the local vision model

  • richer vision-model benchmarking

  • additional MLX-VLM model profiles

  • equation-level confidence reporting

  • optional semantic retrieval

  • richer document provenance inspection


Licence

The source code in this repository is licensed under the MIT License.

See LICENSE for details and for information about licenses of optional or required extraction dependencies.

The project orchestrates external extraction engines behind interfaces so that the document-library layer remains independent of any single extractor.

Available Tools

9 tools
document_statusA

Processing state and quality report for one document: whether it is complete, which pages are scanned or low quality, how many equations were found, and which pages would benefit from re-extraction. Poll this after import_pdf.

ParametersJSON Schema
NameRequiredDescriptionDefault
documentYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of disclosing behavior. It describes what the tool reports and implies it is a read-only status check, but does not explicitly state that it does not modify data, nor does it mention error conditions (e.g., document not found, still processing). The output schema exists, so return format is covered, but behavioral details are partially implicit.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded: it leads with the core purpose, lists the report contents, and ends with a clear usage instruction. Both sentences earn their place, with no redundant wording or filler. It is well-structured for quick parsing by an agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's simplicity (one parameter) and the existence of an output schema, the description is fairly complete. It tells the agent what the tool reports and when to use it. It does not cover edge cases like partial results while processing, but that is a minor omission given the polling context and the availability of the output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, meaning the description must compensate for parameter meaning. The description mentions 'one document' and the parameter is named 'document', making it clear that it expects an identifier for a document. However, it does not specify the format (e.g., ID vs. path) or how to obtain it, relying on shared context from sibling tools. It adds some meaning beyond the raw schema but leaves room for ambiguity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: it provides a processing state and quality report for a single document, enumerating specific details (completeness, page quality, equation count, re-extraction suggestions). This is specific and distinguishes it from siblings like search_library or get_pages, which serve different functions. The phrase 'Poll this after import_pdf' also ties it to a specific workflow.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit usage guidance: 'Poll this after import_pdf.' This tells the agent when to call it and implies it is for checking the result of an import operation. It does not explicitly state when not to use it or mention alternatives, but the context is clear given the sibling tools and the specific trigger.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_chunkA

Return one chunk in full, by the id shown in search results. This is the cheapest way to read a single theorem, proof or definition.

ParametersJSON Schema
NameRequiredDescriptionDefault
chunk_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the behavioral burden. It adds useful traits: the result is a 'full' chunk, the operation is the 'cheapest way' to read a single item, and the chunk represents a theorem, proof, or definition. However, it does not disclose error behavior, authentication requirements, or consequences of an invalid/stale id.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler. The action and resource are front-loaded in the first sentence, and the cost/use-case rationale is in the second. Every phrase earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter read operation with an output schema, the description is largely complete: it identifies the parameter source, the content type, and the cost profile. It could add explicit not-found behavior or a pointer to alternatives, but nothing essential is missing for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does by explaining that chunk_id is the id shown in search results, which is the key semantic an agent needs to use the tool. For a single integer parameter, this is nearly sufficient.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Return') and resource ('one chunk'), and adds the scope 'in full' and the id provenance 'shown in search results'. It also names the content type (theorem, proof, or definition), which clearly differentiates it from sibling tools like get_pages and get_section.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'This is the cheapest way to read a single theorem, proof or definition' gives a clear when-to-use signal, and 'by the id shown in search results' tells the agent where to obtain the argument. It does not explicitly name alternatives or state when not to use it, but the context is clear enough for correct selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_page_imageA

Show the original scanned page, or one region of it, as an image. This is the expensive tool and the last resort: an image costs several times what the same page's text costs, and it stays in context for the rest of the conversation. Use it only when the extracted text is evidently corrupted — nonsense words in a formula, an equation that does not parse — and always pass a block number if you can, since cropping to one equation is both cheaper and sharper than the whole page. get_pages lists the block numbers for OCR'd pages.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageYes
blockNo
documentYes
max_tokensNo

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It explicitly warns about cost ('costs several times what the same page's text costs') and context retention ('stays in context for the rest of the conversation'), and notes that cropping to a block is cheaper and sharper. It does not mention whether the tool is read-only or error behaviors, but the key operational traits are covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense yet efficient. Each sentence earns its place: purpose, cost/context warning, usage condition, and block guidance. It is front-loaded with the core purpose and immediately gives actionable constraints, with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema, the description does not specify the return format, but the agent can infer it will receive an image. It covers the critical cost/context trade-off and integrates with sibling get_pages for block numbers. The main gap is a lack of error or fallback details, but overall it provides sufficient information for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It explains the block parameter's purpose and how to obtain block numbers via get_pages, and implicitly clarifies page and document are required. It does not mention max_tokens, but that parameter is likely self-explanatory. The description adds meaning beyond the schema for the most critical parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb and resource: 'Show the original scanned page, or one region of it, as an image.' It distinguishes itself from text-focused siblings like get_pages, get_chunk, and get_section by explicitly framing it as an image tool and as the expensive last resort.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-to-use guidance: only when extracted text is corrupted, with concrete examples (nonsense words in a formula, equation that does not parse). It also provides when-not-to-use (last resort) and points to get_pages for block numbers, effectively routing the agent to the right alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_pagesA

Return the Markdown of specific pages of one document, with LaTeX preserved. Ask for the few pages a search pointed at, plus a neighbouring page if context is missing. Requesting a wide range is how you flood your own context; the reply is truncated at the configured token budget.

ParametersJSON Schema
NameRequiredDescriptionDefault
pagesYes
documentYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It discloses that the reply is truncated at the configured token budget and warns about flooding context, which is critical for an agent to avoid misuse. It does not detail error handling or edge cases, but the disclosed behaviors are significant and well-covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, front-loaded with the core purpose, and each sentence adds value—purpose, usage guidance, and a warning. There is no redundancy or wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema exists, the description does not need to explain return values. It covers the essential usage, the token-budget warning, and the context strategy. The main gap is the lack of explicit parameter constraints like page indexing, but overall it is sufficient for an agent to make correct calls.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage, so the description must compensate. It implies 'document' identifies the document and 'pages' are the specific page numbers, but it does not explicitly state indexing (e.g., 1-based) or the expected format. The guidance to 'ask for the few pages a search pointed at' gives practical context but lacks precise parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool returns the Markdown of specific pages of one document, preserving LaTeX. It specifies the exact resource (pages of a document) and the action (return Markdown), making it easily distinguishable from siblings like get_chunk or get_section.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives practical guidance on when to use it: after a search, request the few relevant pages plus a neighboring page if context is missing. It also warns against requesting wide ranges due to token truncation. However, it does not explicitly name alternative tools for cases like retrieving a full section or a chunk, so the routing guidance is incomplete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_sectionA

Return every chunk filed under one heading of a document, joined in order. Use when you know the section name; otherwise search first.

ParametersJSON Schema
NameRequiredDescriptionDefault
sectionYes
documentYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, but the description implies a read-only operation. It does not mention side effects or error conditions, but the tool name and description strongly suggest it is non-mutating.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured, combining the core functionality and a usage hint in a single sentence. It avoids unnecessary detail while remaining informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description and sibling context provide a good overall understanding, but the lack of parameter descriptions leaves some gaps regarding input formats. The output schema exists but is not shown, so the return type is not fully specified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides no descriptions for the 'document' and 'section' parameters. While parameter names are self-explanatory, the format of 'section' (exact heading, partial match, case sensitivity) is ambiguous, leaving the agent to infer expected input.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: returning every chunk under a specific heading of a document, joined in order. This distinguishes it from sibling tools like get_chunk (single chunk) and get_pages (pages).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly advises when to use this tool: 'Use when you know the section name; otherwise search first.' This provides clear guidance and points to the alternative search_library tool when the section name is unknown.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

import_pdfA

Add a PDF to the library, or confirm it is already there. Extraction runs in the background and returns a job id immediately; poll document_status until status is 'complete'. A file whose content hash is already known is an instant cache hit and is never reprocessed, so there is no cost to calling this, but also no reason to call it twice.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYes
forceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does an excellent job: it discloses background extraction, immediate job-id return, cache-hit idempotency, no reprocessing, and zero cost for duplicates. This is comprehensive behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core purpose, and every clause adds value (async behavior, polling, cache semantics, cost implications). No redundancy or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the async workflow, polling requirement, idempotency, and cost trade-offs—all essential for correct invocation. Parameter meanings are missing, but that is a separate dimension; for the tool's behavior, the description is thorough and leaves no critical ambiguity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, meaning the description does not explain either 'path' or 'force'. The purpose of 'force' is left entirely to the schema, and even 'path' is not contextualized beyond being a file path. The description fails to compensate for the low coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states 'Add a PDF to the library' with a clear verb and resource, and immediately clarifies the cache-hit behavior that distinguishes it from the sibling get/search/reprocess tools. It fully defines the tool's role in the library workflow.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit operational guidance: poll document_status until complete, and warns against calling twice due to cache hits. It does not name alternative tools, but the context makes the intended usage clear and self-contained.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_documentsA

List every document in the library with its id, page count and processing state. Metadata only, no content.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of disclosing behavior. It explicitly states that only metadata is returned and no content, which is a key behavioral trait. It also implies a read-only operation, though it does not mention pagination or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that conveys all necessary information without any redundancy or fluff. It is appropriately concise for a straightforward list operation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is complete for a tool with no parameters and a simple list operation. It specifies the exact output fields (id, page count, processing state), so the agent knows what to expect. No additional context is needed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the schema description coverage is 100%. According to the rubric, the baseline score of 3 applies, and there is no additional parameter information to add since none exist.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'List', the resource 'documents in the library', and specifies the returned fields (id, page count, processing state). It also distinguishes itself from the sibling tool 'search_library' by emphasizing 'every document' rather than a search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides a clear guideline by stating 'Metadata only, no content,' which informs the agent about the nature of the response. While it does not explicitly contrast with search_library, the purpose itself sufficiently implies when to use this tool for listing all documents.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reprocessA

Re-extract selected pages of a document with the high-quality engine (Marker), which produces real LaTeX for equations. Slow and optional: use it for the specific pages whose math came out badly, never for a whole book. Runs as a background job.

ParametersJSON Schema
NameRequiredDescriptionDefault
pagesNo
engineNo
documentYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden. It does disclose important traits: it is slow, optional, runs as a background job, and produces real LaTeX. However, it does not state whether results overwrite previous extraction, whether the call is idempotent, or how the agent should learn when the background job completes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with no filler. The core action and engine are front-loaded, followed by usage guidance and background behavior. Every sentence adds distinct value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers purpose, usage, and basic behavior, and an output schema exists so return values are covered elsewhere. However, for a background job, it should ideally point the agent to a status-checking sibling like document_status, and clarify whether the re-extraction replaces the prior extraction or creates new data. These gaps are material for an agent deciding how to invoke and monitor the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It maps all three parameters: 'pages' through 'selected pages' and 'specific pages,' 'engine' through 'high-quality engine (Marker),' and 'document' through 'of a document.' It also adds a key usage constraint about not reprocessing a whole book. It does not explain default behavior when pages is null, but it largely covers the parameter meanings.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Re-extract selected pages'), a specific resource ('a document'), and a distinctive method (Marker high-quality engine producing real LaTeX). This clearly separates it from sibling tools like get_pages, which would only fetch or view pages rather than re-run extraction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit guidance: use for 'specific pages whose math came out badly' and 'never for a whole book.' It also labels the tool as 'slow and optional,' which helps an agent decide when to invoke this tool versus cheaper or simpler alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_libraryA

Full-text search across every processed document. This is the entry point for any question about library content: it returns headings, page numbers and short snippets, never full text. Follow up with get_pages or get_chunk for the passages that look right. Accent-insensitive and works for Greek and English.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryYes
documentNo
chunk_typeNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals that results are limited to headings, page numbers, and snippets (never full text), and that search is accent-insensitive and supports Greek and English. This goes beyond a simple 'searches documents' statement. It doesn't mention pagination or behavior on no results, but for a search tool this is solid coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, with the purpose front-loaded in the first sentence. It is concise, every sentence adds value (purpose, output type/follow-up, language behavior), and there is no fluff or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 4 parameters and an output schema, the description is adequate for the core purpose but incomplete for parameters. It explains the return type (headings, page numbers, snippets) but omits any guidance on how to use 'document' or 'chunk_type' filters. The presence of an output schema reduces the need to describe return format, but the parameter semantics gap makes it not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, but it does not. It implies 'query' is the search term via 'Full-text search,' but never explains 'limit,' 'document,' or 'chunk_type.' The 'document' parameter likely filters to a specific document, which is a significant omission given the description says 'every processed document.' The lack of parameter explanation is a clear gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Full-text search across every processed document.' It distinguishes itself from siblings by being 'the entry point for any question about library content' and explicitly notes it returns 'headings, page numbers and short snippets, never full text,' which separates it from get_pages and get_chunk. The verb and resource are specific and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit usage guidance: 'This is the entry point for any question about library content' establishes when to use it, and 'Follow up with get_pages or get_chunk for the passages that look right' names the alternatives and the condition for switching. It also warns 'never full text,' implying those follow-ups are needed for full content. This is clear when-to-use and when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 9 tool updatesv0.1.0
    • First observeddocument_status
    • First observedget_chunk
    • First observedget_page_image
    • First observedget_pages
    • First observedget_section
    • First observedimport_pdf
    • First observedlist_documents
    • First observedreprocess
    • First observedsearch_library

TDQS

A4.3/5.0

Scored across 9 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: import, search, retrieve pages/chunks/sections, image retrieval, listing, status, and reprocessing. No overlap in functionality, and descriptions make the boundaries explicit.

Naming Consistency4/5

Most tools follow a consistent verb_noun pattern (import_pdf, search_library, get_pages, get_chunk, get_section, get_page_image, list_documents). Two tools deviate: document_status (noun) and reprocess (bare verb), but these are minor and the overall pattern is recognizable.

Tool Count5/5

9 tools is well within the ideal 3-15 range. Each tool serves a specific step in the PDF library workflow—import, search, retrieval, status, and reprocessing—without redundancy or bloat.

Completeness4/5

The set covers the full reading workflow: import, search, retrieve, status, and reprocess. The only notable gap is the lack of a delete/remove tool, but given the read-focused nature of the server, this is a minor omission that agents can work around.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • F
    license
    Not graded
    quality
    D
    maintenance
    A local academic research assistant that indexes PDFs into a searchable vector library and exposes MCP tools for semantic search, claim extraction, contradiction detection, and multi-step research synthesis.
    -
  • A
    license
    A
    quality
    A
    maintenance
    A local MCP server for searching scientific papers, retrieving metadata and abstracts, and legally downloading Open Access PDFs via OpenAlex, CrossRef, and Unpaywall APIs.
    5
    3
    MIT
  • F
    license
    A
    quality
    B
    maintenance
    Enables semantic search across personal PDF paper collections with page-level citations, allowing users to query their library from any MCP-capable client.
    9
    -
  • A
    license
    Not graded
    quality
    D
    maintenance
    Remotely-callable MCP server for academic paper search, full-text retrieval and image to LaTeX conversion across arXiv, Semantic Scholar, and OpenAlex.
    1
    MIT