benspdf
An MCP server that gives AI agents local, privacy-preserving PDF tools — inspect, render, create, and manage PDFs without uploading them.
PDF inspection: count pages, read metadata (title/author/dates/producer), check whether pages are text or scanned, extract text page by page, and OCR scans (optionally adding a searchable text layer).
Layout and access: report page sizes, orientation, rotation, and page boxes; check encryption and permissions such as printing, copying, and editing.
Rendering: render PDF pages to PNG images for visual inspection.
Test files: create blank test PDFs with configurable page count and title.
Artifact workflow: temporary results are stored as artifacts; list them, discard them, or export them to real disk paths — the only tool that writes to the user's filesystem.
Local/offline friendly: files stay on the machine, and with Ollama the whole stack can run offline.
Pairs the PDF tools with a local Ollama model so questions about PDFs are answered entirely offline. The bundled benspdf-cli connects to Ollama (e.g. llama3.1, selectable via --model or BENSPDF_MODEL) and drives the PDF tools — page counts, metadata, layout, rendering — without any data leaving the machine.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@benspdfhow many pages are in ~/Downloads/report.pdf?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Ben's PDF tools for AI agents, exposed over the Model Context Protocol (MCP).
Your PDFs are read on your own machine and never uploaded. With a hosted model your questions still reach that model; pair the tools with a local Ollama model and nothing leaves the machine at all.
Works with Claude Desktop, Claude Code, VS Code, Kiro, Cursor, the ChatGPT desktop app, and any other MCP client. The one-click buttons above need uv; for every other client, or to do it by hand, see Setup.
Tools
Tool | What it does |
| Counts the pages in a PDF |
| Reads document properties: title, author, dates, producer |
| Says whether a PDF is readable text or a scan that needs OCR |
| Reads the text a PDF holds, page by page |
| Page sizes, orientation, rotation and page boxes |
| Encryption, and what the file permits: printing, copying, editing |
| Renders pages to images, so a page can be looked at |
| Reads a scan with OCR, and can add a searchable text layer to it |
| Generates a throwaway PDF, handy for trying things out |
| Saves results to a real location on disk |
| Lists recent temporary results |
| Deletes temporary results now |
Every tool but one needs nothing beyond the package. pdf_ocr uses
tesseract, a system program rather
than a Python package, and only looks for it when you actually call it — so
install it if and when you want OCR (brew install tesseract, sudo apt install tesseract-ocr, or winget install UB-Mannheim.TesseractOCR), and everything else
works either way.
Related MCP server: pdfeverything
Where results go
When a tool makes a new PDF, it goes into a scratch folder instead of your own
folders, and you get back a short id like art_a1b2c3d4.pdf. Tools accept those
ids anywhere they accept a file path, so several steps can be chained together.
Results carry the artifact's path as well as its id, so you can open a rendered
page or an intermediate file straight away without exporting it first.
export is the only tool that writes into your folders, so nothing shows up
until you ask for it. Each time the server starts it clears out scratch files
older than 7 days. Set BENSTOOLS_WORKSPACE to put the scratch folder somewhere
other than ~/.benstools/work.
Setup
Install uv
once, then point your client at uvx benspdf-mcp and uv fetches the package,
plus a suitable Python, on first run.
# macOS and Linux
curl -LsSf https://astral.sh/uv/install.sh | sh
# Windows
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"On macOS, brew install uv works too.
Add the server to your client with one of the configs below, then restart the server from your client's UI. Once connected, just ask in plain language:
How many pages are in ~/Downloads/report.pdf?
Claude Desktop
Edit claude_desktop_config.json, which lives at
~/Library/Application Support/Claude/ on macOS and %APPDATA%\Claude\ on
Windows. You can also open it from Settings → Developer → Edit Config.
{
"mcpServers": {
"benspdf": {
"command": "uvx",
"args": ["benspdf-mcp"]
}
}
}Claude Code
One command, no config file. Add --scope user to enable it everywhere rather
than just the current project.
claude mcp add benspdf -- uvx benspdf-mcpCheck it connected with claude mcp list.
VS Code
.vscode/mcp.json in your workspace, or the same file in your user profile.
Note VS Code uses servers rather than mcpServers, and wants an explicit
type.
{
"servers": {
"benspdf": {
"type": "stdio",
"command": "uvx",
"args": ["benspdf-mcp"]
}
}
}Kiro
.kiro/settings/mcp.json in your workspace, or ~/.kiro/settings/mcp.json to
enable it everywhere. autoApprove skips the confirmation prompt for tools you
trust.
{
"mcpServers": {
"benspdf": {
"command": "uvx",
"args": ["benspdf-mcp"],
"autoApprove": ["pdf_page_count"]
}
}
}Cursor
The button at the top does this for you. By hand, it's ~/.cursor/mcp.json to
enable it everywhere, or .cursor/mcp.json in a project.
{
"mcpServers": {
"benspdf": {
"type": "stdio",
"command": "uvx",
"args": ["benspdf-mcp"]
}
}
}Windsurf
~/.codeium/windsurf/mcp_config.json. You can also reach it from Cascade:
Settings → Cascade → Manage MCPs → View raw config, which is worth using
since the path has moved between versions.
{
"mcpServers": {
"benspdf": {
"command": "uvx",
"args": ["benspdf-mcp"]
}
}
}Continue
~/.continue/config.yaml, or a .yaml file under .continue/mcpServers/ in a
project. Continue is the one client here that doesn't take the JSON shape above:
its config is YAML, and mcpServers is a list rather than an object keyed by name.
mcpServers:
- name: benspdf
command: uvx
args:
- benspdf-mcpChatGPT desktop app, Codex CLI, Codex IDE extension
All three are Codex clients and share one config file, ~/.codex/config.toml,
so adding the server once covers all of them. Note this one is TOML, not JSON.
[mcp_servers.benspdf]
command = "uvx"
args = ["benspdf-mcp"]Any other MCP client
Almost every client uses the same JSON as Claude Desktop above — an mcpServers
object, with command set to uvx and args to ["benspdf-mcp"]. Some want an
explicit "type": "stdio"; adding it is harmless where it isn't required.
If a client just asks for a command to run, it's:
uvx benspdf-mcpFully offline with Ollama
The clients above keep your PDFs local, but they answer using a hosted model. Pair the tools with a local model instead and nothing leaves your machine.
You'll need Ollama with a model pulled, plus this package:
pip install benspdf-mcp ollama
ollama pull llama3.1Then run the bundled CLI:
benspdf-cli # uses the first model you have
benspdf-cli --model llama3.1 # or pick one
BENSPDF_MODEL=llama3.1 benspdf-cli # or set it oncepython -m benspdf.cli does the same thing, handy from a source checkout.
You: how many pages in ~/Downloads/report.pdf?
[Using tool: pdf_page_count]
[Result: 12 pages in report.pdf]
Assistant: The PDF has 12 pages.Reference
Each tool's own description tells your client what it does and when to use it, so in normal use there is nothing to look up. If you want the detail — every field a tool returns, and the reasoning behind the answers it gives — see the tool reference.
To work on the code, see CONTRIBUTING.md.
License
Apache 2.0
Available Tools
12 toolscreate_test_pdf_fileA
Create a test PDF, handy for trying the other tools without hunting for one.
The pages are blank, so this is for exercising tools rather than for anything that needs real content.
Args: output_path: Optional path to also write the PDF to. Omit to keep it as a temporary artifact. num_pages: Number of blank pages to create (default: 3) title: Optional title for the PDF metadata
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | ||
| num_pages | No | ||
| output_path | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses that the PDF pages are blank, and that omitting output_path keeps it as a temporary artifact while providing a path writes the PDF there. It doesn't mention cleanup or permissions, but for a simple test-file generator these are minor given the self-contained nature.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two intro sentences establish purpose and limitation, then a compact Args list documents parameters. No redundant or filler sentences; the most important scoping detail ('blank pages') is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity tool with three optional parameters and an output schema, the description covers purpose, use context, behavior, and parameter semantics. There is nothing an agent needs in order to select and invoke it correctly that is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Although schema description coverage is 0%, the description's Args section fully explains every parameter: output_path optional path or temporary artifact, num_pages as the count of blank pages, and title as metadata. This adds meaning well beyond the bare schema properties and defaults.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Create a test PDF.' It also explains the intended niche ('handy for trying the other tools') and clarifies the pages are blank, which distinguishes this generator from the sibling PDF-analysis/processing tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says this is for 'exercising tools' and 'rather than for anything that needs real content,' giving clear when-to-use guidance. Although it doesn't name an alternative tool, no sibling creates PDFs, so the intended context is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
discardA
Delete temporary artifacts from the workspace now.
Only accepts artifact ids, never file paths, so this cannot delete the user's own files. Artifacts expire on their own, so this is only needed when the user explicitly asks to clean up.
Args: refs: Artifact id, or list of artifact ids, to delete.
| Name | Required | Description | Default |
|---|---|---|---|
| refs | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full responsibility for behavioral disclosure. It clearly communicates that this is a destructive deletion action, limits its scope to temporary artifacts, and emphasizes that it cannot affect the user's own files, which is essential safety information.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence earns its place: the first states the action, the second explains a safety constraint, the third provides usage context, and the final line documents the parameter. No fluff or unnecessary repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple single-parameter destructive tool, the description provides everything an agent needs to invoke it correctly and safely: what it deletes, what it cannot delete, when it is appropriate, and how to specify the targets. The output schema exists, so return-value details are not needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must explain the parameter, and it does: 'refs: Artifact id, or list of artifact ids, to delete.' This adds the crucial semantic that refs are artifact ids rather than file paths, which is not present in the input schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a precise verb and resource: 'Delete temporary artifacts from the workspace.' It clearly distinguishes the tool from file-related siblings by stating it only accepts artifact ids, never file paths, and therefore cannot delete the user's own files.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when the tool is needed ('only needed when the user explicitly asks to clean up') and when it is not ('Artifacts expire on their own'). It also gives a hard constraint: inputs must be artifact ids, not file paths.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
exportA
Save one or more results to a real location on disk.
The only tool that writes to the user's filesystem, so call it once at the end, when they have said where the output should go. Never hand an artifact id back as if it were a finished file; artifacts expire.
Args:
refs: Artifact id, or list of artifact ids, to export.
dest: Where to write. For a single artifact this may be a full file
path. For several artifacts, or when name is given, this is a
directory and is created if missing.
name: Filename template used when writing into a directory.
Placeholders: {n} 1-based index (e.g. "page_{n:03d}{ext}"),
{ext} extension, {id} artifact id, {stem} id without extension.
Defaults to "{n:03d}{ext}".
overwrite: Replace existing files instead of failing. Defaults to
False so nothing is destroyed by accident.
| Name | Required | Description | Default |
|---|---|---|---|
| dest | Yes | ||
| name | No | ||
| refs | Yes | ||
| overwrite | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it does so well. It discloses filesystem writes, directory creation when missing, overwrite behavior, and the default safety intent ('nothing is destroyed by accident'). It also warns about artifact expiration, which is important behavioral context for the agent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and usage rule, then moves to a structured Args explanation. It is longer than minimal, but every sentence earns its place: usage timing, filesystem exclusivity, artifact expiry warning, and detailed parameter instructions. The structure is easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the 4-parameter complexity and an output schema that likely describes the return value, the description is complete. It covers when to use the tool, what it does, how parameter interactions work, directory creation, overwrite safety, and naming templating. An agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description fully compensates. It explains each parameter in the Args section: refs as artifact id or list, dest as file path or directory, name as a filename template with placeholders like '{n:03d}{ext}', and overwrite semantics with defaults. This goes well beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Save one or more results to a real location on disk.' It also differentiates the tool from its siblings by explicitly stating it is 'the only tool that writes to the user's filesystem.' This makes the tool's purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit usage timing: 'call it once at the end, when they have said where the output should go.' It also provides a clear when-not-to-use rule by warning never to hand an artifact id back as a finished file and explaining that artifacts expire. This effectively distinguishes when export is needed versus when it is not.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_artifactsA
List recent temporary artifacts in the workspace, newest first.
Useful when you have lost track of an artifact id from earlier in the conversation, so you can find it again instead of redoing the work.
Args: limit: Maximum number of artifacts to return.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses that results are ordered newest first, and the term 'temporary' implies the artifacts are ephemeral. It doesn't mention side effects, but listing is inherently read-only, and the description aligns with that expectation. It adds useful behavioral context beyond a generic 'list' operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences plus a one-line Args section. The purpose is front-loaded, the usage scenario follows immediately, and the parameter explanation is concise. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return format is covered. The description explains what the tool does and when to use it, and the parameter is documented. It lacks only minor details like whether the list is exhaustive or paginated, but for a simple list with a limit, this is adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must explain parameters. It explicitly states 'limit: Maximum number of artifacts to return,' which adds meaning beyond the schema's simple integer type and default value. This fully compensates for the lack of schema description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'list', the resource 'temporary artifacts', and the ordering 'newest first'. It distinguishes itself from the sibling PDF tools and other operations like discard/export by specifying exactly what it does and the context of workspace artifacts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a clear scenario for when to use this tool: when you've lost track of an artifact ID and want to find it again. While it doesn't explicitly name alternatives or say when not to use it, the context is specific enough to guide an agent effectively.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pdf_check_accessA
Check a PDF's encryption and what it permits: printing, copying, editing.
Answers "is this locked, and what am I allowed to do with it?".
encrypted and needs_password are separate answers and the difference
matters: most encrypted files have no user password, so they open silently and
only carry restrictions. permissions gives a boolean per action,
restrictions lists what is denied, and an unencrypted file permits
everything — a PDF has nowhere to keep restrictions but its encryption
dictionary.
Treat restrictions as what the file asks of viewers, not as enforcement: once a document is open nothing stops the bits being ignored, and a file that denies copying still yields its text to pdf_extract_text. Report them as the author's intent, never as an action being impossible.
Worth reaching for when another tool reports a file as encrypted; this one still answers, since the encryption dictionary is readable when the pages are not.
Args: ref: PDF file path, or a workspace artifact id.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and does so thoroughly. It explains the distinct meanings of encrypted and needs_password, defines permissions vs. restrictions, covers the unencrypted case, and clarifies that restrictions are statements of intent rather than enforced barriers. It also notes that it can answer even when pages cannot be read.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than average but every section earns its place: purpose, key distinctions, interpretation guidance, and when-to-use context. The main purpose is front-loaded in the first two sentences, and the argument documentation is cleanly separated at the end.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the conceptual complexity of PDF encryption, the description is complete. It explains edge cases, defines the output semantics without relying on the output schema, gives usage context relative to other tools, and documents the single parameter. Nothing an agent needs to decide whether and how to call this tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema only describes ref as a string with no explanation, and schema coverage is 0%. The description compensates fully by stating 'ref: PDF file path, or a workspace artifact id,' which gives the agent the two acceptable forms of the argument. This is exactly the guidance the schema lacks.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action and resource: 'Check a PDF's encryption and what it permits: printing, copying, editing.' It also gives the core question it answers, 'is this locked, and what am I allowed to do with it?', which clearly separates it from metadata or text extraction siblings. The nuance about encrypted vs. needs_password further sharpens its identity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says when to reach for this tool: when another tool reports a file as encrypted, this one still works because the encryption dictionary is readable. It also connects to pdf_extract_text to warn that denied permissions are not enforcement, teaching the agent how to use the result correctly in a multi-tool workflow.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pdf_check_textA
Check whether a PDF has a text layer, looks scanned, or needs OCR.
Answers "can I read this, or is it a picture of a document?". Worth running before extracting text from a file you have not seen.
verdict is "text", "scanned", "mixed", or "no_text" — the last meaning
nothing readable but nothing scan-like either, so blank, vector-only, or
illustrated pages that OCR cannot help. needs_ocr follows from it, summary
is a sentence worth quoting, and the per page evidence behind the verdict
comes back alongside, including each page's image_coverage (the largest
image's share of the page area).
Samples up to 10 pages spread across the document, so a true sampled means
the answer is an estimate for the pages in between. It does not return the
text (use pdf_extract_text) and it does not run OCR.
Args: ref: PDF file path, or a workspace artifact id.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it does so exceptionally. It discloses the sampling behavior (up to 10 pages), what a 'sampled' verdict means, the meaning of each verdict value, and that it returns per-page evidence such as image_coverage.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the purpose consuming no fluff. Every sentence adds meaningful detail — verdicts, sampling, output evidence, and exclusions — while staying compact and scannable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the decision this tool answers, the important caveat about sampling, the shape of the result, and the clear boundary against text extraction and OCR. An agent has everything it needs to select and invoke this tool correctly, and the output schema covers the remaining return details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, but the description fully compensates by defining `ref` as 'PDF file path, or a workspace artifact id.' For a single-parameter tool, this is sufficient and adds exactly the semantic meaning the schema lacks.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Check whether a PDF has a text layer, looks scanned, or needs OCR.' It clearly distinguishes itself from siblings like pdf_extract_text and pdf_ocr by stating what it does not do.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says when to run the tool: 'Worth running before extracting text from a file you have not seen.' It also names the alternative for text extraction (pdf_extract_text) and states that it does not run OCR, giving clear routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pdf_extract_textA
Read the text of a PDF that has one, page by page.
The cheap and exact way to read a document: milliseconds a page, and the characters are the ones the file holds. Prefer it to pdf_ocr, which is for pages carrying no text, and to pdf_render_pages, which costs an image a page. pdf_check_text says which path a file needs; page numbers come back with the text, so quote page 8 rather than "the document".
A long document arrives in ranges: one call returns about 50,000 characters
and names the pages it did not reach, so when truncated, do what the summary
says. output="txt" writes the whole extraction to an artifact and returns
the counts alone, which is the way to read a book.
Text on a text_suspect page came back partly undecodable, because the fonts
carry no character map: say so rather than quoting it, and look at the page
with pdf_render_pages.
Args: ref: PDF file path, or a workspace artifact id. pages: Which pages: "1-10", "3", "1,5,9-12" or "all". Defaults to all. output: "text", "txt" for the text as an artifact, or "both". layout: Keep the page's spacing, for forms and tables. Costs characters.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | ||
| pages | No | ||
| layout | No | ||
| output | No | text |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does so thoroughly. It discloses performance characteristics (milliseconds per page), exactness of characters, truncation behavior around ~50,000 characters, what `truncated` means, artifact-writing behavior for output='txt', and the text_suspect caveat about undecodable fonts. This is exemplary behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is detailed but every sentence earns its place: each clause either defines a behavior, gives usage guidance, or explains a parameter. The first sentence is a clear one-line purpose, and the nuanced details are logically organized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity, the description is complete: it covers how to invoke it, what to expect, how to handle truncation, when to choose alternatives, and how to handle suspicious output. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. The Args section adds meaningful semantics for all four parameters: ref accepts a file path or artifact id, pages gives concrete syntax examples, output enumerates valid values and their effects, and layout explains when it matters. This fully compensates for the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: "Read the text of a PDF... page by page." It clearly differentiates the tool from pdf_ocr, pdf_render_pages, and pdf_check_text, so an agent knows exactly what this tool does and what it does not do.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage guidance is explicit and actionable: prefer this over pdf_ocr for pages with text, over pdf_render_pages when exact cheap text is needed, consult pdf_check_text to choose a path, use output='txt' for long books, and use pdf_render_pages for text_suspect pages. This leaves no ambiguity about when to select this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pdf_metadataA
Read a PDF's document properties: title, author, dates, producer, keywords.
Answers "who made this, when, and with what". It does not count pages (use pdf_page_count) and says nothing about whether the pages hold readable text (use pdf_check_text).
A PDF can store these fields in two independent places, the legacy Info
dictionary and an XMP packet, and the two often disagree. The normalized
answer is at the top level, preferring XMP, with sources naming the store
each value came from and conflicts listing every field where the two differ,
both values included. Both stores also come back verbatim as info and xmp.
When has_conflicts is true, say so rather than quoting one value as fact.
Args: ref: PDF file path, or a workspace artifact id.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses the dual-storage nature (Info vs XMP), the normalized output with sources and conflicts, and the presence of has_conflicts. It even instructs the agent on how to behave when conflicts exist, which is beyond typical transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured: a clear purpose sentence, a negative definition that routes to alternatives, a detailed but necessary explanation of the output behavior, and a final args block. It is efficient and front-loads the most important information without fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (metadata conflicts) and the existence of an output schema, the description still adds critical context about how to interpret the returned data (sources, conflicts, has_conflicts). It covers what the tool returns, how conflicts are resolved, and the dual stores. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter ref is explained as 'PDF file path, or a workspace artifact id.' This adds semantic meaning beyond the schema's bare 'string' type. Since schema coverage is 0%, this explanation fully compensates for the absence of parameter descriptions in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Read a PDF's document properties: title, author, dates, producer, keywords.' It clearly distinguishes itself from siblings by naming pdf_page_count and pdf_check_text as alternatives for different tasks, so an agent can immediately tell what this tool does and what it does not.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when NOT to use this tool: 'It does not count pages (use pdf_page_count) and says nothing about whether the pages hold readable text (use pdf_check_text).' It also gives operational guidance on how to handle conflicting metadata, advising to report has_conflicts rather than quoting a single value as fact. This covers both exclusions and usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pdf_ocrA
Read a scanned PDF with OCR, and optionally add a real text layer to it.
The way to read a page with no text: a scan, or an export that lost its text. Prefer it to pdf_render_pages for reading; run pdf_check_text when unsure. Tesseract is a system install, and a missing binary or language errors with how to install it.
OCR can be confidently wrong, so pages report mean_confidence and
low_confidence_words: say when confidence is low rather than presenting text as
certain, and look at a bad page with pdf_render_pages.
Text pages are skipped unless force. A long document may stop early; when
truncated, do exactly what the summary says, artifact and range included.
Args: ref: PDF file path, or a workspace artifact id. pages: Which pages: "1-10", "3", "1,5,9-12" or "all". Defaults to all. lang: Tesseract language code, or several as "eng+deu". dpi: Resolution, 72 to 600. 200 suits printed text; higher is slower, not better. output: "text", "pdf" for a searchable copy as an artifact, or "both". force: OCR pages that already have text instead of skipping them.
| Name | Required | Description | Default |
|---|---|---|---|
| dpi | No | ||
| ref | Yes | ||
| lang | No | eng | |
| force | No | ||
| pages | No | ||
| output | No | text |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and does it well: it discloses Tesseract system dependency and error handling, warns that OCR can be confidently wrong, explains confidence reporting, notes that text pages are skipped unless force, and describes early-stop/truncation behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is detailed but every sentence serves a purpose. It front-loads the core action, then provides targeted caveats, usage routing, and parameter semantics. No wasted words, and the structure makes the content easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex OCR tool with no annotations, the description is complete: it addresses external dependencies, failure modes, confidence interpretation, edge cases (text pages, truncation), and all parameters. The output schema presumably covers return values, so the description does not need to duplicate that.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must fully compensate. The Args section explains every parameter: ref format, pages syntax and default, lang format, dpi range and suitability, output options, and force semantics. This goes well beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a clear verb and resource: 'Read a scanned PDF with OCR, and optionally add a real text layer to it.' It also differentiates from siblings by saying to prefer it to pdf_render_pages for reading and to run pdf_check_text when unsure, so an agent can distinguish it without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use this tool ('a scan, or an export that lost its text') and gives direct alternatives: 'Prefer it to pdf_render_pages for reading; run pdf_check_text when unsure.' This gives the agent clear routing guidance relative to sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pdf_page_countA
Count the pages in a PDF. Answers "how many pages is this?".
Reads only the document's page tree, so it stays cheap on large files. It does not read page text, page sizes, or document properties.
Args: ref: PDF file path, or a workspace artifact id.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses that it reads only the page tree, stays cheap on large files, and lists what it does not read. It does not mention error conditions or explicitly state read-only behavior, but the phrasing implies no side effects. This is adequate for a simple count operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, with a clear first sentence stating the purpose, followed by a brief explanation of scope and the parameter. No redundant or filler content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with one parameter and an output schema, the description covers the core behavior and scope. It could mention potential errors (e.g., invalid file path) or the exact return format, but the output schema likely handles return details. Overall, it is sufficiently complete for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides only the parameter name 'ref' with no description, and schema coverage is 0%. The description compensates fully by explaining that 'ref' is a PDF file path or a workspace artifact ID, adding essential meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool counts pages in a PDF and explicitly distinguishes it from siblings by noting it reads only the page tree, not text, sizes, or properties. This makes the purpose unambiguous and separates it from pdf_metadata, pdf_extract_text, and pdf_page_layout.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context on what the tool does and does not do ('does not read page text, page sizes, or document properties'), which implies when not to use it. However, it does not name specific alternative tools like pdf_metadata or pdf_extract_text, so the guidance is implicit rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pdf_page_layoutA
Measure a PDF's page sizes, orientation and rotation.
Answers "what size is this, is it all the same, and why does one page come out sideways?". Reads page geometry only, so it stays cheap on long documents.
sizes groups the pages by shape, most pages first, each carrying the page
numbers it covers as a range like "1-16,18", so a 300 page document answers in
one entry instead of 300 rows.
Sizes are the page as a reader sees it: CropBox clipped to MediaBox with
/Rotate applied, so a landscape box rotated 90 degrees reports as portrait.
Pages within 3pt of a known paper are grouped and named together, since real A4
varies by a millimetre.
Pass pages for per page rows carrying the raw boxes: "1-20", "3", "1,5,9-12"
or "all", capped at 100 rows.
It does not look at page content; for whether the pages hold readable text use pdf_check_text.
Args: ref: PDF file path, or a workspace artifact id. pages: Optional page range for per page detail. Omit for groups only.
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | ||
| pages | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden, and it does so thoroughly. It explains that only page geometry is read, that pages are grouped by shape with page-range notation, that sizes are computed by clipping CropBox to MediaBox and applying /Rotate, and that a 3pt tolerance is used for grouping. It also discloses that per-page output is capped at 100 rows.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and then delivers dense, useful details about behavior, grouping, parameter formats, and alternatives. Each sentence contributes necessary information, and the final Args block cleanly summarizes the two parameters without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers what the tool returns conceptually, how it handles real-world page size variance, how rotation affects reported sizes, how to request per-page details, and what to use instead for text content. Since an output schema is present, the absence of exhaustive return-field documentation is not a gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does. It explains that 'ref' accepts either a PDF file path or a workspace artifact id, and it defines 'pages' with concrete range formats like '1-20', '3', '1,5,9-12', or 'all', plus the 100-row cap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Measure a PDF's page sizes, orientation and rotation.' It also frames the tool around concrete questions, which makes its purpose unmistakable. It further distinguishes itself from text-related siblings by explicitly stating it does not look at page content and points to pdf_check_text for that.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly states when to use this tool: for page geometry and understanding inconsistent sizes or rotation, while noting it stays cheap on long documents. It also gives a direct alternative, pdf_check_text, for readable-text questions, which is an explicit when-not-to-use condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pdf_render_pagesA
Render PDF pages to images, so a page can be looked at rather than read.
Use it to see what a page looks like: checking a change landed (did a redaction remove content, did a split cut where intended), previewing, and pages where appearance is the content — handwriting, signatures, charts, checkbox state.
For reading a scan prefer pdf_ocr, whose text layer is searchable and needs no vision model; render when OCR is unavailable or would mangle what matters.
Every page is saved as a PNG artifact and its id returned. With view, the
default, the first few images also come back to be looked at, since images are
expensive in context. Limits are explicit: 20 pages per call, and resolution is
reduced for a page that would be enormous — both reported, never silent.
Args: ref: PDF file path, or a workspace artifact id. pages: Which pages: "1-20", "3", "1,5,9-12" or "all". Defaults to all, up to the per-call limit. dpi: Resolution, 36 to 600. 150 suits reading and previewing. view: Return the images to look at, not just artifact ids.
| Name | Required | Description | Default |
|---|---|---|---|
| dpi | No | ||
| ref | Yes | ||
| view | No | ||
| pages | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden. It discloses that every page is saved as a PNG artifact, that with view=true images are returned, that limits exist (20 pages per call), and that resolution reduction is reported. It even explains why images are expensive in context. This is highly transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured and front-loaded. It starts with the core purpose, then usage guidance, then behavioral details, then parameter explanations. Each sentence adds value; there is no redundancy or fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema and no annotations, the description is complete: it covers what the tool does, when to use it, how it behaves (including limits and side effects), and all parameters with defaults and examples. An agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains each parameter: ref (file path or artifact id), pages (formats like '1-20', 'all', defaults to all), dpi (range 36-600, suggests 150), and view (returns images vs ids). This adds significant meaning beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb (Render) and resource (PDF pages to images) and explains the purpose: to look at a page rather than read it. It distinguishes from siblings by explicitly naming pdf_ocr as the alternative for reading, and gives concrete use cases (checking redaction, split, handwriting, signatures, charts).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to use this tool ('Use it to see what a page looks like') and when not ('For reading a scan prefer pdf_ocr'), with conditions (when OCR is unavailable or would mangle). This is clear guidance with alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.3.0- Added
pdf_extract_text - Added
pdf_ocr
10 tool updates
v0.1.1- First observed
create_test_pdf_file - First observed
discard - First observed
export - First observed
list_artifacts - First observed
pdf_check_access - First observed
pdf_check_text - First observed
pdf_metadata - First observed
pdf_page_count - First observed
pdf_page_layout - First observed
pdf_render_pages
TDQS
Scored across 12 tools
Each tool targets a distinct PDF inspection or artifact management action. The pdf_* tools are clearly separated by function (page count, metadata, text check, extraction, access, layout, render, OCR), and the artifact tools (discard, list, export) serve unique purposes. No overlapping responsibilities.
PDF tools follow a consistent 'pdf_' prefix with descriptive names (pdf_page_count, pdf_extract_text), but artifact management tools (discard, export, list_artifacts, create_test_pdf_file) lack the prefix and follow a different verb/noun style. The split is understandable but slightly inconsistent.
12 tools is well-scoped for a PDF utility server. Each tool fills a distinct role in inspecting and processing PDFs, plus artifact lifecycle management. The count is neither bloated nor thin for the apparent domain.
The surface covers the essential PDF inspection workflows: determining page count, text presence, extraction, OCR, metadata, access restrictions, page layout, and rendering. Artifact handling (create, list, discard, export) is complete. No obvious missing operations for the server's stated purpose.
Maintenance
Related MCP Connectors
Generate and read PDFs for AI agents: a generate_pdf and a read_pdf tool, priced per document.
PDF, image, video, OCR, screenshot, SQL, QR and text tools for agents. No API key, no signup.
Image & PDF tools for AI agents: compress, convert, resize, PDF, AI vision, pipeline.
Merge, split, extract, rotate, reorder and stamp PDF pages from your AI chat, all offline.
Related MCP Servers
- AlicenseAqualityDmaintenanceEnables AI agents to securely read and extract information from PDF files including text content, metadata, and page counts from both local files and URLs within the project context.12,180 npmMIT
- FlicenseNot gradedqualityCmaintenanceEnables AI agents to perform 13 PDF operations (merge, split, compress, watermark, encrypt, and more) on local files via MCP.3-
- AlicenseAqualityCmaintenanceEnables AI agents to process and inspect PDFs by detecting document type, extracting text, and generating Markdown with layout information.71MIT
- FlicenseNot gradedqualityAmaintenanceEnables AI agents to perform comprehensive PDF operations locally, including compression, text extraction, PII redaction, page organization, splitting, merging, watermarking, creation, and form filling, all without cloud uploads.6 npm-