Skip to main content
Glama

r7_pdf_inspect

Measure a PDF's text fidelity: inventory pages/fonts, count drawn glyphs, test text extraction, and return a text/outlined/mixed verdict; optionally repair malformed ToUnicode CMap entries.

Instructions

Measure the text fidelity of a PDF: page and font inventory, how many glyphs are really drawn by text operators, how many glyph fills painted nothing, whether the text is extractable, and a verdict (text/outlined/mixed). Use it to prove a PDF export kept its text instead of silently losing it; set repair to fix malformed ToUnicode CMap entry counts that make R7-exported body text copy out as garbage.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
repairNoRepair malformed ToUnicode CMap entry counts in place (rewrites filePath) before reporting. Default false.
filePathYesPath to the PDF file to inspect.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.1

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided the description carries the full burden, and it discloses the important behavioral trait that setting repair mutates the file in place to fix CMap entry counts, plus the concrete symptom (text copies out as garbage). It does not explicitly state that the default path is non-mutating or mention permission/auth needs, leaving a small gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with purpose before the use case and parameter guidance, so nothing is buried. It is dense with enumerated clauses, which slightly taxes readability, but every clause carries distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Though there is no output schema, the description effectively specifies the return content (page/font inventory, drawn-glyph counts, empty-fill counts, extractability, verdict), and it covers purpose, usage, and the single non-trivial parameter. Nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already 100%, so the baseline is 3, but the description adds meaning beyond the schema by explaining what repair actually fixes (malformed ToUnicode CMap entry counts) and the real-world consequence it addresses (R7-exported body text copying out as garbage). That extra context exceeds the schema's drier 'rewrites filePath' wording.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Measure the text fidelity of a PDF') and enumerates exactly what is measured (page/font inventory, drawn glyph counts, empty glyph fills, extractability, a text/outlined/mixed verdict). The diagnostic scope distinguishes it from the generic sibling r7_inspect without needing to open either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear triggering scenario: 'Use it to prove a PDF export kept its text instead of silently losing it,' and names the condition that selects the repair mode (malformed ToUnicode CMap making body text copy as garbage). It does not explicitly contrast with the sibling r7_inspect or state exclusions, so it falls just short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.