Skip to main content
Glama

pdf_extract_text

Extract the text content of a PDF page by page, returning the exact characters from the file. Use it to read or quote from text-based PDFs, with optional page ranges and layout preservation.

Instructions

Read the text of a PDF that has one, page by page.

The cheap and exact way to read a document: milliseconds a page, and the characters are the ones the file holds. Prefer it to pdf_ocr, which is for pages carrying no text, and to pdf_render_pages, which costs an image a page. pdf_check_text says which path a file needs; page numbers come back with the text, so quote page 8 rather than "the document".

A long document arrives in ranges: one call returns about 50,000 characters and names the pages it did not reach, so when truncated, do what the summary says. output="txt" writes the whole extraction to an artifact and returns the counts alone, which is the way to read a book.

Text on a text_suspect page came back partly undecodable, because the fonts carry no character map: say so rather than quoting it, and look at the page with pdf_render_pages.

Args: ref: PDF file path, or a workspace artifact id. pages: Which pages: "1-10", "3", "1,5,9-12" or "all". Defaults to all. output: "text", "txt" for the text as an artifact, or "both". layout: Keep the page's spacing, for forms and tables. Costs characters.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
refYes
pagesNo
layoutNo
outputNotext

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv0.3.0

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does so thoroughly. It discloses performance characteristics (milliseconds per page), exactness of characters, truncation behavior around ~50,000 characters, what `truncated` means, artifact-writing behavior for output='txt', and the text_suspect caveat about undecodable fonts. This is exemplary behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is detailed but every sentence earns its place: each clause either defines a behavior, gives usage guidance, or explains a parameter. The first sentence is a clear one-line purpose, and the nuanced details are logically organized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity, the description is complete: it covers how to invoke it, what to expect, how to handle truncation, when to choose alternatives, and how to handle suspicious output. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. The Args section adds meaningful semantics for all four parameters: ref accepts a file path or artifact id, pages gives concrete syntax examples, output enumerates valid values and their effects, and layout explains when it matters. This fully compensates for the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: "Read the text of a PDF... page by page." It clearly differentiates the tool from pdf_ocr, pdf_render_pages, and pdf_check_text, so an agent knows exactly what this tool does and what it does not do.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage guidance is explicit and actionable: prefer this over pdf_ocr for pages with text, over pdf_render_pages when exact cheap text is needed, consult pdf_check_text to choose a path, use output='txt' for long books, and use pdf_render_pages for text_suspect pages. This leaves no ambiguity about when to select this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.