pdf_extract_text
Extract the text content of a PDF page by page, returning the exact characters from the file. Use it to read or quote from text-based PDFs, with optional page ranges and layout preservation.
Instructions
Read the text of a PDF that has one, page by page.
The cheap and exact way to read a document: milliseconds a page, and the characters are the ones the file holds. Prefer it to pdf_ocr, which is for pages carrying no text, and to pdf_render_pages, which costs an image a page. pdf_check_text says which path a file needs; page numbers come back with the text, so quote page 8 rather than "the document".
A long document arrives in ranges: one call returns about 50,000 characters
and names the pages it did not reach, so when truncated, do what the summary
says. output="txt" writes the whole extraction to an artifact and returns
the counts alone, which is the way to read a book.
Text on a text_suspect page came back partly undecodable, because the fonts
carry no character map: say so rather than quoting it, and look at the page
with pdf_render_pages.
Args: ref: PDF file path, or a workspace artifact id. pages: Which pages: "1-10", "3", "1,5,9-12" or "all". Defaults to all. output: "text", "txt" for the text as an artifact, or "both". layout: Keep the page's spacing, for forms and tables. Costs characters.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| ref | Yes | ||
| pages | No | ||
| layout | No | ||
| output | No | text |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| result | Yes |