document_extract
Count all elements and list requested element types from a PDF/DOCX/MD/TXT file, paginate results, and optionally save the inventory as JSON.
Instructions
A mechanical inventory of one PDF/DOCX/MD/TXT - not a substitute for reading it. Counts always; lists only for the kinds asked (numbers, identifiers, dates, urls, emails, headings, tables, links, emphasis, images, pages, blocks, warnings, unsupported, normalized_text), paged by offset/max_items with the rest stated. output_path writes everything to JSON. [docbridge schema 0.2.4]
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| kinds | No | ||
| offset | No | ||
| encoding | No | Text encoding of a .md/.txt input. Omit for UTF-8; docbridge never guesses. | |
| max_items | No | ||
| overwrite | No | ||
| input_path | Yes | Absolute file path. | |
| output_path | No | ||
| text_offset | No |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| tool | Yes | ||
| error | No | ||
| document | No | ||
| disclaimer | No | docbridge never alters, summarizes or silently truncates source evidence. docbridge reports what it compared. PASS on an axis covers only that axis's stated scope; NOT_CHECKED and UNSUPPORTED are never passes. PDF text is the text layer as PyMuPDF decodes it, not the rendered glyphs. Nothing here interprets meaning: numbers are compared as characters, not as values. | |
| written_to | No | ||
| schema_version | Yes | ||
| operation_completed | Yes |