Get Document Text
get_document_contentReturn a document's OCR'd text, limited to 20,000 characters per request. Provide the offset from a truncation marker to read subsequent sections.
Instructions
Return the OCR'd text content of a document.
Documents such as books and technical standards can run to millions of
characters, so each call is capped at 20,000. A partial result opens
with a marker naming the character range returned, the document's full
length, and the offset to pass to read the next section; text that
fits under the cap is returned whole with no marker.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| offset | No | Character position to start reading from. Pass the value named in a truncation marker to continue from where it stopped. | |
| max_chars | No | Maximum number of characters to return, up to 20,000. | |
| document_id | Yes | ID of the document to read. |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| result | Yes |