get_extracted_text
Fetch machine-extracted OCR text from a NARA Catalog record's page images to identify relevant pages, then verify against the original images.
Instructions
Read the OCR text NARA machine-extracted from a record's page images.
This is a lead, not evidence. OCR was run over scans of handwriting,
carbon copies and microfilm: it drops handwritten pages entirely, and it
turns one surname into another silently. Use it to find which page matters,
then open that page image with get_record_images and cite what you saw
there -- never cite the OCR text itself.
Returns one entry per digital object, in page order, with the text
truncated to max_chars.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| naid | Yes | The record's NAID. | |
| page | No | Page of results, 1-based. | |
| limit | No | Maximum digital objects to return (1-100). | |
| refresh | No | True re-reads from the Catalog instead of the cache, spending one call and replacing the cached copy. The cache never expires on its own, so use this when the answer may have changed since you last asked. | |
| max_chars | No | Characters of text to return per page. 0 returns all of it, which for a long file unit is tens of thousands. | |
| object_id | No | Limit to one digital object (one scanned page). Omit it to get the text of every page of the record. |