ia_grep_text
Search an Internet Archive item's OCR text for a word or phrase and return every match with surrounding context and character offsets, including across OCR line breaks.
Instructions
Find every occurrence of a word or phrase in an item's OCR text, with surrounding context and offsets.
The text is fetched once and held in memory, so repeated searches of the same book are fast.
Args:
identifier: The Internet Archive identifier.
pattern: Text to find. A plain phrase matches across the line breaks, repeated spaces and split
hyphens that OCR leaves in scanned books ("twenty-five prisoners" finds "twenty-five prisoners" and
"twenty- five prisoners"). A regular expression only if regex is true.
regex: Treat the pattern as a regular expression.
ignore_case: Ignore capitalisation.
context: Characters of context on each side (0 to 1000).
max_matches: Most matches to return (1 to 100); the total is always reported.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| regex | No | ||
| context | No | ||
| pattern | Yes | ||
| identifier | Yes | ||
| ignore_case | No | ||
| max_matches | No |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||