Skip to main content
Glama
magicianmarty

heritage-research-mcp

ia_grep_text

Search an Internet Archive item's OCR text for a word or phrase and return every match with surrounding context and character offsets, including across OCR line breaks.

Instructions

Find every occurrence of a word or phrase in an item's OCR text, with surrounding context and offsets.

The text is fetched once and held in memory, so repeated searches of the same book are fast.

Args: identifier: The Internet Archive identifier. pattern: Text to find. A plain phrase matches across the line breaks, repeated spaces and split hyphens that OCR leaves in scanned books ("twenty-five prisoners" finds "twenty-five prisoners" and "twenty- five prisoners"). A regular expression only if regex is true. regex: Treat the pattern as a regular expression. ignore_case: Ignore capitalisation. context: Characters of context on each side (0 to 1000). max_matches: Most matches to return (1 to 100); the total is always reported.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
regexNo
contextNo
patternYes
identifierYes
ignore_caseNo
max_matchesNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.3

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does well: it discloses the in-memory fetch-and-cache behavior (performance trait), and that "the total is always reported" even when matches are truncated. It omits auth/rate-limit details, but for a read-only grep these are minor.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded purpose sentence followed by a performance note, then a clean Args block. Every sentence earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values needn't be explained, and all six parameters are documented. The only gap is the lack of explicit routing guidance versus sibling tools, which the description could have added.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: it explains normalization behavior for the pattern (line breaks, repeated spaces, split hyphens), the meaning of the regex flag, the context range (0–1000), and the max_matches range plus total reporting. Every parameter gains meaning beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states a specific verb ("Find every occurrence") and resource ("a word or phrase in an item's OCR text") plus the return shape ("with surrounding context and offsets"). This clearly distinguishes it from siblings like ia_read_text and ia_fulltext_search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage through "repeated searches of the same book are fast," suggesting it's for iterative querying of one item, but it never explicitly says when to choose this over ia_read_text or ia_fulltext_search. Usage context is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.