Skip to main content
Glama
ianderso

nara-catalog-mcp

by ianderso

search_extracted_text

Read-only

Find US National Archives Catalog records by words in OCR or contributed text, then check page images before citing.

Instructions

Find records whose extracted text mentions something.

This reaches text contributed by NARA's digitisation partners over the digital objects -- the searchable layer under the scans. It is machine output and is not evidence: it misses handwriting, it mangles names, and a hit means a page probably says this, not that it does. Read the page image before citing anything you find here.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
pageNoPage of results, 1-based.
limitNoMaximum records to return (1-100).
queryYesWords to find in extracted text. Accepts AND, OR, NOT, wildcards and "exact phrases".

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A3.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint and openWorldHint, but the description adds substantial behavioral context the annotations cannot: this is machine output that misses handwriting and mangles names, and a hit means a page 'probably' says this rather than proving it. It sets a clear expectation about result reliability and mandates verifying against the page image – genuinely useful disclosure beyond the structured fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first sentence, and the subsequent caveat about reliability is well placed and each clause carries distinct information (misses handwriting, mangles names, probabilistic hits). It is a touch long but nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only search tool with no output schema and full schema coverage on the three params, the definition covers purpose and the critical reliability caveat an agent needs before citing results. It does not address result ordering or how pages/pagination interplay, but those are minor given the schema and annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so query/page/limit (including query syntax and the 1-100 limit range) are fully documented in the schema. The description adds no parameter-level meaning beyond that, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence gives a specific verb and resource: 'Find records whose extracted text mentions something.' The follow-up explains the nature of the corpus (OCR/searchable layer under scans contributed by digitisation partners), which helps position it. It does not, however, explicitly distinguish it from close siblings like search_transcriptions or search_by_contribution_text.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: it is for finding text mentions, with the caveat that hits are not evidence and the page image should be read before citing. There is no explicit when-to-use-vs-alternatives routing to the many sibling search tools (search_records, search_transcriptions, search_by_contribution_text), which is a real gap given how crowded this family is.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.