nara-catalog-mcp
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@nara-catalog-mcpfind Civil War pension records for John Murphy and show transcriptions"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
nara-catalog-mcp
An MCP server for genealogical research in the US National Archives Catalog. Find an ancestor's pension file, service record or census page, read what machines and volunteers have transcribed from it, and fetch the page images you will cite.
It works the way a careful genealogist does. A catalogue description is a finding aid, and OCR text, transcriptions and tags are somebody else's reading of the page: all of them are leads. The page image is the source. The tools say which is which, and the server tells the model to read the image before citing it.
Nothing here writes to the Catalog, and nothing here keeps a family tree. The server finds and reads records; what you conclude from them belongs in your genealogy software, or in a family-tree MCP server running alongside this one.
The tools work on any record the Catalog describes, so a historian or a journalist can use them too. The examples throughout are about tracing people.
This is an independent project. It is not affiliated with, endorsed by, or supported by the National Archives and Records Administration.
Tools
The server publishes 18 tools. Seventeen only read, and are annotated
read-only for the client. One, download_page_image, writes a new file to
local disk; it never overwrites one.
Finding a record
Tool | Purpose |
| Search by title or full text. Returns totals plus compact summaries: NAID, hierarchy, holding unit, date coverage, image count. |
| The same search with the filters that narrow a common name: date range, record group, collection, ancestor NAID, level of description, microfilm publication number, local identifier, any control number, Congress number, reference unit, creator, person or organisation, place, type of materials, recurring date, digitised-only, "has transcriptions/tags/comments", exact matching of identifiers, |
| The immediate children of a node: record group to series to file unit to item. Walk down from a series you trust. |
Reading a record
Tool | Purpose |
| Read one record in full by NAID: scope and content note, hierarchy, reference units, every page image with its object id. |
| Page images in page order, each with its object id and URL — the evidence to read before citing. |
| Save one page to a new local file, by page number or by the object id an OCR or transcription entry names. The media host is open, so this spends no API quota. |
| Where else the record is published online, including by commercial partners. |
| Digital object ids a commercial partner (Ancestry among them) has matched to this NAID. An empty list is the common answer. |
What machines and other people read in it
Everything in this group is somebody else's reading of the document. It tells you which page to open. It does not tell you what the page says.
Tool | Purpose |
| OCR text extracted from the page images, per digital object, with who produced it, and the object total and next results page when a file runs longer than one call. |
| Citizen transcriptions — how you get into a handwritten pension file that OCR cannot touch. |
| Citizen tags, which on genealogical records are usually the names of the people inside the file. |
| Other researchers' notes on the record. |
Searching the documents rather than the catalogue
Tool | Purpose |
| Search transcriptions, tags, comments and OCR text — and get full record summaries back. |
| Find records whose transcribed text mentions a name or phrase. |
| Find records carrying a tag, exactly or by word. |
| Find records whose partner-contributed extracted text matches. |
| Find records other researchers have commented on. |
Housekeeping
Tool | Purpose |
| Live calls and cache hits this session, the month's live calls from a ledger beside the cache, and what remains of the key's monthly allowance. |
Related MCP server: nara-mcp-server
Setup
You need Python 3.11 or later, uv, and a Catalog API key. Request a key from the Catalog API team via the API documentation; the default allowance is 10,000 calls per month.
There are two ways to run the server.
Without cloning. uvx fetches it from PyPI and runs it in one step, and
caches the result:
NARA_API_KEY=your-key uvx nara-catalog-mcpFrom a clone, which is what you want if you will change it:
git clone https://github.com/ianderso/nara-catalog-mcp
cd nara-catalog-mcp
uv sync
cp .env.example .env # then put your key in it
uv run nara-catalog-mcp # stdio server, usually launched by the clientEither way the server speaks MCP over stdio, so you will normally let an MCP client start it rather than run it by hand.
Claude Desktop
Without cloning:
{
"mcpServers": {
"nara": {
"command": "uvx",
"args": ["nara-catalog-mcp"],
"env": { "NARA_API_KEY": "your-key" }
}
}
}From a clone (the env block can be dropped if the key is in the clone's
.env):
{
"mcpServers": {
"nara": {
"command": "uv",
"args": ["--directory", "/path/to/nara-catalog-mcp", "run", "nara-catalog-mcp"],
"env": { "NARA_API_KEY": "your-key" }
}
}
}A desktop app does not always inherit your shell's PATH. If the server fails
to start because uv or uvx cannot be found, give the full path that
which uvx prints as the command.
Claude Code
claude mcp add nara --env NARA_API_KEY=your-key -- uvx nara-catalog-mcpor, from a clone whose .env holds the key:
claude mcp add nara -- uv --directory /path/to/nara-catalog-mcp run nara-catalog-mcpConfiguration
All settings come from the environment. A .env file in the directory the
server starts in supplies any that the environment does not; with
uv --directory that is the clone. Only that directory is read — not its
parents, and not the directory the package is installed in.
Variable | Meaning |
| Your Catalog API key. Required. |
| Response cache directory. Default |
| HTTP timeout in seconds. Default 60. |
| Calls per month the key allows, reported by |
A missing key, or an unusable value, is reported on the first tool call as a
not_configured result naming the variable.
The call budget
A Catalog key is capped per month, and a sweep across many names will exhaust
it faster than expected. Every successful response is cached on disk, keyed by
endpoint and query, so repeating a call costs nothing. Live calls are also
written to a ledger, budget.json beside the cache, so api_budget reports
the month's spend across sessions as well as this session's live calls and
cache hits. A rejected call counts as live, since it reached the API. The
ledger sees one machine and one cache directory; the Catalog's own count is
the authority.
The cache never expires. Pass refresh=true to get_record,
get_record_images, get_extracted_text, get_transcriptions, get_tags
or get_comments to re-read one answer from the Catalog: that spends one
call and replaces the cached copy. New transcriptions and tags arrive over
time, and a file that was paper-only can be digitised, so a cached "nothing
here" is the answer most worth refreshing. Deleting NARA_CACHE_DIR still
works, and also resets the month's ledger.
Searching past the first 10,000 hits
page stops working beyond 10,000 results, which a common surname passes.
search_records_advanced reports a paging_note when you hit that boundary;
either narrow the search or re-run it with search_after="*" and follow the
next_search_after value from each response.
How a search behaves
title matches words in the record title and is the precise option — case
files are usually titled with the person's name. query searches the full
description: broader, noisier, and worth reaching for only when a title search
finds nothing.
Results come back as summaries rather than whole records. A raw Catalog record is large and mostly irrelevant to the decision you are making, which is whether this record is worth opening.
Notes
A description is not evidence. It summarises a file; it says nothing about what any individual page contains. Read the images before citing.
Neither is OCR, and neither is a transcription.
get_extracted_textis a machine reading a scan of handwriting;get_transcriptionsis one volunteer's typing, unreviewed. Both are the fastest way to find the page that matters, and neither is a source. Open the image and cite that.Record the NAID. It is the stable identifier that makes a citation refindable.
Not every record is digitised.
image_count: 0means the description exists but the pages are not online — thereference_unitsfield tells you which archive holds the paper.Open is not unrestricted. The media host needs no key, and most of NARA's holdings are in the public domain, but not all: some descriptions record use restrictions. Check before republishing an image.
Deliberately not here
The Catalog API can post tags and comments and put transcriptions. This server does not, and will not by default: a contribution publishes under whoever's key is configured, and nothing here can take it back.
Security
The key is read from the environment, sent only to
catalog.archives.govas thex-api-keyheader, and never sent to the media host. It is not written to the response cache.Contributions are untrusted text. Tags, comments and transcriptions are written by members of the public and reach the model verbatim, so one could contain instructions aimed at it. The server's instructions tell the model to treat that text as material to weigh, never as instructions; the model is still the one deciding, so review what it proposes to do.
download_page_imagewrites files. It creates a new file wherever the server's user can write, and refuses to overwrite an existing one — so a mistaken or injected path cannot destroy anything. It is annotated as not read-only, so a client can ask before each call.
To report a vulnerability, see SECURITY.md.
Development
uv sync --extra dev
uv run pytest # mocked with respx; no API key needed
uv run ruff check .
uv run ruff format --check .
uv run python -m tests.live_check # needs a key: asks the Catalog what the mocks cannotThe live check asks the Catalog what the mocks cannot: where the two text flags put their text, and which search parameters the server does not send. Run it after a change on either side; a difference shows up in its output.
See CONTRIBUTING.md for how the suite is organised and what a change is expected to carry.
API notes
The published OpenAPI spec is unreliable in specific, repeatable ways — wrong
required flags, maximum used where maxLength is meant, and two
conflicting definitions of the pagination cursor. Corrections verified against
the live Catalog are in docs/API-NOTES.md. Why the server
is shaped the way it is, what is out of scope by decision, and how the test
suite is built are in docs/DESIGN.md.
License
MIT.
Available Tools
18 toolsapi_budgetARead-only
Report the key's Catalog API spend, this session and this month.
A key is capped per month and a long sweep can exhaust it. Live calls are kept in a ledger beside the cache, so the month's figure survives restarts. Cached repeats cost nothing and are not counted.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds real behavioral context beyond that: the monthly cap, ledger persistence surviving restarts, and the fact that cached repeats are free and uncounted — quota and accounting semantics the annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose in the first sentence, then supporting context. Three sentences with essentially no waste, though 'Live calls are kept in a ledger beside the cache' is mildly atmospheric for the information it conveys.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must convey what comes back; it does so at the level of 'spend, this session and this month'. It omits the return format (currency, numeric vs. text) and any cap-remaining detail, which is a minor gap for a simple reporting tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. Nothing in the description misleads about inputs, and there is nothing further to document.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (report) and resource (the key's Catalog API spend) with scope qualifiers (this session and this month). It is unmistakably distinct from every sibling, all of which are search/get/download tools rather than budget introspection.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description supplies the condition that makes the tool relevant ('a key is capped per month and a long sweep can exhaust it'), which implicitly tells the agent to check spend before large operations. It stops short of naming an alternative tool or an explicit when-not-to-use rule, but the single-purpose, zero-parameter nature leaves little routing ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
browse_childrenARead-only
List a record's immediate children, one level down the hierarchy.
The Catalog nests record group, then series, then file unit, then item.
Searching finds a node; this walks down from it, which is how you get from
a series you trust to the file unit for one person, and how you find out
what else sits alongside a file you already have. A record's own ancestors
come back from get_record.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page of results, 1-based. | |
| limit | No | Maximum children to return (1-100). | |
| parent_naid | Yes | NAID of the parent: a record group, collection, series or file unit. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, so safety is covered. The description adds real domain context beyond that: the four-level Catalog hierarchy (record group > series > file unit > item) and the traversal direction, which tells the agent what 'children' actually returns. It stops short of describing result size or ordering behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action in the first clause, then supporting context. The illustrative 'one person' example is mildly verbose but still earns its place by anchoring the hierarchy in a real use case.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only list tool with no output schema, the description gives the agent enough to know what comes back (immediate children, one level) and the hierarchy semantics. Pagination defaults live in the schema, so nothing critical is missing, though return-shape details are left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so parent_naid, page, and limit are already documented in the schema. The description adds no syntax, format, or constraint detail for those parameters, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('List a record's immediate children') with the scope qualifier 'one level down the hierarchy,' which immediately separates it from the recursive/ancestor sibling get_record. An agent can distinguish it from search_records without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly frames the workflow: 'Searching finds a node; this walks down from it,' and names the opposite operation ('A record's own ancestors come back from `get_record`'). Both the when-to-use and the alternative are stated, not implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
download_page_imageAIdempotent
Download one page of a record so it can actually be read.
This is the step that turns a catalogue hit into evidence. The image is written to disk rather than returned inline — a page scan runs to several megabytes.
NARA's media URLs are open and need no key, so this costs nothing against your API allowance. One catalogue call resolves the page list; the download itself is not an API call.
A pension file can run to sixty pages and the page you need is rarely the first. Use get_extracted_text or get_transcriptions to find which page carries the fact, then fetch it by the object_id they name.
| Name | Required | Description | Default |
|---|---|---|---|
| naid | Yes | The record's NAID, e.g. '54765873'. | |
| page | No | Which page to download, 1-based, in the order get_record_images lists them. Defaults to 1 when object_id is not given. | |
| object_id | No | The digital object to download, as get_extracted_text, get_transcriptions and get_record_images name it. An alternative to page: pass one or the other. | |
| destination | Yes | Absolute path of a new file to write, e.g. '/tmp/hall-w17050-p51.jpg'. The directory must already exist, and the file must not: nothing is ever overwritten. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With readOnlyHint=false, openWorldHint=true and idempotentHint=true already declared, the description still adds substantive context: the output is a multi-megabyte file written to disk rather than returned inline, the media URLs require no key, and the download does not consume API allowance. That is exactly the side-effect and cost information an agent needs before invoking.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then layered with cost, disk-write, and page-selection guidance in a logical order; every sentence carries information. Slightly prose-heavy ('turns a catalogue hit into evidence') but not padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, but the description covers what is produced (a file on disk, not a return value), its size, the absence of auth or budget cost, and how to choose the right page. Nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema already documents naid, page, object_id and destination in detail, so the baseline is 3. The description goes slightly further by framing page and object_id as mutually exclusive alternatives and pointing at the sibling tools that produce a usable object_id, adding a little routing value on top of the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Download one page of a record') and clarifies the scope ('so it can actually be read', 'the image is written to disk rather than returned inline'). It differentiates itself from catalogue-oriented siblings by noting that a catalogue call only resolves the page list, not the image itself.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit routing: 'Use get_extracted_text or get_transcriptions to find which page carries the fact, then fetch it by the object_id they name.' It also supplies the motivating condition (a pension file can run to sixty pages, so the needed page is rarely the first), which tells the agent when this tool is the right step versus when to search first.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_commentsARead-only
Read other researchers' comments on a record.
Comments often say what a file actually contains, where a related file sits, or that the description is wrong. None of it is verified by NARA: it is correspondence between researchers, useful for finding the next record and never a source in itself.
| Name | Required | Description | Default |
|---|---|---|---|
| naid | Yes | The record's NAID. | |
| refresh | No | True re-reads from the Catalog instead of the cache, spending one call and replacing the cached copy. The cache never expires on its own, so use this when the answer may have changed since you last asked. | |
| max_chars | No | Characters of each comment to return. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint=true, openWorldHint=true), but the description adds genuinely non-obvious context: the comments are user-generated correspondence, unverified by NARA, and not authoritative. That trust caveat is exactly the kind of information an agent would otherwise have to infer.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences with the purpose front-loaded, followed by content value and then the trust caveat. No filler; every sentence carries information an agent needs to use the result responsibly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description does the important work of characterizing the return content and warning against treating it as a source. It stops short of describing volume, paging, or what an empty result means, but nothing critical is missing for a read-only lookup.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the schema itself documents naid, refresh (cache behavior and cost) and max_chars. The description adds no parameter syntax or defaulting detail beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read other researchers' comments on a record') and immediately frames what the comments contain, so an agent can distinguish it from search_comments without opening either schema. The scope is narrow and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains how to interpret the output ('useful for finding the next record and never a source in itself') but never says when to call this tool versus alternatives such as search_comments or get_record, nor any prerequisite. Usage is only implied by the content framing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_extracted_textARead-only
Read the OCR text NARA machine-extracted from a record's page images.
This is a lead, not evidence. OCR was run over scans of handwriting,
carbon copies and microfilm: it drops handwritten pages entirely, and it
turns one surname into another silently. Use it to find which page matters,
then open that page image with get_record_images and cite what you saw
there -- never cite the OCR text itself.
Returns one entry per digital object, in page order, with the text
truncated to max_chars.
| Name | Required | Description | Default |
|---|---|---|---|
| naid | Yes | The record's NAID. | |
| page | No | Page of results, 1-based. | |
| limit | No | Maximum digital objects to return (1-100). | |
| refresh | No | True re-reads from the Catalog instead of the cache, spending one call and replacing the cached copy. The cache never expires on its own, so use this when the answer may have changed since you last asked. | |
| max_chars | No | Characters of text to return per page. 0 returns all of it, which for a long file unit is tens of thousands. | |
| object_id | No | Limit to one digital object (one scanned page). Omit it to get the text of every page of the record. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover read-only/open-world safety, and the description adds substantial unannotated behavior: OCR drops handwritten pages entirely, silently mangles names, output is one entry per digital object in page order, and text is truncated to max_chars. These limitations are exactly what an agent needs to avoid misusing the results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then the critical caveat, then the recommended workflow, then return shape. Every sentence carries distinct useful information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, and the description compensates by describing the return structure (one entry per digital object, page order, truncated text). Combined with rich annotations and a fully documented schema, an agent has everything needed to call and interpret this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all six parameters are already documented, including the refresh cache semantics and object_id scoping. The description only reinforces max_chars truncation and page ordering, adding marginal meaning beyond the schema, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource — reads machine-extracted OCR text from a record's page images — and implicitly distinguishes itself from siblings like search_extracted_text (search vs. read) and get_transcriptions (different source). An agent can tell what it retrieves without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly frames the tool as a lead, not evidence, and gives a concrete workflow: use it to locate the page that matters, then open that page with `get_record_images`, and never cite the OCR text itself. This names both the when and the when-not plus the alternative tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_online_availabilityARead-only
List where else this record is available online.
NARA records what has been digitised and published elsewhere, including by commercial partners. Use it to resolve a hint from a subscription site back to the archival original: the NAID and the reference unit are what make a citation refindable, and a partner's index entry is not a substitute for the page.
| Name | Required | Description | Default |
|---|---|---|---|
| naid | Yes | The record's NAID. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover safety (readOnlyHint) and scope (openWorldHint), so the bar is lower; the description still adds real interpretive context — that NARA tracks digitisation by commercial partners and that a partner index entry should not replace the original page. That is useful behavioral framing not present in the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose is front-loaded in the opening sentence and the follow-up earns its place by explaining provenance rationale. The closing aside about citations is slightly discursive but not wasteful for a tool with no output schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, read-only lookup with no output schema, the description adequately conveys what is returned (other online holdings, partner copies) and why it matters. Only the ambiguity with get_partner_digital_objects keeps it from being fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for the single naid parameter, so the baseline is 3. The description gestures at NAID's importance for citations but adds no format, syntax, or sourcing detail beyond what the schema already states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states a specific verb and resource — listing alternative online locations for a record — which is clearer than a tautology. It does not, however, differentiate itself from the near-twin sibling get_partner_digital_objects, so an agent must still reason about which one to call.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It supplies one concrete scenario ('resolve a hint from a subscription site back to the archival original'), which is better than nothing, but it gives no when-not-to-use guidance and never mentions get_partner_digital_objects as the alternative for partner-hosted objects.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_partner_digital_objectsARead-only
List digital object ids a partner has matched to this record.
NARA indexes metadata supplied by commercial partners — Ancestry among them — against its own records. A hit tells you the partner holds imagery for this NAID, which is worth knowing when NARA's own pages are not online.
An empty list is the common answer and is not an error. It means no partner metadata has been matched, not that no partner holds the record.
The ids are pointers into the partner's index, not a citation. Cite the archival record by NAID and reference unit.
| Name | Required | Description | Default |
|---|---|---|---|
| naid | Yes | The record's NAID, e.g. '54765873'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnlyHint, openWorldHint), and the description adds genuinely non-obvious behavior: an empty list is the common, non-error outcome, and it means no partner metadata was matched rather than no partner holding the record. It also warns the ids are pointers into a partner index, not a citable reference. Return format is described in prose since no output schema exists.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action and result, then two short paragraphs of caveats. Every sentence carries information, though the four-paragraph layout is slightly heavier than the tool's complexity warrants.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden itself and does so well: it states what the list contains, what an empty result means, and how the ids should (not) be used for citation. Nothing needed to call or interpret this tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% for the single 'naid' parameter, and the schema already gives an example value. The description only refers to it obliquely as 'this record' and adds no format or lookup guidance beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence gives a specific verb and resource ('List digital object ids a partner has matched to this record') and scopes it to partner-supplied metadata, which cleanly separates it from siblings like get_online_availability, get_record_images, and get_extracted_text. An agent can pick this tool without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description supplies a clear use condition — it is 'worth knowing when NARA's own pages are not online' — which implicitly routes the agent to this tool after checking NARA's own availability. It does not name the sibling tool (e.g., get_online_availability) explicitly, so the alternative is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_recordARead-only
Read one Catalog record in full, by NAID.
Returns the scope and content note, the full hierarchy, the holding reference units, and every page-image URL. Use this once a search has given you a NAID worth pursuing.
| Name | Required | Description | Default |
|---|---|---|---|
| naid | Yes | The record's NAID, e.g. '54765873'. | |
| refresh | No | True re-reads from the Catalog instead of the cache, spending one call and replacing the cached copy. The cache never expires on its own, so use this when the answer may have changed since you last asked. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and openWorldHint, so the safety profile is covered; the description adds useful behavioral context by disclosing the full payload it returns rather than a summary. The caching/re-read behavior of refresh is described, though only in the schema, not the description.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, and the core action and its trigger ('once a search has given you a NAID') are front-loaded. The list of returned content is compact rather than padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so by enumerating the four content categories. Combined with the annotation-provided safety profile and fully covered parameters, nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both naid and refresh are fully documented in the schema, including the cache-replacement semantics of refresh. The description only echoes that lookup is by NAID, adding no syntax or format detail beyond structured data, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read one Catalog record in full, by NAID') and enumerates what 'in full' means (scope/content note, hierarchy, holding reference units, page-image URLs). An agent can distinguish this single-record fetch from the sibling search_* tools without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly sequences the tool: 'Use this once a search has given you a NAID worth pursuing,' which ties it to the search_* siblings as the precondition. It gives clear when-to-use context but does not name a specific alternative or state exclusions (e.g. when to prefer browse_children or get_record_images for sub-parts).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_record_imagesARead-only
List a record's page images in page order, each with its object id.
These are the evidence. A catalog description summarises a file; it does not tell you what any individual page says, so read the images before citing anything to this record. The object id is what get_extracted_text and get_transcriptions entries point at; pass it, or the page number, to download_page_image.
| Name | Required | Description | Default |
|---|---|---|---|
| naid | Yes | The record's NAID. | |
| refresh | No | True re-reads from the Catalog instead of the cache, spending one call and replacing the cached copy. The cache never expires on its own, so use this when the answer may have changed since you last asked. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, so the safety profile is covered. The description adds real context beyond that: it frames the images as evidence, explains that the catalog description does not describe individual pages, and states the return shape (ordered entries carrying an object id). It omits pagination or size limits, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The load-bearing statement is front-loaded in one sentence, and the remaining sentences explain why and what to do next. The middle clause about catalog descriptions is slightly discursive but earns its place by justifying the call.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must sketch the return value, and it does: a page-ordered list, each item carrying an object id usable by get_extracted_text and get_transcriptions. It does not mention pagination, empty results, or failure modes, so it is good but not exhaustive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and both parameters (naid, refresh) are documented there, including the cache-replacement behavior of refresh. The description adds no direct detail about naid or refresh, so the baseline of 3 applies; its object-id talk concerns the return value rather than the inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence gives a specific verb (List), resource (a record's page images), and ordering guarantee (page order), plus the key field returned (object id). It is immediately distinguishable from siblings like get_record, get_extracted_text, and download_page_image.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description tells the agent when this tool matters ('read the images before citing anything to this record') and clarifies why the catalog description is insufficient. It also routes forward to download_page_image by explaining what to pass. It stops short of stating when to skip this tool or naming a directly competing alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_tagsARead-only
Read the citizen tags on a record.
On genealogical records tags are very often the names of the people who appear inside the file -- the widow, the children, the witnesses -- which the title does not carry. A tag is a stranger's reading of the document and is not evidence: follow it to the page image and cite what the image shows.
| Name | Required | Description | Default |
|---|---|---|---|
| naid | Yes | The record's NAID. | |
| refresh | No | True re-reads from the Catalog instead of the cache, spending one call and replacing the cached copy. The cache never expires on its own, so use this when the answer may have changed since you last asked. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only and open-world semantics, but the description adds genuinely useful behavioural context absent from structured fields: tags are user-contributed ('a stranger's reading'), not authoritative evidence, and should be verified against the page image. The refresh parameter's cache-replacement semantics are also documented in the schema, reinforcing cache behaviour disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action is front-loaded in the first sentence, with supporting context after it. The second paragraph is somewhat discursive for a tool description, but every sentence carries real interpretive value rather than filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description would ideally say what a tag result looks like (a list of strings, objects with contributors, etc.), and it does not. It is otherwise adequate: the safety profile is covered by annotations and the cache behaviour by the schema, but the return shape for a read tool remains undescribed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already fully documents both naid and refresh, including the cache behaviour and the cost of a refresh call. The description adds no additional parameter semantics beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read the citizen tags on a record') and clarifies what tags actually contain (names of people inside the file not carried by the title). This makes the operation's value concrete, though it never names the adjacent sibling search_tags to distinguish record-scoped reads from catalogue-wide searches.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the naid requirement and by the cautionary note about treating tags as pointers rather than evidence, which hints at when to trust them. However there is no explicit guidance on when to call this versus search_tags, get_record, or get_extracted_text, and no stated prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_transcriptionsARead-only
Read the citizen transcriptions of a record's pages.
Volunteers have transcribed handwritten pension files, service records and letters that OCR cannot touch, which makes this the fastest way into a document in copperplate. It is still a lead, not evidence: a transcription is one stranger's reading, unreviewed, and names are exactly where such a reading goes wrong. Check the page image before citing a name or date you found here, and cite the image.
Returns one entry per transcription with its text, its author and the page it belongs to.
| Name | Required | Description | Default |
|---|---|---|---|
| naid | Yes | The record's NAID. | |
| refresh | No | True re-reads from the Catalog instead of the cache, spending one call and replacing the cached copy. The cache never expires on its own, so use this when the answer may have changed since you last asked. | |
| max_chars | No | Characters of each transcription to return. 0 returns the whole thing. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover readOnly/openWorld, but the description adds real behavioral context the annotations cannot: the content is unreviewed volunteer output, names are the most error-prone element, and the correct workflow is to verify against the page image. It does not restate the mutation/caching behavior that the schema already documents.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence, followed by a caveat and a return summary. The prose is somewhat embellished ('copperplate', 'one stranger's reading'), which costs a point against strict economy, but each sentence carries usable information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description supplies the return structure, the trustworthiness caveat, and the verification workflow. For a simple three-parameter read tool this is fully sufficient to call and interpret the results correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so naid, refresh and max_chars are already fully documented in the schema. The description adds return-shape detail (one entry per transcription with text, author, page) but no additional parameter semantics, matching the baseline for full schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Read the citizen transcriptions of a record's pages') and scopes it to a single record via the required naid. The explanation that OCR cannot touch handwritten files clearly distinguishes it from the OCR-oriented sibling get_extracted_text.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives strong contextual guidance: this is the fastest route into handwritten documents, but it is 'still a lead, not evidence' and the image should be checked before citing names or dates. It lacks an explicit when-not-to-use or a named alternative (e.g. get_extracted_text for printed text), which keeps it shy of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_by_contribution_textARead-only
Search what people wrote on records, and get the records back.
This is the difference between searching a catalogue and searching the documents. A pension file titled only with the veteran's name will name his widow, his children and his witnesses in its transcribed text — none of which a title search reaches.
Unlike search_transcriptions and its siblings, which return the
contributions themselves, this filters the main index and returns full
record summaries. Use this when you want the record; use those when you
want to read what a particular volunteer wrote.
A transcription is one volunteer's reading and OCR is a machine's. Both are leads, not evidence — open the page image before citing anything.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page of results, 1-based. | |
| limit | No | Maximum records (1-100). | |
| tag_text | No | Words to find in citizen tags, which on genealogical records are usually the names of people appearing in them. | |
| comment_text | No | Words to find in researchers' comments on records. | |
| extracted_text | No | Words to find in OCR text, including text NARA's partners contributed. | |
| available_online | No | Restrict to digitised records only. | |
| transcription_text | No | Words to find in volunteers' transcriptions of the handwriting. Accepts AND, OR, NOT, wildcards and "exact phrases". |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint and openWorldHint, so the safety profile is already covered. The description adds genuinely useful behavioral context beyond that: it returns 'full record summaries' rather than contributions, and it warns that transcriptions and OCR are 'leads, not evidence' that should be verified against the page image. It stays silent on pagination behavior, but that is partially covered by the schema's page/limit fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence and the routing rule is crisp. The catalogue-versus-documents metaphor and the evidence caveat consume space, but both carry real decision value; the prose is slightly richer than strictly necessary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description supplies the return shape ('full record summaries'), the distinction from sibling tools, and a data-quality caveat about citations. Combined with annotations covering read-only/open-world behavior, an agent has everything needed to select and call this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so each of the seven parameters is already fully documented in the schema, including query syntax support. The description's conceptual framing of transcription vs OCR vs tags vs comments adds interpretive value but no per-parameter syntax or format detail beyond what the schema states. Baseline 3 applies when the schema carries the parameter burden.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first sentence states a specific verb and resource ('Search what people wrote on records, and get the records back') and the second makes the scope concrete: it filters the main index by contributed text rather than titles. It names and distinguishes itself from specific siblings (search_transcriptions and its siblings), so an agent can route correctly without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit routing rule: 'Use this when you want the record; use those when you want to read what a particular volunteer wrote.' It states both the when and the when-not, and names the competing siblings, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_commentsARead-only
Find records other researchers have commented on.
Useful for picking up where someone else stopped: a comment naming a surname often marks a file that a researcher has already read. Unverified by NARA, so it points at records rather than settling anything.
| Name | Required | Description | Default |
|---|---|---|---|
| naid | No | Restrict to comments on one record. | |
| page | No | Page of results, 1-based. | |
| limit | No | Maximum records to return (1-100). | |
| query | No | Words to find in comments. | |
| contributor | No | Restrict to one contributor's screen name. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and openWorldHint, so safety is covered. The description adds genuinely useful behavioral context beyond them: the comments are 'Unverified by NARA' and 'point at records rather than settling anything,' which tells the agent how much trust to place in results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the purpose and each sentence earning its place (use case, caveat). The second sentence is slightly narrative but carries real workflow value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only search with no output schema, the description adequately conveys that results are records surfaced via comments and that the data is unverified. It does not describe pagination or result shape, but with 100% schema coverage and annotations present this is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters (naid, page, limit, query, contributor) are already documented in the schema. The description only hints at query semantics via the 'surname' example and contributes no syntax or format detail, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Find records other researchers have commented on'), which is clear and distinct from the sibling get_comments (comments on a known record). It does not, however, name any sibling to route against, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete when-to-use scenario ('picking up where someone else stopped: a comment naming a surname often marks a file that a researcher has already read'). This is clear context, but there is no explicit exclusion or comparison to alternatives like get_comments or search_records.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_extracted_textARead-only
Find records whose extracted text mentions something.
This reaches text contributed by NARA's digitisation partners over the digital objects -- the searchable layer under the scans. It is machine output and is not evidence: it misses handwriting, it mangles names, and a hit means a page probably says this, not that it does. Read the page image before citing anything you find here.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page of results, 1-based. | |
| limit | No | Maximum records to return (1-100). | |
| query | Yes | Words to find in extracted text. Accepts AND, OR, NOT, wildcards and "exact phrases". |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnlyHint and openWorldHint, but the description adds substantial behavioral context the annotations cannot: this is machine output that misses handwriting and mangles names, and a hit means a page 'probably' says this rather than proving it. It sets a clear expectation about result reliability and mandates verifying against the page image – genuinely useful disclosure beyond the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence, and the subsequent caveat about reliability is well placed and each clause carries distinct information (misses handwriting, mangles names, probabilistic hits). It is a touch long but nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only search tool with no output schema and full schema coverage on the three params, the definition covers purpose and the critical reliability caveat an agent needs before citing results. It does not address result ordering or how pages/pagination interplay, but those are minor given the schema and annotations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so query/page/limit (including query syntax and the 1-100 limit range) are fully documented in the schema. The description adds no parameter-level meaning beyond that, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence gives a specific verb and resource: 'Find records whose extracted text mentions something.' The follow-up explains the nature of the corpus (OCR/searchable layer under scans contributed by digitisation partners), which helps position it. It does not, however, explicitly distinguish it from close siblings like search_transcriptions or search_by_contribution_text.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: it is for finding text mentions, with the caveat that hits are not evidence and the page image should be read before citing. There is no explicit when-to-use-vs-alternatives routing to the many sibling search tools (search_records, search_transcriptions, search_by_contribution_text), which is a real gap given how crowded this family is.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_recordsARead-only
Search the National Archives Catalog for records.
Start with title and a specific phrase; the Catalog holds tens of
millions of descriptions and a broad query will bury the useful hit. A
result is a lead: check the hierarchy and dates against what you already
know before reading the images.
Returns the total number of matches and a page of summaries, each with its
NAID, hierarchy, holding unit and image count. Use search_records_advanced
when you need dates, a record group, an M-number or digitised-only.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page of results, 1-based. The API pages rather than offsetting; beyond 10,000 results it needs cursor pagination. | |
| limit | No | Maximum records to return (1-100). | |
| query | No | Full-text search across the description. Broader and noisier than title; use it when a title search finds nothing. | |
| title | No | Words to match in the record title, e.g. 'Hall pension' or a person's name. Titles of case files usually carry the name. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint, openWorldHint), so the bar is lower, yet the description still adds value: it warns that results are leads needing verification, and describes the return shape (total matches plus page of summaries with NAID, hierarchy, holding unit, image count). It does not repeat the annotations but adds practical caution about trusting hits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then usage guidance, then return format and sibling routing across three short paragraphs. Every sentence carries distinct information with no padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by describing the return value (match count plus page of summaries with key fields). Combined with the 100%-covered input schema, the annotation safety profile, and explicit sibling routing, an agent has everything needed to call and interpret this tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds strategic meaning beyond the schema: it advises starting with `title` and a specific phrase and frames `query` as the fallback when a title search fails. That semantic framing of when to prefer one parameter over the other is not present in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Search the National Archives Catalog for records') and explicitly differentiates itself from the sibling search_records_advanced by naming the conditions that select the other tool. An agent can tell the two search tools apart without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete when-to-use guidance ('Start with title and a specific phrase'), warns that a broad query buries hits, and names the alternative (search_records_advanced) with the exact triggers for switching (dates, record group, M-number, digitised-only). This is explicit routing rather than implied context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_records_advancedARead-only
Search the Catalog with the filters that narrow a common name.
Every parameter is optional but at least one is required. The filters that
earn their keep for research are the date range, microform_publication
for an M-number you already cite, record_group_number or ancestor_naid
to stay inside one body of records, and available_online when you intend
to read pages rather than order copies.
Results are the same summaries search_records returns. Past 10,000 hits
use the next_search_after cursor rather than page; a common surname
passes that boundary easily.
| Name | Required | Description | Default |
|---|---|---|---|
| page | No | Page of results, 1-based. Ignored when search_after is given, and unusable past 10,000 results. | |
| exact | No | Match title, local_identifier and microform_publication in full and exactly, instead of by words within them. Use it when a words match returns too much, or you hold the complete title. | |
| limit | No | Maximum records to return (1-100). | |
| query | No | Full-text search across the whole description. | |
| title | No | Words to match in the record title. | |
| creators | No | The agency or person who created the records, matched against the creator headings. | |
| end_date | No | Latest date to include, in the same format as start_date. | |
| exact_date | No | A single date, YYYY-MM-DD. Cannot be combined with start_date or end_date. | |
| start_date | No | Earliest date to include, as YYYY, YYYY-MM or YYYY-MM-DD. Use the same precision as end_date. A surname search is usually only workable once it is bounded to a lifetime. | |
| tags_exist | No | True for records carrying citizen tags. | |
| data_source | No | 'description' for archival descriptions, 'authority' for authority records (people, organisations, topics). This narrows a search and cannot be one on its own. | |
| search_after | No | Cursor for paging past 10,000 results. You MUST pass '*' for the first page, then the 'next_search_after' value from each response. Starting from an ordinary search does not work: without '*' the results are relevance-sorted and their cursor is not resumable. Cannot be combined with page. | |
| ancestor_naid | No | NAID of an ancestor node: returns only records below it in the hierarchy. Use it to search inside one series. | |
| person_or_org | No | A person or organisation named in the description, either as its subject or in a role such as creator. Distinct from `creators`, which is the record's creating body only. | |
| recurring_day | No | Day as DD. Normally used with recurring_month. | |
| comments_exist | No | True for records carrying researcher comments. | |
| congress_number | No | Records of one numbered Congress, e.g. 55 for 1897-99. Private relief bills, petitions and claims naming individuals sit in the records of Congress. | |
| control_numbers | No | Any identifier NARA attaches to a record: accession number, local identifier, microfilm publication, NAID, transfer number or variant control number. For a citation whose kind you cannot name. | |
| recurring_month | No | Month as MM. With recurring_day, finds records dated to that day in any year -- a birthday across every census. | |
| reference_units | No | Name of the archive holding the paper, e.g. 'National Archives at St. Louis'. Comma-separate several. | |
| available_online | No | True for digitised records only -- those whose pages you can read now rather than order from a reading room. | |
| local_identifier | No | The archives' own identifier for the record, as printed in finding aids. | |
| type_of_materials | No | Material type, e.g. 'Textual Records', 'Photographs and other Graphic Materials', 'Maps and Charts', 'Moving Images'. | |
| contributions_exist | No | True for records carrying any contribution at all: a transcription, tag or comment. Someone has already worked on them. | |
| record_group_number | No | Record group number, e.g. '15' for Veterans Affairs. Scopes the search to one agency's records. | |
| geographic_reference | No | Place the records are about, matched against geographic subject headings, e.g. 'Franklin County (Pa.)'. | |
| level_of_description | No | One of recordGroup, collection, series, fileUnit, item. A case file is usually a fileUnit; a single page is an item. | |
| transcriptions_exist | No | True for records someone has transcribed; False for ones nobody has. Transcribed records are searchable by their text. | |
| collection_identifier | No | Collection identifier, the Presidential-library equivalent of a record group. | |
| microform_publication | No | Microfilm publication number, e.g. 'M804' for Revolutionary War pension applications or 'T624' for the 1910 census. This is the citation genealogists actually carry. | |
| include_extracted_text | No | Fold each hit's OCR text into the response, saving a call per hit: each record gains an extracted_text list with one entry per page that carries text, NARA's own OCR or a partner's, capped at 2000 characters each. It makes the response much larger, so use it on a narrowed search rather than a broad one; get_extracted_text returns whole pages. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and openWorldHint, so safety is covered. The description adds real behavioral context beyond that: results are the same summaries `search_records` returns, and a hard 10,000-hit boundary forces the `next_search_after` cursor instead of `page`. It stops short of describing ordering or response size tradeoffs beyond the paging note.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short paragraphs, front-loaded with purpose, then filter selection, then paging. No filler sentences; each one carries a distinct instruction.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 31 parameters, no output schema and no required fields, the description does well to explain the at-least-one-rule, result shape and cursor mechanics. It could go further on what the returned record summaries contain, but for a tool whose schema is fully documented this is close to sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description still earns above that by adding strategic meaning the schema lacks — e.g. using `available_online` only when the user intends to read pages rather than order copies, and bounding a surname search by date. It adds selection rationale rather than restating field definitions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (search) and resource (the Catalog) plus the distinguishing scope — the filter-heavy variant that narrows a common name. However, it never explicitly positions itself against the sibling `search_records`; it only mentions that sibling to describe the return shape, not to explain which tool to pick. A clear purpose, but sibling differentiation is left implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete research-oriented guidance on which filters matter (date range, microform_publication, record_group_number/ancestor_naid, available_online) and the rule that at least one parameter is required. It does not, however, say when NOT to use this tool or explicitly route the agent to `search_records` for the simpler case.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_tagsARead-only
Find records that carry a given citizen tag.
Tags are short, so an exact tag_is on a surname is often sharper than a
title search: a volunteer who read the file tagged the people in it. What
comes back is what a stranger thought the document said -- a lead to the
page image, and not evidence.
| Name | Required | Description | Default |
|---|---|---|---|
| naid | No | Restrict to tags on one record. | |
| page | No | Page of results, 1-based. | |
| limit | No | Maximum records to return (1-100). | |
| query | No | Words to find in citizen tags. | |
| tag_is | No | Match one exact tag rather than words within tags. | |
| contributor | No | Restrict to one contributor's screen name. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=true, so safety is covered. The description adds genuinely valuable interpretive context beyond that: results are 'what a stranger thought the document said -- a lead to the page image, and not evidence,' warning the agent not to treat tags as authoritative. Pagination/limits are left to the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence, with supporting rationale following. The prose is slightly discursive but each sentence carries useful guidance; nothing is pure filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by characterizing the return ('a lead to the page image, and not evidence'). Combined with read-only annotations and full schema coverage, an agent has enough to invoke and interpret correctly, though result shape and pagination are only loosely implied.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description goes further by explaining the tradeoff between tag_is and query (exact tag vs words within tags) and the rationale that tags are short, adding real meaning beyond the schema's terse field descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Find records that carry a given citizen tag.' An agent can distinguish this from record/text searches. It does not explicitly name the closest sibling (get_tags, which likely returns tags on a record rather than searching across them), so it falls just short of full sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Offers a useful heuristic ('an exact tag_is on a surname is often sharper than a title search'), which implicitly contrasts with title/record search and guides parameter choice. However, it never explicitly says when to prefer search_tags over search_records or get_tags, so usage is only implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_transcriptionsARead-only
Find records whose citizen transcriptions mention something.
This searches the documents rather than the catalogue, which is the difference between finding a pension file titled with the veteran's name and finding the file that names his widow, his children and the neighbours who swore to the marriage. Only transcribed records are reachable this way, so silence here means nobody has transcribed it, not that it does not exist.
Returns record summaries. A transcription is a stranger's reading and is not evidence: open the matching page images to see what was written.
| Name | Required | Description | Default |
|---|---|---|---|
| naid | No | Restrict to transcriptions of one record. | |
| page | No | Page of results, 1-based. | |
| limit | No | Maximum records to return (1-100). | |
| query | No | Words to find in transcribed text. Accepts AND, OR, NOT, wildcards and "exact phrases". | |
| contributor | No | Restrict to one contributor's screen name. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/openWorld, so the bar is lower, yet the description adds real behavioral context: only transcribed records are reachable, absence of results reflects transcription coverage rather than existence, returns are summaries not page images, and transcriptions are fallible readings requiring verification against page images.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose, then layers coverage caveat and reliability caveat. Slightly prose-heavy, but each sentence carries distinct information that changes how an agent uses the results; nothing is redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description supplies the return shape ('record summaries') plus the two facts an agent most needs to interpret results correctly: partial transcription coverage and the unreliability of transcribed text. Nothing essential is missing for correct invocation and interpretation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all five parameters (naid, page, limit, query, contributor) are already documented in the schema, including query operator syntax. The description adds no parameter-level guidance beyond 'returns record summaries', so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb and resource ('Find records whose citizen transcriptions mention something') and immediately distinguishes itself from catalogue search: 'This searches the documents rather than the catalogue.' An agent can tell this apart from search_records without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives strong when-to-use framing (documents vs catalogue, transcribed-only coverage, silence ≠ nonexistence) with a concrete example of the difference in results. It stops short of naming the obvious sibling alternative, search_extracted_text, which an agent must choose between for document-text search.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
18 tool updates
v0.1.0- First observed
api_budget - First observed
browse_children - First observed
download_page_image - First observed
get_comments - First observed
get_extracted_text - First observed
get_online_availability - First observed
get_partner_digital_objects - First observed
get_record - First observed
get_record_images - First observed
get_tags - First observed
get_transcriptions - First observed
search_by_contribution_text - First observed
search_comments - First observed
search_extracted_text - First observed
search_records - First observed
search_records_advanced - First observed
search_tags - First observed
search_transcriptions
TDQS
Scored across 18 tools
There is a large cluster of search tools (search_records, search_records_advanced, search_transcriptions, search_tags, search_extracted_text, search_comments, search_by_contribution_text) plus mirroring getters, which could cause misselection. However, the descriptions carefully draw the line — notably search_by_contribution_text returning records vs. the contribution-returning siblings — so an attentive agent can tell them apart.
Nearly everything follows a clean snake_case verb_noun pattern (get_record, search_records, browse_children, download_page_image). The lone outlier is api_budget, a bare noun that breaks the convention, but it is a minor deviation.
18 tools is on the heavier side but each maps to a distinct resource or surface (records, images, OCR, transcriptions, tags, comments, partner objects, hierarchy, budget), so most earn their place. It sits just past the comfortable 3-15 range rather than being bloated.
The surface covers search, full-record read, page-image listing/download, machine and human text, citizen tags/comments, partner objects, hierarchy descent, and budget tracking. Only minor gaps exist, such as no explicit ancestor-walk or citation-export tool, though ancestors are folded into get_record.
Maintenance
Related MCP Connectors
US National Archives Catalog (NARA): search the federal government's permanent records…
Read-only OpenHeritage search for genealogy and cultural heritage records.
Search and trace US federal rules across the Federal Register, eCFR, and Regulations.gov.
Search experimental video history, read public records, follow sources and export citations.
Related MCP Servers
- AlicenseAqualityAmaintenanceEnables users to search and access digital collections from the Swedish National Archives (Riksarkivet) through multiple APIs. Supports searching records by keywords, exploring collections, and downloading historical images and documents.224Apache 2.0
- AlicenseAqualityCmaintenanceRead-only Model Context Protocol server for the US National Archives Catalog API, enabling search and retrieval of archival records, child records, extracted text, comments, and tags.71MIT
- FlicenseNot gradedqualityBmaintenanceEnables querying genealogical records, archive statistics, historical weather, and full-text transcriptions from Open Archives via natural language.8 npm4-
- FlicenseAqualityCmaintenanceEnables users to search catalog metadata, read captured pages, verify quotations, and detect changes in accepted catalog captures for New York City's September 11th Document Portal using local, read-only snapshots.10-