Skip to main content
Glama
ianderso

nara-catalog-mcp

by ianderso

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault
NARA_API_KEYYesYour Catalog API key. Required.
NARA_TIMEOUTNoHTTP timeout in seconds. Default 60.60
NARA_CACHE_DIRNoResponse cache directory. Default `~/.cache/nara-catalog-mcp`.~/.cache/nara-catalog-mcp
NARA_MONTHLY_CALL_BUDGETNoCalls per month the key allows, reported by `api_budget`. Default 10000.10000

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": false
}
prompts
{
  "listChanged": false
}
resources
{
  "subscribe": false,
  "listChanged": false
}
experimental
{}

Tools

Functions exposed to the LLM to take actions

NameDescription
search_recordsA

Search the National Archives Catalog for records.

Start with title and a specific phrase; the Catalog holds tens of millions of descriptions and a broad query will bury the useful hit. A result is a lead: check the hierarchy and dates against what you already know before reading the images.

Returns the total number of matches and a page of summaries, each with its NAID, hierarchy, holding unit and image count. Use search_records_advanced when you need dates, a record group, an M-number or digitised-only.

search_records_advancedA

Search the Catalog with the filters that narrow a common name.

Every parameter is optional but at least one is required. The filters that earn their keep for research are the date range, microform_publication for an M-number you already cite, record_group_number or ancestor_naid to stay inside one body of records, and available_online when you intend to read pages rather than order copies.

Results are the same summaries search_records returns. Past 10,000 hits use the next_search_after cursor rather than page; a common surname passes that boundary easily.

get_recordA

Read one Catalog record in full, by NAID.

Returns the scope and content note, the full hierarchy, the holding reference units, and every page-image URL. Use this once a search has given you a NAID worth pursuing.

get_record_imagesA

List a record's page images in page order, each with its object id.

These are the evidence. A catalog description summarises a file; it does not tell you what any individual page says, so read the images before citing anything to this record. The object id is what get_extracted_text and get_transcriptions entries point at; pass it, or the page number, to download_page_image.

get_extracted_textA

Read the OCR text NARA machine-extracted from a record's page images.

This is a lead, not evidence. OCR was run over scans of handwriting, carbon copies and microfilm: it drops handwritten pages entirely, and it turns one surname into another silently. Use it to find which page matters, then open that page image with get_record_images and cite what you saw there -- never cite the OCR text itself.

Returns one entry per digital object, in page order, with the text truncated to max_chars.

get_transcriptionsA

Read the citizen transcriptions of a record's pages.

Volunteers have transcribed handwritten pension files, service records and letters that OCR cannot touch, which makes this the fastest way into a document in copperplate. It is still a lead, not evidence: a transcription is one stranger's reading, unreviewed, and names are exactly where such a reading goes wrong. Check the page image before citing a name or date you found here, and cite the image.

Returns one entry per transcription with its text, its author and the page it belongs to.

get_tagsA

Read the citizen tags on a record.

On genealogical records tags are very often the names of the people who appear inside the file -- the widow, the children, the witnesses -- which the title does not carry. A tag is a stranger's reading of the document and is not evidence: follow it to the page image and cite what the image shows.

get_commentsA

Read other researchers' comments on a record.

Comments often say what a file actually contains, where a related file sits, or that the description is wrong. None of it is verified by NARA: it is correspondence between researchers, useful for finding the next record and never a source in itself.

search_transcriptionsA

Find records whose citizen transcriptions mention something.

This searches the documents rather than the catalogue, which is the difference between finding a pension file titled with the veteran's name and finding the file that names his widow, his children and the neighbours who swore to the marriage. Only transcribed records are reachable this way, so silence here means nobody has transcribed it, not that it does not exist.

Returns record summaries. A transcription is a stranger's reading and is not evidence: open the matching page images to see what was written.

search_tagsA

Find records that carry a given citizen tag.

Tags are short, so an exact tag_is on a surname is often sharper than a title search: a volunteer who read the file tagged the people in it. What comes back is what a stranger thought the document said -- a lead to the page image, and not evidence.

search_extracted_textA

Find records whose extracted text mentions something.

This reaches text contributed by NARA's digitisation partners over the digital objects -- the searchable layer under the scans. It is machine output and is not evidence: it misses handwriting, it mangles names, and a hit means a page probably says this, not that it does. Read the page image before citing anything you find here.

search_commentsA

Find records other researchers have commented on.

Useful for picking up where someone else stopped: a comment naming a surname often marks a file that a researcher has already read. Unverified by NARA, so it points at records rather than settling anything.

browse_childrenA

List a record's immediate children, one level down the hierarchy.

The Catalog nests record group, then series, then file unit, then item. Searching finds a node; this walks down from it, which is how you get from a series you trust to the file unit for one person, and how you find out what else sits alongside a file you already have. A record's own ancestors come back from get_record.

get_online_availabilityA

List where else this record is available online.

NARA records what has been digitised and published elsewhere, including by commercial partners. Use it to resolve a hint from a subscription site back to the archival original: the NAID and the reference unit are what make a citation refindable, and a partner's index entry is not a substitute for the page.

get_partner_digital_objectsA

List digital object ids a partner has matched to this record.

NARA indexes metadata supplied by commercial partners — Ancestry among them — against its own records. A hit tells you the partner holds imagery for this NAID, which is worth knowing when NARA's own pages are not online.

An empty list is the common answer and is not an error. It means no partner metadata has been matched, not that no partner holds the record.

The ids are pointers into the partner's index, not a citation. Cite the archival record by NAID and reference unit.

search_by_contribution_textA

Search what people wrote on records, and get the records back.

This is the difference between searching a catalogue and searching the documents. A pension file titled only with the veteran's name will name his widow, his children and his witnesses in its transcribed text — none of which a title search reaches.

Unlike search_transcriptions and its siblings, which return the contributions themselves, this filters the main index and returns full record summaries. Use this when you want the record; use those when you want to read what a particular volunteer wrote.

A transcription is one volunteer's reading and OCR is a machine's. Both are leads, not evidence — open the page image before citing anything.

download_page_imageA

Download one page of a record so it can actually be read.

This is the step that turns a catalogue hit into evidence. The image is written to disk rather than returned inline — a page scan runs to several megabytes.

NARA's media URLs are open and need no key, so this costs nothing against your API allowance. One catalogue call resolves the page list; the download itself is not an API call.

A pension file can run to sixty pages and the page you need is rarely the first. Use get_extracted_text or get_transcriptions to find which page carries the fact, then fetch it by the object_id they name.

api_budgetA

Report the key's Catalog API spend, this session and this month.

A key is capped per month and a long sweep can exhaust it. Live calls are kept in a ledger beside the cache, so the month's figure survives restarts. Cached repeats cost nothing and are not counted.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

A4/5.0

Scored across 18 tools

Disambiguation4/5

There is a large cluster of search tools (search_records, search_records_advanced, search_transcriptions, search_tags, search_extracted_text, search_comments, search_by_contribution_text) plus mirroring getters, which could cause misselection. However, the descriptions carefully draw the line — notably search_by_contribution_text returning records vs. the contribution-returning siblings — so an attentive agent can tell them apart.

Naming Consistency4/5

Nearly everything follows a clean snake_case verb_noun pattern (get_record, search_records, browse_children, download_page_image). The lone outlier is api_budget, a bare noun that breaks the convention, but it is a minor deviation.

Tool Count4/5

18 tools is on the heavier side but each maps to a distinct resource or surface (records, images, OCR, transcriptions, tags, comments, partner objects, hierarchy, budget), so most earn their place. It sits just past the comfortable 3-15 range rather than being bloated.

Completeness4/5

The surface covers search, full-record read, page-image listing/download, machine and human text, citizen tags/comments, partner objects, hierarchy descent, and budget tracking. Only minor gaps exist, such as no explicit ancestor-walk or citation-export tool, though ancestors are folded into get_record.

Maintenance

ActivityNo data
ResponsivenessNo issues