Skip to main content
Glama
magicianmarty

heritage-research-mcp

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault
DPLA_API_KEYNoDPLA API key (or place it on one line in ~/.config/heritage-research-mcp/keys/dpla, mode 600). DPLA is skipped until the key is set.
NARA_API_KEYNoUS National Archives Catalog API key (or place it on one line in ~/.config/heritage-research-mcp/keys/nara, mode 600). NARA is skipped until the key is set.
NARA_API_VERSIONNov3 to use NARA's newer search.v2
NARA_MONTHLY_LIMITNoStop before NARA's monthly quota.10000
SMITHSONIAN_API_KEYNoSmithsonian Open Access API key (or place it on one line in ~/.config/heritage-research-mcp/keys/smithsonian, mode 600). Smithsonian is skipped until the key is set.
HERITAGE_MCP_CONTACTNoAdded to the User-Agent, so providers can reach you.
HERITAGE_MCP_CACHE_DIRNoWhere download_media writes.~/.cache/heritage-research-mcp
HERITAGE_MCP_MAX_DOWNLOAD_MBNoRefuse larger downloads.250
HERITAGE_MCP_DISABLE_DOWNLOADSNoRemove download_media (hosted use). Set to 1 to disable downloads, since a hosted agent cannot read files the server downloads.off

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": false
}
prompts
{
  "listChanged": false
}
resources
{
  "subscribe": false,
  "listChanged": false
}

Tools

Functions exposed to the LLM to take actions

NameDescription
list_sourcesA

List the archives this server can search, whether each is ready to use, and what it is good for.

Sources that need a free API key show configured: false until a key is set; the Internet Archive and Wikimedia Commons need none.

usage_reportA

Requests made this month and this session per source, and the provider rate limit last seen.

searchA

Search several archives at once and return normalised records grouped by source.

Every record carries rights.reuse (free, attribution, share_alike, non_commercial, no_derivatives, restricted or unknown). This is a filtering aid, not legal advice: unknown means the holder stated nothing.

Args: query: What to look for. Plain words work everywhere. sources: Limit to some of internet_archive, commons, dpla, nara, smithsonian. Default: every source that is ready to use (see list_sources). limit: Records per source (1 to 25). date_from: Earliest year. Applied by the Internet Archive, DPLA and NARA; ignored by the others. date_to: Latest year, applied the same way. fulltext: Also search the text inside Internet Archive books (an experimental endpoint). Returned under fulltext; this is the way to find a name or place inside a memoir or official report. kind: Only this kind of material: text (books, reports, manuscripts), image (photographs, prints), map, audio or video. Each archive is asked in its own vocabulary, and kind_applied in each result says how. Maps are the least uniform: Commons matches on the word "map" in the title and the Internet Archive on its map subjects and collections, so expect some misses and some noise. Every record also carries a best-effort kind.

get_recordA

Fetch one record in the normalised shape, including its media files and rights.

Args: source: internet_archive, commons, dpla, nara or smithsonian. record_id: The id from a search result (an Internet Archive identifier, a Commons "File:" title, a DPLA id, a NARA naId, or a Smithsonian id).

download_mediaA

Download one original file from a record into the local cache and write a provenance sidecar.

The sidecar (<file>.provenance.json) records the source, URL, retrieval time, SHA-256 and the rights the holder stated, so the file can be traced later. Only https URLs on public addresses are fetched. If rights.reuse is not free the result carries a rights_warning: treat the file as reference only.

Args: source: internet_archive, commons, dpla, nara or smithsonian. record_id: The record's id, as returned by search. media_index: Which entry of the record's media list to fetch (default the first). media_kind: Pick the first media of this kind instead (pdf, text, image, audio, video, archive). overwrite: Fetch again even if the file is already cached. max_mb: Refuse files larger than this (default 250, or HERITAGE_MCP_MAX_DOWNLOAD_MB).

ia_searchA

Search Internet Archive item metadata (title, author, subject, description). Not the text inside books.

Use ia_fulltext_search to find a phrase inside scanned books. The query accepts Lucene syntax, for example creator:(mosby) AND subject:"Virginia".

Args: query: The search. mediatype: texts, image, audio, movies, software, data, etc. year_from: Earliest publication year. year_to: Latest publication year. collection: Restrict to a collection identifier, e.g. americana. rows: Results per page (1 to 100). page: Page number from 1. sort: For example "downloads desc" or "date asc". kind: text, image, map, audio or video, translated into Internet Archive terms (maps are matched on the "maps" subject and the map collections, so results are noisy).

ia_fulltext_searchA

Search the OCR text inside Internet Archive books and reports, with snippets and page numbers.

This is the way to find a name, place or phrase inside a memoir, regimental history or official report. Matches are wrapped in ** in the snippets. It uses an experimental endpoint that may change.

Args: query: Words or a "quoted phrase". hits: Results to return (1 to 50). page: Page number from 1.

ia_get_itemB

Get one item's metadata, rights and downloadable files (PDF, OCR text, page images).

ia_read_textA

Read a slice of an item's OCR text by character offset (up to 20,000 characters at a time).

Use ia_grep_text first to find the offset you want.

ia_grep_textA

Find every occurrence of a word or phrase in an item's OCR text, with surrounding context and offsets.

The text is fetched once and held in memory, so repeated searches of the same book are fast.

Args: identifier: The Internet Archive identifier. pattern: Text to find. A plain phrase matches across the line breaks, repeated spaces and split hyphens that OCR leaves in scanned books ("twenty-five prisoners" finds "twenty-five prisoners" and "twenty- five prisoners"). A regular expression only if regex is true. regex: Treat the pattern as a regular expression. ignore_case: Ignore capitalisation. context: Characters of context on each side (0 to 1000). max_matches: Most matches to return (1 to 100); the total is always reported.

commons_searchB

Search Wikimedia Commons files. Each record includes the licence read from the file's metadata.

Args: query: Search words. limit: Results to return (1 to 50). filetype: bitmap, drawing, audio, video, office or multimedia. kind: text, image, map, audio or video. Maps are found by the word "map" in the file title.

commons_file_infoA

Get one Commons file's licence, author, description, categories and original size.

Args: title: The file name, with or without the "File:" prefix.

commons_category_membersA

List the files in a Commons category (for example "Mosby's Rangers").

Args: category: The category name, with or without the "Category:" prefix. limit: Results to return (1 to 50). filetype: Optionally keep only bitmap, drawing, audio or video files.

dpla_searchA

Search the Digital Public Library of America for items held by US libraries, archives and museums.

Needs DPLA_API_KEY. Give a search term, any filter, or both. DPLA indexes descriptions, not the item itself: follow landing_url to the holding institution, and read rights before reusing anything.

Args: q: Free-text search across the record. title: Words in the title. creator: Creator or author. subject: Subject heading. place_state: US state name, e.g. Virginia. date_after: Earliest date (YYYY or YYYY-MM-DD). date_before: Latest date. type: image, text, sound, moving image, physical object, etc. provider: DPLA hub name. data_provider: Contributing institution name. page: Page number from 1. page_size: Results per page (1 to 100). sort_by: A DPLA field such as sourceResource.date.begin. kind: text, image, map, audio or video. Maps are images with the subject "Maps".

dpla_get_itemB

Get one DPLA item by its id. Needs DPLA_API_KEY.

nara_searchA

Search the US National Archives Catalog. Needs NARA_API_KEY (10,000 queries a month by default).

One page per call, at most 100 records: the API's terms forbid scraping or bulk download. The remaining monthly budget is visible in usage_report. Results include the attribution NARA requires.

Args: q: Search words; supports AND, OR, NOT, wildcards (*) and "exact phrases". title: Words in the title. start_date: Earliest date (YYYY, YYYY-MM or YYYY-MM-DD). end_date: Latest date. available_online: Only records with digitised objects. type_of_materials: e.g. Photographs and other Graphic Materials, Textual Records, Maps. level: series, fileUnit, item, recordGroup, etc. record_group: Record group number, e.g. 109 (Confederate records). ancestor_na_id: Only records beneath this naId. geographic: Geographic subject heading. creators: Creator heading. include_extracted_text: Include OCR text in the results where NARA has it. limit: Results per page (1 to 100). page: Page number from 1. kind: text, image, map, audio or video, mapped to NARA's type of materials (unverified until a key has been used against the live service).

nara_get_recordA

Get one Catalog record by its numeric naId, with digital objects and ancestry. Needs NARA_API_KEY.

nara_childrenB

List the immediate children of a series or file unit (one level down). Needs NARA_API_KEY.

nara_extracted_textB

Get OCR text extracted from a record's digital objects. Needs NARA_API_KEY.

The response is passed through from NARA with long strings shortened to 4,000 characters.

si_searchA

Search Smithsonian Open Access records. Needs SMITHSONIAN_API_KEY.

Records whose media are all marked CC0 come back with rights.reuse "free"; otherwise "Usage conditions apply" is surfaced as restricted.

Args: q: Search words; supports AND/OR and fielded terms such as topic:Cavalry (see si_terms). rows: Results to return (1 to 100). start: Offset of the first row. sort: relevancy (default), newest or updated. type: EDAN record type, e.g. edanmdm. row_group: objects or archives. category: Search within art_design, history_culture or science_technology instead. kind: text, image, map, audio or video, added to the query as object_type or online_media_type.

si_get_contentA

Get one Smithsonian record by id (for example edanmdm-nmaahc_2012.36.4ab). Needs SMITHSONIAN_API_KEY.

si_termsB

List the vocabulary for a search field: culture, data_source, date, online_media_type, place, topic or unit_code. Useful for building fielded queries. Needs SMITHSONIAN_API_KEY.

si_statsA

Counts of CC0 objects and media in the collection. Needs SMITHSONIAN_API_KEY.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

A3.5/5.0

Scored across 23 tools

Disambiguation3/5

The general `search` and `get_record` tools overlap with the five per-source search and get tools (si_search, nara_search, ia_search, commons_search, dpla_search, si_get_content, nara_get_record, ia_get_item, commons_file_info, dpla_get_item). Descriptions help by distinguishing normalized output from source-specific details, but an agent could still misselect for simple queries.

Naming Consistency4/5

Most tools follow a predictable snake_case pattern with source prefixes (nara_, si_, ia_, commons_, dpla_) for source-specific operations and unprefixed names for general operations. Minor deviations are noun-only names like `si_terms`, `si_stats`, and `usage_report`, but overall the scheme is clear and consistent.

Tool Count3/5

23 tools is on the heavy side for a single server, even one covering five archives. While many tools serve distinct purposes (e.g., `ia_grep_text`, `nara_children`, `si_terms`), the overlapping general and source-specific search/get tools add redundancy.

Completeness4/5

The set covers discovery, retrieval, fulltext search, media download, rights, provenance, and usage reporting across five archives. Minor gaps remain, such as cross-source fulltext search (only Internet Archive supports it) and OCR retrieval for some sources, but core research workflows are well supported.

Maintenance

ActivityMaintained
ResponsivenessNo issues