heritage-research-mcp
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| DPLA_API_KEY | No | DPLA API key (or place it on one line in ~/.config/heritage-research-mcp/keys/dpla, mode 600). DPLA is skipped until the key is set. | |
| NARA_API_KEY | No | US National Archives Catalog API key (or place it on one line in ~/.config/heritage-research-mcp/keys/nara, mode 600). NARA is skipped until the key is set. | |
| NARA_API_VERSION | No | v3 to use NARA's newer search. | v2 |
| NARA_MONTHLY_LIMIT | No | Stop before NARA's monthly quota. | 10000 |
| SMITHSONIAN_API_KEY | No | Smithsonian Open Access API key (or place it on one line in ~/.config/heritage-research-mcp/keys/smithsonian, mode 600). Smithsonian is skipped until the key is set. | |
| HERITAGE_MCP_CONTACT | No | Added to the User-Agent, so providers can reach you. | |
| HERITAGE_MCP_CACHE_DIR | No | Where download_media writes. | ~/.cache/heritage-research-mcp |
| HERITAGE_MCP_MAX_DOWNLOAD_MB | No | Refuse larger downloads. | 250 |
| HERITAGE_MCP_DISABLE_DOWNLOADS | No | Remove download_media (hosted use). Set to 1 to disable downloads, since a hosted agent cannot read files the server downloads. | off |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| list_sourcesA | List the archives this server can search, whether each is ready to use, and what it is good for. Sources that need a free API key show |
| usage_reportA | Requests made this month and this session per source, and the provider rate limit last seen. |
| searchA | Search several archives at once and return normalised records grouped by source. Every record carries Args:
query: What to look for. Plain words work everywhere.
sources: Limit to some of internet_archive, commons, dpla, nara, smithsonian. Default: every source
that is ready to use (see list_sources).
limit: Records per source (1 to 25).
date_from: Earliest year. Applied by the Internet Archive, DPLA and NARA; ignored by the others.
date_to: Latest year, applied the same way.
fulltext: Also search the text inside Internet Archive books (an experimental endpoint). Returned
under |
| get_recordA | Fetch one record in the normalised shape, including its media files and rights. Args: source: internet_archive, commons, dpla, nara or smithsonian. record_id: The id from a search result (an Internet Archive identifier, a Commons "File:" title, a DPLA id, a NARA naId, or a Smithsonian id). |
| download_mediaA | Download one original file from a record into the local cache and write a provenance sidecar. The sidecar ( Args:
source: internet_archive, commons, dpla, nara or smithsonian.
record_id: The record's id, as returned by search.
media_index: Which entry of the record's |
| ia_searchA | Search Internet Archive item metadata (title, author, subject, description). Not the text inside books. Use ia_fulltext_search to find a phrase inside scanned books. The query accepts Lucene syntax, for
example Args: query: The search. mediatype: texts, image, audio, movies, software, data, etc. year_from: Earliest publication year. year_to: Latest publication year. collection: Restrict to a collection identifier, e.g. americana. rows: Results per page (1 to 100). page: Page number from 1. sort: For example "downloads desc" or "date asc". kind: text, image, map, audio or video, translated into Internet Archive terms (maps are matched on the "maps" subject and the map collections, so results are noisy). |
| ia_fulltext_searchA | Search the OCR text inside Internet Archive books and reports, with snippets and page numbers. This is the way to find a name, place or phrase inside a memoir, regimental history or official report. Matches are wrapped in ** in the snippets. It uses an experimental endpoint that may change. Args: query: Words or a "quoted phrase". hits: Results to return (1 to 50). page: Page number from 1. |
| ia_get_itemB | Get one item's metadata, rights and downloadable files (PDF, OCR text, page images). |
| ia_read_textA | Read a slice of an item's OCR text by character offset (up to 20,000 characters at a time). Use ia_grep_text first to find the offset you want. |
| ia_grep_textA | Find every occurrence of a word or phrase in an item's OCR text, with surrounding context and offsets. The text is fetched once and held in memory, so repeated searches of the same book are fast. Args:
identifier: The Internet Archive identifier.
pattern: Text to find. A plain phrase matches across the line breaks, repeated spaces and split
hyphens that OCR leaves in scanned books ("twenty-five prisoners" finds "twenty-five prisoners" and
"twenty- five prisoners"). A regular expression only if |
| commons_searchB | Search Wikimedia Commons files. Each record includes the licence read from the file's metadata. Args: query: Search words. limit: Results to return (1 to 50). filetype: bitmap, drawing, audio, video, office or multimedia. kind: text, image, map, audio or video. Maps are found by the word "map" in the file title. |
| commons_file_infoA | Get one Commons file's licence, author, description, categories and original size. Args: title: The file name, with or without the "File:" prefix. |
| commons_category_membersA | List the files in a Commons category (for example "Mosby's Rangers"). Args: category: The category name, with or without the "Category:" prefix. limit: Results to return (1 to 50). filetype: Optionally keep only bitmap, drawing, audio or video files. |
| dpla_searchA | Search the Digital Public Library of America for items held by US libraries, archives and museums. Needs DPLA_API_KEY. Give a search term, any filter, or both. DPLA indexes descriptions, not the item
itself: follow Args: q: Free-text search across the record. title: Words in the title. creator: Creator or author. subject: Subject heading. place_state: US state name, e.g. Virginia. date_after: Earliest date (YYYY or YYYY-MM-DD). date_before: Latest date. type: image, text, sound, moving image, physical object, etc. provider: DPLA hub name. data_provider: Contributing institution name. page: Page number from 1. page_size: Results per page (1 to 100). sort_by: A DPLA field such as sourceResource.date.begin. kind: text, image, map, audio or video. Maps are images with the subject "Maps". |
| dpla_get_itemB | Get one DPLA item by its id. Needs DPLA_API_KEY. |
| nara_searchA | Search the US National Archives Catalog. Needs NARA_API_KEY (10,000 queries a month by default). One page per call, at most 100 records: the API's terms forbid scraping or bulk download. The remaining monthly budget is visible in usage_report. Results include the attribution NARA requires. Args: q: Search words; supports AND, OR, NOT, wildcards (*) and "exact phrases". title: Words in the title. start_date: Earliest date (YYYY, YYYY-MM or YYYY-MM-DD). end_date: Latest date. available_online: Only records with digitised objects. type_of_materials: e.g. Photographs and other Graphic Materials, Textual Records, Maps. level: series, fileUnit, item, recordGroup, etc. record_group: Record group number, e.g. 109 (Confederate records). ancestor_na_id: Only records beneath this naId. geographic: Geographic subject heading. creators: Creator heading. include_extracted_text: Include OCR text in the results where NARA has it. limit: Results per page (1 to 100). page: Page number from 1. kind: text, image, map, audio or video, mapped to NARA's type of materials (unverified until a key has been used against the live service). |
| nara_get_recordA | Get one Catalog record by its numeric naId, with digital objects and ancestry. Needs NARA_API_KEY. |
| nara_childrenB | List the immediate children of a series or file unit (one level down). Needs NARA_API_KEY. |
| nara_extracted_textB | Get OCR text extracted from a record's digital objects. Needs NARA_API_KEY. The response is passed through from NARA with long strings shortened to 4,000 characters. |
| si_searchA | Search Smithsonian Open Access records. Needs SMITHSONIAN_API_KEY. Records whose media are all marked CC0 come back with rights.reuse "free"; otherwise "Usage conditions apply" is surfaced as restricted. Args: q: Search words; supports AND/OR and fielded terms such as topic:Cavalry (see si_terms). rows: Results to return (1 to 100). start: Offset of the first row. sort: relevancy (default), newest or updated. type: EDAN record type, e.g. edanmdm. row_group: objects or archives. category: Search within art_design, history_culture or science_technology instead. kind: text, image, map, audio or video, added to the query as object_type or online_media_type. |
| si_get_contentA | Get one Smithsonian record by id (for example edanmdm-nmaahc_2012.36.4ab). Needs SMITHSONIAN_API_KEY. |
| si_termsB | List the vocabulary for a search field: culture, data_source, date, online_media_type, place, topic or unit_code. Useful for building fielded queries. Needs SMITHSONIAN_API_KEY. |
| si_statsA | Counts of CC0 objects and media in the collection. Needs SMITHSONIAN_API_KEY. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 23 tools
The general `search` and `get_record` tools overlap with the five per-source search and get tools (si_search, nara_search, ia_search, commons_search, dpla_search, si_get_content, nara_get_record, ia_get_item, commons_file_info, dpla_get_item). Descriptions help by distinguishing normalized output from source-specific details, but an agent could still misselect for simple queries.
Most tools follow a predictable snake_case pattern with source prefixes (nara_, si_, ia_, commons_, dpla_) for source-specific operations and unprefixed names for general operations. Minor deviations are noun-only names like `si_terms`, `si_stats`, and `usage_report`, but overall the scheme is clear and consistent.
23 tools is on the heavy side for a single server, even one covering five archives. While many tools serve distinct purposes (e.g., `ia_grep_text`, `nara_children`, `si_terms`), the overlapping general and source-specific search/get tools add redundancy.
The set covers discovery, retrieval, fulltext search, media download, rights, provenance, and usage reporting across five archives. Minor gaps remain, such as cross-source fulltext search (only Internet Archive supports it) and OCR retrieval for some sources, but core research workflows are well supported.