Skip to main content
Glama
magicianmarty

heritage-research-mcp

heritage-research-mcp

CI License: MIT Python 3.11+

One MCP server for historical research across five public archives. Ask an AI assistant to find a period map, a photograph, a memoir passage or a federal record, and get back records in one shape, each with the holder's rights statement and a provenance trail for anything downloaded.

Source

Good for

Key

Internet Archive

Scanned books, memoirs, regimental histories, official reports, and the full text inside them

none

Wikimedia Commons

Freely licensed photographs, maps and prints, with the licence read from each file

none

DPLA

Finding items held by US libraries, archives and museums

free

US National Archives Catalog

Federal records, military and census material, with OCR text for many scans

free

Smithsonian Open Access

Objects, photographs, archives and library items; CC0 images with direct high-resolution files

free

It was built for sourcing reference material where "where did this come from, and may I reuse it?" matters as much as finding it.

What it looks like

Ask for a Civil War map of Fairfax County, Virginia, and search with kind: "map" returns, among others:

{
  "source": "commons",
  "title": "A Civil War field map of Fairfax County, Virginia with Fort Coccoran. LOC 2014588392",
  "date": "1861",
  "description": "Shows paths and roads that no longer exist and names of some landowners. …",
  "kind": "map",
  "rights": { "reuse": "free", "label": "Public domain", "basis": "holder" },
  "media": [{ "kind": "image", "mime": "image/jpeg", "bytes": 3960781, "width": 4615, "height": 6169 }],
  "landing_url": "https://commons.wikimedia.org/wiki/File:A_Civil_War_field_map_of_Fairfax_County,_Virginia_with_Fort_Coccoran._LOC_2014588392.jpg"
}

And ia_fulltext_search finds a name inside a scanned book, with the page and a snippet (matches in **bold**):

Mosby's War Reminiscences (1887): "…Capture of a Federal Picket at Herndon Station. The Dash and Excitement of a Cavalry Skirmish…"

ia_grep_text then finds every occurrence in that book and ia_read_text reads around one. download_media caches an original (a 7 MB high-resolution CC0 photograph, say) and writes a .provenance.json beside it.

Related MCP server: Arke Institute MCP Server

Searching by type

search takes kind: text, image, map, audio or video. Each archive is asked in its own vocabulary, every result says how (kind_applied), and every record carries a best-effort kind of its own.

kind

Internet Archive

Commons

DPLA

Smithsonian

NARA

text

mediatype:texts

filetype:office (PDF, DjVu)

type=text

scanned books, full-text documents, books, manuscripts

Textual Records ¹

image

mediatype:image

filetype:bitmap

type=image

online_media_type:"Images"

Photographs and other Graphic Materials ¹

map

the "maps" subject and the map collections

map in the file title

images with the subject "Maps"

object_type:"Maps"

Maps and Charts ¹

audio

mediatype:audio

filetype:audio

type=sound

online_media_type:"Sound recordings"

Sound Recordings ¹

video

mediatype:movies

filetype:video

type=moving image

online_media_type:"Video recordings"

Moving Images ¹

¹ NARA's vocabulary comes from its documentation and has not yet been run against a live key.

Do maps work? Yes, with honest limits. Commons, DPLA and the Smithsonian have solid map cataloguing and return real maps (the example above came from this filter). The Internet Archive has no map type at all, so map there is a subject-and-collection match and is noisy. Because archives catalogue differently, treat kind as a strong hint rather than a guarantee, and run a query both with and without it when something seems missing. Multi-word queries need every word to match at the four archives tested (NARA is not yet verified), so fewer words find more.

Text versus images. kind selects whole records. To search inside documents, use ia_fulltext_search (the OCR text of Internet Archive books), and nara_search with include_extracted_text or nara_extracted_text for NARA scans.

Tools

Tool

What it does

list_sources, usage_report

What is ready to use, and how many requests have been made

search

Query every ready archive at once, optionally by kind, date range, and including full text inside books

get_record

One record, normalised, with its media files and rights

download_media

Cache one original locally with a provenance sidecar

ia_search, ia_fulltext_search, ia_get_item, ia_read_text, ia_grep_text

Internet Archive: metadata search, search inside books, files, and reading or grepping OCR text

commons_search, commons_file_info, commons_category_members

Commons files and categories

dpla_search, dpla_get_item

DPLA

nara_search, nara_get_record, nara_children, nara_extracted_text

National Archives Catalog

si_search, si_get_content, si_terms, si_stats

Smithsonian Open Access

A source that needs a key is skipped, with a message saying where to get one, until the key is set. The Internet Archive and Commons work immediately.

Search results are compact (a short description, the first two files, no thumbnails) so that many can be read cheaply; get_record and the *_get_* and *_file_info tools return the full record.

Install

You need uv (or any Python 3.11+ environment).

Claude Code

claude mcp add heritage-research -- uvx --from git+https://github.com/magicianmarty/heritage-research-mcp heritage-research-mcp

Claude Desktop and other MCP clients

{
  "mcpServers": {
    "heritage-research": {
      "command": "uvx",
      "args": ["--from", "git+https://github.com/magicianmarty/heritage-research-mcp", "heritage-research-mcp"]
    }
  }
}

Pin a release for reproducible installs: git+https://github.com/magicianmarty/heritage-research-mcp@v0.1.0.

Hosted agent platforms. Any platform that runs stdio MCP servers can run it. Use the same uvx command, pass the keys as environment variables from the platform's secret store, and set HERITAGE_MCP_DISABLE_DOWNLOADS=1 (a hosted agent cannot read files the server downloads, so download_media is switched off and the text-returning tools do the work).

From source

git clone https://github.com/magicianmarty/heritage-research-mcp && cd heritage-research-mcp
uv venv && uv pip install -e ".[dev]"
.venv/bin/heritage-research-mcp doctor

Works with both the 1.x and 2.x generations of the MCP Python SDK.

Keys

Three sources need a free key. docs/KEYS.md walks through each. Put a key in an environment variable or on one line in ~/.config/heritage-research-mcp/keys/<dpla|nara|smithsonian> (mode 600). Files are the better choice: they stay out of client configs and shell history.

heritage-research-mcp doctor --live     # shows what is configured, then makes one small real request to each source

Configuration

Variable

Default

Purpose

DPLA_API_KEY, NARA_API_KEY, SMITHSONIAN_API_KEY

none

API keys (or key files, above)

HERITAGE_MCP_CONTACT

none

Added to the User-Agent, so providers can reach you

HERITAGE_MCP_CACHE_DIR

~/.cache/heritage-research-mcp

Where download_media writes

HERITAGE_MCP_MAX_DOWNLOAD_MB

250

Refuse larger downloads

HERITAGE_MCP_DISABLE_DOWNLOADS

off

Remove download_media (hosted use)

NARA_MONTHLY_LIMIT

10000

Stop before NARA's monthly quota

NARA_API_VERSION

v2

v3 to use NARA's newer search

Rights

Archives state rights in many vocabularies: Creative Commons URLs, rightsstatements.org URIs, free text, or nothing. Each record's rights reduces that to a reuse value:

reuse

Meaning

free

Public domain, CC0, or a no-copyright statement

attribution

Free if credited (rights.attribution has the text)

share_alike

Reuse must carry the same licence

non_commercial, no_derivatives, restricted

Conditions that limit reuse

unknown

The holder stated nothing. This is not permission.

rights.basis says where the value came from: holder, or date-heuristic for one narrow case (a US text first published 96 or more years ago). This is a filtering aid, not legal advice. download_media adds a rights_warning when an item is not marked free.

Being a good citizen

  • Requests are paced per source and retried with backoff; Retry-After is honoured. Every request carries a User-Agent naming this project.

  • NARA allows 10,000 queries per key per month and forbids scraping or bulk download through the API. This server counts requests, stops at the limit, returns one page of at most 100 records per call, and includes the attribution NARA requires.

  • The Internet Archive full-text endpoint is experimental and may change.

Security

  • Keys are never returned in results and are redacted from errors and logs (HTTP request logging is filtered, because two providers take the key as a query parameter).

  • download_media fetches only https URLs that resolve to public addresses, refuses credentials and non-standard ports, re-checks every redirect, and enforces a size limit. The URL always comes from a record a source returned, never directly from the model.

  • See SECURITY.md to report a problem.

Status

Verified against the live services on 6 October 2026: the Internet Archive, Wikimedia Commons, DPLA and Smithsonian Open Access. The NARA adapter follows that provider's published API specification and is tested against fixtures of the documented response shapes; it has not yet been exercised with a real key. If you find a response shape that differs, please open an issue. The Live smoke workflow re-checks every configured source weekly.

Development

uv venv && uv pip install -e ".[dev]"
.venv/bin/pytest            # offline: every request is mocked
.venv/bin/ruff check src tests && .venv/bin/pyright src

CI runs the suite on Python 3.11 and 3.13 against both SDK generations. See docs/DESIGN.md for how it is put together.

Licence

MIT. Records and files returned by this server remain subject to their holders' terms.

Available Tools

23 tools
commons_category_membersA

List the files in a Commons category (for example "Mosby's Rangers").

Args: category: The category name, with or without the "Category:" prefix. limit: Results to return (1 to 50). filetype: Optionally keep only bitmap, drawing, audio or video files.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
categoryYes
filetypeNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses the limit range (1 to 50) and that filetype acts as a filter, but says nothing about ordering, pagination beyond the limit cap, or what happens with an unknown category. For a low-risk read tool this is adequate but thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence stating the operation plus a compact Args block. Every line carries information; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With only 3 parameters, an output schema present, and all parameters explained, an agent has enough to invoke it correctly. Minor gaps remain around result ordering and behavior past the 50-item cap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it does: it explains the 'Category:' prefix is optional, gives the limit range 1-50, and supplies the filetype vocabulary (bitmap, drawing, audio, video) that the schema lacks. The limit default (20) is not mentioned, keeping it short of a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List the files in a Commons category') and anchors it with a concrete example ('Mosby's Rangers'). It is clearly distinct from siblings like commons_search (finding) and commons_file_info (metadata for one file).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by naming what it lists, but never says when to prefer it over commons_search or commons_file_info, nor does it mention any prerequisites. No exclusions or alternative-routing guidance is present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

commons_file_infoA

Get one Commons file's licence, author, description, categories and original size.

Args: title: The file name, with or without the "File:" prefix.

ParametersJSON Schema
NameRequiredDescriptionDefault
titleYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses the returned fields, which implicitly signals a non-mutating metadata read, but it says nothing about authentication, rate limits, failure modes (e.g. missing file), or whether the lookup is exact-match or normalized.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence listing the returned fields, followed by a compact Args block for the sole parameter. Nothing is padded and the essential information comes first.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the description need not explain return structure, yet it usefully previews the returned fields and covers the only parameter's format. The main remaining gap is the absence of any error/empty-result behavior for a nonexistent title.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, but with a single parameter the description supplies the key semantic detail the schema lacks: the title is a file name that may be given with or without the 'File:' prefix. That normalization rule is exactly the kind of information an agent needs and would not get from the bare string type.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Get) and resource (one Commons file) and enumerates exactly what is returned: licence, author, description, categories, original size. This is clearly distinguishable from siblings like commons_search and commons_category_members, which operate on collections rather than a single file.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the singular 'one Commons file', which suggests it is the detail lookup paired with commons_search, but no sibling is named and there is no explicit when-to-use or when-not-to-use guidance. Adequate but leaves routing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

download_mediaA

Download one original file from a record into the local cache and write a provenance sidecar.

The sidecar (<file>.provenance.json) records the source, URL, retrieval time, SHA-256 and the rights the holder stated, so the file can be traced later. Only https URLs on public addresses are fetched. If rights.reuse is not free the result carries a rights_warning: treat the file as reference only.

Args: source: internet_archive, commons, dpla, nara or smithsonian. record_id: The record's id, as returned by search. media_index: Which entry of the record's media list to fetch (default the first). media_kind: Pick the first media of this kind instead (pdf, text, image, audio, video, archive). overwrite: Fetch again even if the file is already cached. max_mb: Refuse files larger than this (default 250, or HERITAGE_MCP_MAX_DOWNLOAD_MB).

ParametersJSON Schema
NameRequiredDescriptionDefault
max_mbNo
sourceYes
overwriteNo
record_idYes
media_kindNo
media_indexNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the https/public-address fetch restriction, the provenance sidecar contents (source, URL, retrieval time, SHA-256, stated rights), the rights_warning behavior when rights.reuse is not free, and cache/overwrite semantics. It omits auth requirements, rate limits, and failure modes, so it stops short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action and side-effect, then an Args block that is dense and earns its place. Slightly verbose in the sidecar sentence but no filler; structure makes the contract easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists so return values need not be restated, and the description still covers the important non-schema behavior: sidecar writing, rights warnings, URL restrictions, and per-parameter meaning. An agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does for all six parameters: source is enumerated (internet_archive, commons, dpla, nara, smithsonian), media_kind lists valid kinds, media_index default is explained, overwrite semantics are given, and max_mb's default and HERITAGE_MCP_MAX_DOWNLOAD_MB env override are documented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+scope: 'Download one original file from a record into the local cache and write a provenance sidecar.' The 'one original file' scope and the sidecar side-effect clearly separate it from sibling read tools like get_record, ia_get_item, or commons_file_info.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: the description explains the media_index/media_kind selection interplay and overwrite behavior, which helps an agent call it correctly, but never says when to prefer this tool over siblings or what prerequisites (e.g., a prior search to obtain record_id) apply. No alternatives or exclusions are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

dpla_get_itemB

Get one DPLA item by its id. Needs DPLA_API_KEY.

ParametersJSON Schema
NameRequiredDescriptionDefault
item_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it does disclose the auth requirement (API key), which is genuinely useful and not present in structured fields. It says nothing about error behavior, rate limits, or what happens with an invalid/missing id, so it is only partially transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short, front-loaded sentences with zero filler. The core action comes first and the prerequisite second.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple single-parameter retrieval tool with an output schema present, the description covers the essential action and the auth prerequisite. The return shape is handled by the output schema, so the only real gap is the lack of id-format detail and sibling routing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single parameter item_id has no schema-level documentation. The phrase 'by its id' hints that item_id is an identifier, but nothing about format, source, or example is added, so the description does little to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Get one DPLA item by its id'), which is enough for an agent to know it retrieves a single item rather than listing or searching. It does not explicitly distinguish itself from the sibling dpla_search, but the singular 'one item by its id' does imply the difference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a prerequisite ('Needs DPLA_API_KEY') but no guidance on when to use this versus dpla_search, or when an item id is the right input. No alternatives or exclusions are named, so the agent must infer usage from the tool name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_recordA

Fetch one record in the normalised shape, including its media files and rights.

Args: source: internet_archive, commons, dpla, nara or smithsonian. record_id: The id from a search result (an Internet Archive identifier, a Commons "File:" title, a DPLA id, a NARA naId, or a Smithsonian id).

ParametersJSON Schema
NameRequiredDescriptionDefault
sourceYes
record_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses that the payload includes media files and rights, and the verb 'Fetch' implies a read-only operation, but it says nothing about auth, rate limits, error behavior for invalid source/record_id, or whether the fetch is cached.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded one-sentence purpose followed by a compact Args block; no filler. Slightly redundant restating of parameter names already visible in the schema, but overall efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be re-explained, and the parameter documentation is solid. The gap is routing: with ~20 siblings including several source-specific 'get' tools, the description never says when this generic getter is preferable to nara_get_record/dpla_get_item/ia_get_item.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the schema provides no enum for 'source', so the description does the heavy lifting: it enumerates the five valid sources and explains the id format for each source (IA identifier, Commons 'File:' title, DPLA id, NARA naId, Smithsonian id). That is substantially more than the schema offers, though it doesn't state what happens on an unrecognized source or mismatched id format.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Fetch one record') plus what the returned object contains (normalised shape, media files, rights). It does not, however, distinguish itself from the source-specific getters in the sibling list (nara_get_record, dpla_get_item, ia_get_item), so an agent can't tell from the description alone which getter to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: 'record_id: The id from a search result' hints that this follows a search, but there is no explicit when-to-use/when-not statement and no comparison against the numerous per-source getter siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ia_get_itemB

Get one item's metadata, rights and downloadable files (PDF, OCR text, page images).

ParametersJSON Schema
NameRequiredDescriptionDefault
identifierYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It usefully discloses what content comes back (PDF, OCR text, page images), but says nothing about whether files are actually downloaded, permission/authentication needs, rate limits, or response size limits for large items.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One front-loaded sentence with zero filler; the verb and resource lead, and the parenthetical enumerates return content compactly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the return shape need not be documented, and for a single-parameter retrieval tool the description is close to sufficient. The remaining gap is the identifier's expected form, which neither schema nor description explains.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one required parameter with 0% schema description coverage, and the description adds nothing about it — no identifier format, no example (e.g., an Internet Archive item slug), no statement that it must be an exact item id rather than a query. The single self-evident name does not fully compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (get) and resource (one item) and enumerates what the caller receives: metadata, rights, downloadable files. That distinguishes it from the sibling search tools (ia_search, ia_fulltext_search) and the text tools (ia_read_text, ia_grep_text), though it never names an alternative explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: a caller with an item identifier who wants the item's metadata and files would reach for this. There is no statement of when to prefer it over ia_read_text or ia_grep_text, nor any prerequisite such as needing an identifier obtained from ia_search.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ia_grep_textA

Find every occurrence of a word or phrase in an item's OCR text, with surrounding context and offsets.

The text is fetched once and held in memory, so repeated searches of the same book are fast.

Args: identifier: The Internet Archive identifier. pattern: Text to find. A plain phrase matches across the line breaks, repeated spaces and split hyphens that OCR leaves in scanned books ("twenty-five prisoners" finds "twenty-five prisoners" and "twenty- five prisoners"). A regular expression only if regex is true. regex: Treat the pattern as a regular expression. ignore_case: Ignore capitalisation. context: Characters of context on each side (0 to 1000). max_matches: Most matches to return (1 to 100); the total is always reported.

ParametersJSON Schema
NameRequiredDescriptionDefault
regexNo
contextNo
patternYes
identifierYes
ignore_caseNo
max_matchesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does well: it discloses the in-memory fetch-and-cache behavior (performance trait), and that "the total is always reported" even when matches are truncated. It omits auth/rate-limit details, but for a read-only grep these are minor.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded purpose sentence followed by a performance note, then a clean Args block. Every sentence earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values needn't be explained, and all six parameters are documented. The only gap is the lack of explicit routing guidance versus sibling tools, which the description could have added.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: it explains normalization behavior for the pattern (line breaks, repeated spaces, split hyphens), the meaning of the regex flag, the context range (0–1000), and the max_matches range plus total reporting. Every parameter gains meaning beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states a specific verb ("Find every occurrence") and resource ("a word or phrase in an item's OCR text") plus the return shape ("with surrounding context and offsets"). This clearly distinguishes it from siblings like ia_read_text and ia_fulltext_search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage through "repeated searches of the same book are fast," suggesting it's for iterative querying of one item, but it never explicitly says when to choose this over ia_read_text or ia_fulltext_search. Usage context is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ia_read_textA

Read a slice of an item's OCR text by character offset (up to 20,000 characters at a time).

Use ia_grep_text first to find the offset you want.

ParametersJSON Schema
NameRequiredDescriptionDefault
startNo
lengthNo
identifierYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it does disclose two useful traits: the operation is offset/slice-based and is capped at 20,000 characters per call. However, it omits auth requirements, behavior when start is past the end of the text, whether the cap is enforced or truncated, and any error/pagination semantics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, zero filler, with the core capability front-loaded and the routing hint second. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained. The description supplies the operation, the slicing mechanism, the size cap, and the recommended sibling workflow, which is sufficient for an agent to call this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for all three parameters. It conveys the offset concept and the 20,000-character ceiling, which maps to start/length, but never names the parameters, never mentions the length default of 4000 or start default of 0, and says nothing about the required identifier's expected form.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Read a slice of an item's OCR text') plus the mechanism ('by character offset') and a hard cap ('up to 20,000 characters at a time'). It does not explicitly differentiate itself from similar text-retrieval siblings like nara_extracted_text or si_get_content, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent to a prerequisite workflow: 'Use ia_grep_text first to find the offset you want.' That is clear when-to-use guidance tied to a named sibling. It gives no when-not or alternate-path guidance (e.g., when to use ia_get_item instead), keeping it below 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_sourcesA

List the archives this server can search, whether each is ready to use, and what it is good for.

Sources that need a free API key show configured: false until a key is set; the Internet Archive and Wikimedia Commons need none.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the load and does reasonably well: it explains the `configured: false` state, that a free API key clears it, and that Internet Archive and Wikimedia Commons need no key. It stops short of describing the full response shape, though the output schema covers that.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences: the core purpose first, then the operational nuance about API keys. No filler, and the most important information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a no-parameter discovery tool with an output schema, the description is complete: it says what is returned (archives, readiness, suggested use) and covers the one gotcha (keyless vs keyed sources). Return shape is delegated correctly to the output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, which is the baseline 4 case. Schema coverage is 100% and there is nothing to document, so no descriptive burden exists here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (the archives this server can search) plus the payload contents (readiness and fitness). It is immediately distinguishable from the search/get siblings, which all operate on content rather than the source catalog itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies a discovery/preflight role by explaining readiness and usefulness, but never explicitly says when to call it (e.g., before searching, or to determine which sources are available). No alternative is named or excluded, leaving usage to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nara_childrenB

List the immediate children of a series or file unit (one level down). Needs NARA_API_KEY.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageNo
limitNo
parent_na_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does add one genuinely useful operational fact: an API key is required. However, it says nothing about pagination behavior, ordering, or what happens when the parent has no children or an invalid ID, so behavioral disclosure is partial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with zero filler; the purpose and scope lead, and the auth prerequisite follows. Nothing is redundant or padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. Still, for an un-annotated tool with zero parameter documentation, the description omits pagination semantics and parent-ID format, leaving real gaps an agent would only discover by trial.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across three parameters, so the description must compensate and largely does not. 'parent_na_id' is only obliquely implied by 'children of a series or file unit', and 'page'/'limit' are never mentioned at all — an agent gets no hint that results are paged.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and resource ('immediate children of a series or file unit') with precise scope ('one level down'), which is more than a restatement of the name. It does not explicitly differentiate itself from siblings like nara_get_record or nara_search, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the scope: call this to walk one level below a known parent. There is no explicit when-to-use/when-not guidance and no named alternative (e.g. nara_search for discovery vs. this for hierarchy traversal), leaving routing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nara_extracted_textB

Get OCR text extracted from a record's digital objects. Needs NARA_API_KEY.

The response is passed through from NARA with long strings shortened to 4,000 characters.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageNo
limitNo
na_idYes
object_idNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses the auth requirement and the 4,000-character truncation of long strings, which is real behavioral context. However it says nothing about pagination behavior (despite page/limit params), what happens for records without digital objects or OCR, or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core action and the truncation caveat second. No filler, though the truncation note could be tied to the limit parameter for tighter structure.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value explanation is not required. But for a 4-param tool with 0% schema coverage and no annotations, the description leaves pagination and object selection unexplained, which is a meaningful gap for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 4 parameters. The description only hints at the record identity and says nothing about page, limit, or object_id — the pagination and object-selection semantics are entirely undocumented in both the schema and the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: retrieve OCR text extracted from a record's digital objects. An agent can distinguish this from nara_get_record (metadata) by the 'extracted text' scope, but the description never explicitly names or contrasts the sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the auth prerequisite (NARA_API_KEY) but gives no when-to-use guidance, no condition that selects this over nara_get_record or the other text tools, and no note on when OCR may be unavailable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

nara_get_recordA

Get one Catalog record by its numeric naId, with digital objects and ancestry. Needs NARA_API_KEY.

ParametersJSON Schema
NameRequiredDescriptionDefault
na_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses an auth requirement (NARA_API_KEY) and hints at payload contents (digital objects, ancestry), but says nothing about rate limits, error behavior if the naId is absent/invalid, or rate/scope of the read.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the core action and scope, then the prerequisite. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description still flags payload scope. For a single-parameter read tool the definition is nearly sufficient; only the missing usage routing and edge-case behavior keep it from being complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the only parameter is titled 'Na Id', so the description's phrase 'numeric naId' is what tells the agent this is a NARA catalog identifier and that it must be an integer. That adds real meaning but is minimal for a 1-param tool with no schema documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Get) and resource (Catalog record) and pins the lookup key to the numeric naId, with extra scope ('with digital objects and ancestry'). This distinguishes it from the nara_search sibling, though it does not explicitly contrast with the generic get_record sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the agent must already hold a numeric naId (presumably from nara_search). There is no statement of when to use this over nara_search or nara_children, and no exclusions or prerequisites beyond the key requirement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

si_get_contentA

Get one Smithsonian record by id (for example edanmdm-nmaahc_2012.36.4ab). Needs SMITHSONIAN_API_KEY.

ParametersJSON Schema
NameRequiredDescriptionDefault
content_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses the auth requirement (API key), which is genuine value, but says nothing about read-only nature, failure behavior for unknown ids, or rate limits – though an output schema exists to cover the return shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, no waste, and the core action and identifier format are front-loaded ahead of the auth note.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter getter with an output schema, the description covers the essentials: what it fetches, the id format, and the credential needed. Only marginal gaps remain around error/not-found behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With one parameter at 0% schema coverage, the description must compensate, and it does by supplying a concrete id example ('edanmdm-nmaahc_2012.36.4ab') that shows the expected format. It does not explain whether the id is a Smithsonian EDAN id versus some other key, so it is not fully compensating.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Get one Smithsonian record by id') and gives a concrete example identifier, which implicitly separates it from the search siblings like si_search. It stops short of naming an alternative explicitly, so it lands just below the top tier.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a prerequisite ('Needs SMITHSONIAN_API_KEY') but never states when to reach for this tool versus si_search or get_record. Usage is only implied by 'by id' – an agent with an identifier will infer it, but there is no explicit when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

si_statsA

Counts of CC0 objects and media in the collection. Needs SMITHSONIAN_API_KEY.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It usefully discloses the auth requirement and that counts cover only CC0-licensed content, but says nothing about whether the call is read-only (implied), rate limits, or whether counts are collection-wide or scoped. The output schema presumably covers the return shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the result the tool produces and followed by the auth prerequisite. Nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described, and the description covers the resource scope and the auth requirement. It could say more about whether the stats are global or filterable, but for a zero-parameter counter it is nearly complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4 — there are no argument semantics left for the description to explain. Schema coverage is reported at 100%.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource (CC0 objects and media counts) with the implied verb 'counts', and the 'CC0' scoping makes it distinguishable from the search/fetch siblings. It does not explicitly name an alternative, but no sibling does anything similar, so confusion risk is low.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a prerequisite ('Needs SMITHSONIAN_API_KEY') but no explicit when-to-use or when-not-to-use guidance. Usage is implied by the purpose — call it to get aggregate collection counts — but there is no routing or exclusion language.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

si_termsB

List the vocabulary for a search field: culture, data_source, date, online_media_type, place, topic or unit_code. Useful for building fielded queries. Needs SMITHSONIAN_API_KEY.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
categoryYes
starts_withNo

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses the auth requirement (SMITHSONIAN_API_KEY), but says nothing about pagination/limit behavior, rate limits, or result shape. The auth note is real added value, but the safety/operational profile is thin for a no-annotation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the core action and then the enumerated vocabulary. The category list is long but earns its place as the key semantic payload; nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained, and the auth requirement is covered. However, two of three parameters remain undefined and there is no guidance on result size or how to consume the vocabulary, leaving the definition only adequate for a 3-param tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does document the accepted values for the required `category` parameter (culture, data_source, date, place, topic, unit_code, etc.), which is genuinely useful, but `limit` and `starts_with` are left entirely undocumented in both schema and description. Partial compensation only.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (list) and resource (the vocabulary for a search field), and enumerates the exact categories the field accepts. An agent can tell this is a lookup/controlled-vocabulary tool rather than a search tool like si_search, though it never names a sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Useful for building fielded queries" implies when to reach for it, but there is no explicit when-not guidance or named alternative (e.g. si_search for the actual query). Usage is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

usage_reportA

Requests made this month and this session per source, and the provider rate limit last seen.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full behavioral burden. It does disclose the temporal scope (this month, this session) and the per-source granularity, which is real added context, but it never states that the call is a side-effect-free read, whether it requires auth, or whether it is cheap to call repeatedly.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; it leads with what the caller gets and moves straight to scope details. Nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the description is not obliged to explain the return shape, and it correctly focuses on the data's scope instead. For a zero-parameter read tool this is nearly sufficient; only the auth/side-effect profile is left unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline is 4. There are no inputs whose semantics need explaining, and the schema has no gaps to compensate for.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the specific resource and scope: request counts for this month and this session, broken down per source, plus the last-seen provider rate limit. That is enough to distinguish it from the retrieval-oriented siblings (si_stats aside), though it omits an explicit verb like 'report' and never contrasts itself with si_stats.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use statement or named alternative, but the mention of 'the provider rate limit last seen' strongly implies the motivating scenario: checking quota before or after sustained API use. Usage is inferred rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 23 tool updatesv0.1.3
    • First observedcommons_category_members
    • First observedcommons_file_info
    • First observedcommons_search
    • First observeddownload_media
    • First observeddpla_get_item
    • First observeddpla_search
    • First observedget_record
    • First observedia_fulltext_search
    • First observedia_get_item
    • First observedia_grep_text
    • First observedia_read_text
    • First observedia_search
    • First observedlist_sources
    • First observednara_children
    • First observednara_extracted_text
    • First observednara_get_record
    • First observednara_search
    • First observedsearch
    • First observedsi_get_content
    • First observedsi_search
    • First observedsi_stats
    • First observedsi_terms
    • First observedusage_report

TDQS

A3.5/5.0

Scored across 23 tools

Disambiguation3/5

The general `search` and `get_record` tools overlap with the five per-source search and get tools (si_search, nara_search, ia_search, commons_search, dpla_search, si_get_content, nara_get_record, ia_get_item, commons_file_info, dpla_get_item). Descriptions help by distinguishing normalized output from source-specific details, but an agent could still misselect for simple queries.

Naming Consistency4/5

Most tools follow a predictable snake_case pattern with source prefixes (nara_, si_, ia_, commons_, dpla_) for source-specific operations and unprefixed names for general operations. Minor deviations are noun-only names like `si_terms`, `si_stats`, and `usage_report`, but overall the scheme is clear and consistent.

Tool Count3/5

23 tools is on the heavy side for a single server, even one covering five archives. While many tools serve distinct purposes (e.g., `ia_grep_text`, `nara_children`, `si_terms`), the overlapping general and source-specific search/get tools add redundancy.

Completeness4/5

The set covers discovery, retrieval, fulltext search, media download, rights, provenance, and usage reporting across five archives. Minor gaps remain, such as cross-source fulltext search (only Internet Archive supports it) and OCR retrieval for some sources, but core research workflows are well supported.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    Enables users to search and access digital collections from the Swedish National Archives (Riksarkivet) through multiple APIs. Supports searching records by keywords, exploring collections, and downloading historical images and documents.
    2
    26
    Apache 2.0
  • F
    license
    Not graded
    quality
    Not graded
    maintenance
    Enables semantic search across the Arke Institute's extensive archive of NARA records and presidential libraries using natural language queries. Provides access to millions of historical documents, photographs, and records with OCR'd content and complete metadata.
    -
  • A
    license
    A
    quality
    A
    maintenance
    Provides access to 61 public digital libraries through a single unified interface, enabling users to search and retrieve information from academic papers, books, legal records, and more using natural language.
    11
    35 npm
    2
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables natural language search and exploration of the Dutch WWII Oorlogsbronnen archives, allowing users to query historical documents, photographs, and personal accounts through AI assistants.
    35 npm
    15
    MIT