Skip to main content
Glama
magicianmarty

heritage-research-mcp

download_media

Fetch an original media file from a supported public archive record into the local cache and create a provenance sidecar for traceability.

Instructions

Download one original file from a record into the local cache and write a provenance sidecar.

The sidecar (<file>.provenance.json) records the source, URL, retrieval time, SHA-256 and the rights the holder stated, so the file can be traced later. Only https URLs on public addresses are fetched. If rights.reuse is not free the result carries a rights_warning: treat the file as reference only.

Args: source: internet_archive, commons, dpla, nara or smithsonian. record_id: The record's id, as returned by search. media_index: Which entry of the record's media list to fetch (default the first). media_kind: Pick the first media of this kind instead (pdf, text, image, audio, video, archive). overwrite: Fetch again even if the file is already cached. max_mb: Refuse files larger than this (default 250, or HERITAGE_MCP_MAX_DOWNLOAD_MB).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
max_mbNo
sourceYes
overwriteNo
record_idYes
media_kindNo
media_indexNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.3

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the https/public-address fetch restriction, the provenance sidecar contents (source, URL, retrieval time, SHA-256, stated rights), the rights_warning behavior when rights.reuse is not free, and cache/overwrite semantics. It omits auth requirements, rate limits, and failure modes, so it stops short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action and side-effect, then an Args block that is dense and earns its place. Slightly verbose in the sidecar sentence but no filler; structure makes the contract easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Output schema exists so return values need not be restated, and the description still covers the important non-schema behavior: sidecar writing, rights warnings, URL restrictions, and per-parameter meaning. An agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does for all six parameters: source is enumerated (internet_archive, commons, dpla, nara, smithsonian), media_kind lists valid kinds, media_index default is explained, overwrite semantics are given, and max_mb's default and HERITAGE_MCP_MAX_DOWNLOAD_MB env override are documented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+scope: 'Download one original file from a record into the local cache and write a provenance sidecar.' The 'one original file' scope and the sidecar side-effect clearly separate it from sibling read tools like get_record, ia_get_item, or commons_file_info.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied rather than stated: the description explains the media_index/media_kind selection interplay and overwrite behavior, which helps an agent call it correctly, but never says when to prefer this tool over siblings or what prerequisites (e.g., a prior search to obtain record_id) apply. No alternatives or exclusions are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.