download_media
Fetch an original media file from a supported public archive record into the local cache and create a provenance sidecar for traceability.
Instructions
Download one original file from a record into the local cache and write a provenance sidecar.
The sidecar (<file>.provenance.json) records the source, URL, retrieval time, SHA-256 and the rights
the holder stated, so the file can be traced later. Only https URLs on public addresses are fetched.
If rights.reuse is not free the result carries a rights_warning: treat the file as reference only.
Args:
source: internet_archive, commons, dpla, nara or smithsonian.
record_id: The record's id, as returned by search.
media_index: Which entry of the record's media list to fetch (default the first).
media_kind: Pick the first media of this kind instead (pdf, text, image, audio, video, archive).
overwrite: Fetch again even if the file is already cached.
max_mb: Refuse files larger than this (default 250, or HERITAGE_MCP_MAX_DOWNLOAD_MB).
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| max_mb | No | ||
| source | Yes | ||
| overwrite | No | ||
| record_id | Yes | ||
| media_kind | No | ||
| media_index | No |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||