@cyanheads/internet-archive-mcp-server
Provides tools to search the Wayback Machine and Internet Archive library (40M+ items), fetch archived snapshots, retrieve item metadata and full text, enabling AI agents to access historical web content and digital collections.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@@cyanheads/internet-archive-mcp-serverFind Wayback Machine snapshots of example.com from 2020"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Overview
The Wayback Machine and Internet Archive library (40M+ items). Find and fetch archived snapshots of any URL, search the library by keyword and metadata, and retrieve item metadata, file manifests, and OCR text from any MCP client. Runs as a stdio process or a local Streamable HTTP server.
Tools
Tool | Description |
| Find Wayback Machine snapshots of a URL, by closest timestamp or full capture history |
| Fetch archived page content at a specific Wayback timestamp |
| Search the IA library (40M+ items) by keyword and metadata filters |
| Retrieve full metadata and file manifest for an Archive item |
| Retrieve readable OCR text from a text item, with paging |
Resources
Resource | Description |
| Metadata snapshot for an Archive item — title, creator, mediatype, description, subjects, collections, date, license, and file count |
All resource data is also reachable via ia_get_item.
Related MCP server: MCP Wayback Machine Server
Capability reference
ia_find_snapshots tool
closestmode: returns the nearest capture to a giventimestampvia the Availability API, falling back to one CDX closest-capture query (preferring a200capture, 25 s deadline) when the Availability API has no answerhistorymode: full capture list via the CDX API; filter by date range (from/to), HTTP status (status_filter), and MIME typeDefault
collapseoftimestamp:8(one capture per day); adjustable totimestamp:N, N=1–14Up to 10,000 records per call (
limit, default 100);resume_keypagination for large historiesReplay URLs are always
https://web.archive.org/web/…Typed errors:
missing_timestamp(closest mode without atimestamp, answered before any lookup),no_snapshots(no matches),no_snapshot_available(closest mode, no capture near timestamp from either source),cdx_unavailable(CDX 5xx, 429, or unreadable response, or the closest-mode CDX check did not complete),availability_unavailable(closest mode, Availability API 5xx, 429, or unreadable response); a 429 is answered after one request with a hint to wait
ia_get_snapshot tool
Resolves to the nearest available capture when the exact timestamp has no snapshot; exact 14-digit timestamps skip resolution
Reports the capture Wayback served —
replay_url,resolved_timestamp, andresolved_statusfollow Wayback's redirect when the requested timestamp is not itself a captureDecodes the page with its declared charset (
Content-Type, then<meta>, else UTF-8 when the bytes are valid UTF-8, else Wayback's guessed charset or windows-1252), then removes scripts, styles, comments, and tags and decodes character references, returning readable plain text alongside the replay URLReads at most the first 4 MiB of a page (a
noticesays when a page is longer); output capped atIA_MAX_SNAPSHOT_CHARS(default 50,000 characters)Typed errors:
no_snapshot_available,content_fetch_failed(Wayback unreachable during lookup or fetch)
ia_search_items tool
Solr query syntax plus structured filters:
mediatype,collection,creator,language, and date range (date_from/date_to)mediatypetakes the ten Internet Archive media types —texts,movies,audio,software,image,data,web,collection,etree,account— case-insensitively, and resolves common near-misses (text,book,books→texts;movie,video,videos→movies;images→image;collections→collection)Sort by relevance, date, or downloads (
sort, Solr syntax; defaultdownloads desc)Up to 200 results per page (
rows, default 50), 1-indexedpageOutput carries
total_found,page,rowsfor pagination; an empty page returns anoticerather than an error — naming themediatypeapplied when nothing matched, or the last page whenpageis past the endTyped error
invalid_mediatypefor any othermediatype, listing the accepted values, answered before any search request
ia_get_item tool
Returns
title,creator,description,subject,collection,licenseurl,rights, andlanguagewhen present in upstream metadata;creator,description,subject,collection, andlanguagemay be a string or a listfiles[]is one page of the manifest in upstream order —format,size,md5, and a directdownload_urlper file;file_countis always the full manifest sizemax_files(1–500, default 50) andfile_offset(default 0) page through large items; when files remain, the response setstruncatedand names the nextfile_offsetformatkeeps one file type (exact match, case-insensitive — e.g.DjVuTXT,Text PDF,VBR MP3) before paging and reports the match count astotalCount; a format with no matches returns an empty page and a notice listing the formats the item hasTyped error
item_not_foundfor unknown identifiers
ia_get_text tool
max_chars(defaults toIA_MAX_SNAPSHOT_CHARS) andchar_offsetpage through long documents;has_moresignals additional text remainsLocates the best available text file — DjVuTXT preferred, falls back to plain text;
source_filenames the file fetchedTyped errors:
item_not_found,no_text_file,download_forbidden(restricted collections)
ia://item/{identifier} resource
Returns
application/json—title,creator,mediatype,description,subject,collection,date,licenseurl,rights,language, andfile_countidentifiercomes fromia_search_itemsresultsTyped error
item_not_foundfor unknown identifiers
Features
Built on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.
Internet Archive-specific:
No credentials required — all four APIs are public
Three service layers:
WaybackService(Availability + CDX),ArchiveSearchService(Solr),ArchiveMetadataService(Metadata + downloads)CDX collapse-by-day default and configurable
limitkeep responses tractable for high-capture URLsIdentifies via a custom User-Agent on every request as required by IA's terms of use; configurable via
IA_USER_AGENT
Agent-friendly output:
Pagination context on every list response —
total_found,page,rows(search),resume_key(CDX history), andfile_countplus the nextfile_offset(item files) so agents never have to guess whether results are completeTyped error reasons (
missing_timestamp,no_snapshots,no_snapshot_available,cdx_unavailable,availability_unavailable,content_fetch_failed,invalid_mediatype,item_not_found,no_text_file,download_forbidden) with recovery hints so callers can retry or explain to users without parsing textStructured file manifests —
ia_get_itemreturns file-level metadata (format, size, URL) and aformatfilter, so agents can pick the right file without paging through thumbnails
Getting started
No API key required — the Internet Archive's APIs are fully public.
Add the following to your MCP client configuration file:
{
"mcpServers": {
"internet-archive-mcp-server": {
"type": "stdio",
"command": "bunx",
"args": ["@cyanheads/internet-archive-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info"
}
}
}
}Or with npx (no Bun required):
{
"mcpServers": {
"internet-archive-mcp-server": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@cyanheads/internet-archive-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info"
}
}
}
}Or with Docker:
{
"mcpServers": {
"internet-archive-mcp-server": {
"type": "stdio",
"command": "docker",
"args": [
"run", "-i", "--rm",
"-e", "MCP_TRANSPORT_TYPE=stdio",
"ghcr.io/cyanheads/internet-archive-mcp-server:latest"
]
}
}
}For Streamable HTTP, set the transport and start the server:
MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 bun run start:http
# Server listens at http://localhost:3010/mcpPrerequisites
Bun v1.4.0 or higher (or Node.js v24+).
No external accounts or API keys required.
Installation
Clone the repository:
git clone https://github.com/cyanheads/internet-archive-mcp-server.gitNavigate into the directory:
cd internet-archive-mcp-serverInstall dependencies:
bun installConfigure environment:
cp .env.example .env
# Optional: edit .env for custom User-Agent, timeouts, etc.Configuration
All configuration is validated at startup via Zod schemas in src/config/server-config.ts.
Variable | Description | Default |
| Transport: |
|
| HTTP server port |
|
| Auth mode: |
|
| Log level ( |
|
| Directory for log files (Node.js only) |
|
| Storage backend |
|
| Enable OpenTelemetry instrumentation |
|
| Custom User-Agent for IA API requests |
|
| HTTP request timeout in milliseconds |
|
| Default character cap for |
|
See .env.example for the full list of optional overrides.
Running the server
Local development
Build and run:
# One-time build bun run rebuild # Run the built server bun run start:stdio # or bun run start:httpRun checks and tests:
bun run devcheck # Lint, format, typecheck, security bun run test # Vitest test suite bun run lint:mcp # Validate MCP definitions against spec
Docker
docker build -t internet-archive-mcp-server .
docker run --rm -p 3010:3010 internet-archive-mcp-serverThe Dockerfile defaults to HTTP transport, stateless session mode, and logs to /var/log/internet-archive-mcp-server. OpenTelemetry peer dependencies are installed by default — build with --build-arg OTEL_ENABLED=false to omit them.
Project structure
Directory | Purpose |
|
|
| Server-specific environment variable parsing and validation with Zod. |
| Tool definitions ( |
| Resource definitions. |
|
|
|
|
|
|
| Unit and integration tests mirroring |
Development guide
See CLAUDE.md for development guidelines and architectural rules. The short version:
Handlers throw, framework catches — no
try/catchin tool logicUse
ctx.logfor request-scoped logging,ctx.statefor tenant-scoped storageRegister new tools and resources via the barrels in
src/mcp-server/*/definitions/index.tsWrap external API calls: validate raw → normalize to domain type → return output schema; never fabricate missing fields
Contributing
Issues are welcome. Run checks and tests before submitting:
bun run devcheck
bun run testLicense
Apache-2.0 — see LICENSE for details.
This server cannot be deployed
Maintenance
Related MCP Connectors
Archive MCP — wraps the Internet Archive APIs (free, no auth)
Internet Archive (archive.org) item search & metadata MCP.
Search books and authors across Open Library, the Internet Archive open catalog.
One MCP for the Web. Easily search, crawl, navigate, and extract websites without getting blocked.…
Related MCP Servers
- AlicenseAqualityCmaintenanceMCP server for the Internet Archive's Wayback Machine. Search archived snapshots, extract page text from a specific date, track how a site has changed over time, check if broken links are recoverable, and perform research across Internet Archive collections.623 PyPI3MIT
- AlicenseAqualityBmaintenanceMCP server and CLI tool for interacting with the Internet Archive's Wayback Machine, supporting full CDX search, snapshot retrieval, screenshot listing, snapshot comparison, and optional authentication.82,034 npm54Creative Commons Attribution Non Commercial Share Alike 4.0 International
- AlicenseNot gradedqualityBmaintenanceProvides tools to query the Internet Archive, including enumerating historical captures, listing revisions, reading snapshots as clean text, diffing captures, and searching archived items.52 npmMIT
- AlicenseNot gradedqualityBmaintenanceFull-coverage MCP server for Internet Archive, enabling search, metadata lookup, collection browsing, and Wayback Machine snapshot retrieval via 13 tools.1BSD Zero Clause