@cyanheads/internet-archive-mcp-server
Provides tools to search the Wayback Machine and Internet Archive library (40M+ items), fetch archived snapshots, retrieve item metadata and full text, enabling AI agents to access historical web content and digital collections.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@@cyanheads/internet-archive-mcp-serverFind Wayback Machine snapshots of example.com from 2020"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Tools
Five tools covering two Internet Archive pillars — Wayback Machine snapshot discovery and retrieval, and IA library search and content access:
Tool | Description |
| Find Wayback Machine snapshots of a URL. Mode |
| Fetch the archived content of a URL at a specific Wayback timestamp. Strips HTML to readable text and returns the canonical replay URL. |
| Search the IA library (40M+ items). Filter by media type, collection, creator, date range, and language. Sort by relevance, date, or downloads. Returns identifiers, titles, types, and pagination context ( |
| Retrieve full metadata and the file manifest for an Archive item by identifier — title, creator, description, subjects, collections, license, and every file with its format, size, and direct download URL. |
| Retrieve readable OCR text (DjVuTXT or plain-text) from a text item. Length-aware truncation with continuation pointer ( |
ia_find_snapshots
Discover what the Wayback Machine has captured for any URL.
closestmode: single fast lookup via the Availability API — returns the nearest capture to a given timestamphistorymode: full capture list via the CDX API, filterable by date range (from/to), HTTP status (status_filter), and MIME typeDefault collapse of
timestamp:8(one capture per day) keeps responses tractable for popular URLs; adjust with thecollapseparameter (timestamp:N, N=1–14)Resume-key pagination (
resume_key) for stepping through large CDX histories without re-scanning
ia_get_snapshot
Retrieve what a page actually said at a point in time.
Resolves to the nearest available capture when the exact timestamp has no snapshot
Strips Wayback banner injections and extracts readable text — returns clean content alongside the canonical replay URL for browser access
Useful for fact-checking, citation verification, and tracing how content changed over time
ia_search_items
Search across 40M+ Archive items by keyword and metadata filters.
Full-text Solr query syntax plus structured filters:
mediatype(texts, audio, video, software, image),collection,creator,language, and date rangeSort by relevance, date added, or download count
Pagination via
pageandrows; output includestotal_foundand currentpage/rowsso agents can paginate correctly without guessing
ia_get_item
Fetch the complete metadata and file manifest for any Archive item.
Returns structured fields:
title,creator,description,subjects,collections,date,license, and morefiles[]includes every file in the item with itsformat,size, and direct download URL — the primary way to act on a search resultmetadataresponse{}on unknown identifier → typeditem_not_founderror
ia_get_text
Read the OCR text of public-domain books, documents, and transcripts.
Locates the best available text file in the item's manifest (DjVuTXT preferred, falls back to plain text)
max_charsandchar_offsetenable efficient paging through long documents without re-fetchingSurfaces
download_forbidden(HTTP 403) as a typed error for restricted collections rather than failing silently
Related MCP server: Wayback Machine MCP Server
Resource
Type | Name | Description |
Resource |
| Metadata snapshot for an Archive item — title, creator, mediatype, description, subjects, collections, date, license, and file count. Stable URIs for injectable context. |
All resource data is also reachable via ia_get_item. The resource provides a stable, injectable URI for referencing a specific item across workflows.
Features
Built on @cyanheads/mcp-ts-core:
Declarative tool, resource, and prompt definitions — single file per primitive, framework handles registration and validation
Unified error handling — handlers throw, framework catches, classifies, and formats
Pluggable auth:
none,jwt,oauthSwappable storage backends:
in-memory,filesystem,Supabase,Cloudflare KV/R2/D1Structured logging with optional OpenTelemetry tracing
STDIO and Streamable HTTP transports
Internet Archive-specific:
No credentials required — all four APIs are public
Three service layers:
WaybackService(Availability + CDX),ArchiveSearchService(Solr),ArchiveMetadataService(Metadata + downloads)CDX collapse-by-day default and configurable
limitkeep responses tractable for high-capture URLsIdentifies User-Agent on every request as required by IA's terms; configurable via
IA_USER_AGENT
Agent-friendly output:
Pagination context on every list response —
total_found,page,rows(search) andresume_key(CDX history) so agents never have to guess whether results are completeTyped error reasons (
no_snapshots,no_snapshot_available,item_not_found,no_text_file,download_forbidden) with recovery hints so callers can retry or explain to users without parsing textStructured file manifests — every
ia_get_itemresponse includes file-level metadata (format, size, URL) enabling agents to select the right file without a follow-up call
Getting started
No API key required — the Internet Archive's APIs are fully public.
Add the following to your MCP client configuration file:
{
"mcpServers": {
"internet-archive-mcp-server": {
"type": "stdio",
"command": "bunx",
"args": ["@cyanheads/internet-archive-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info"
}
}
}
}Or with npx (no Bun required):
{
"mcpServers": {
"internet-archive-mcp-server": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@cyanheads/internet-archive-mcp-server@latest"],
"env": {
"MCP_TRANSPORT_TYPE": "stdio",
"MCP_LOG_LEVEL": "info"
}
}
}
}Or with Docker:
{
"mcpServers": {
"internet-archive-mcp-server": {
"type": "stdio",
"command": "docker",
"args": [
"run", "-i", "--rm",
"-e", "MCP_TRANSPORT_TYPE=stdio",
"ghcr.io/cyanheads/internet-archive-mcp-server:latest"
]
}
}
}For Streamable HTTP, set the transport and start the server:
MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 bun run start:http
# Server listens at http://localhost:3010/mcpPrerequisites
Bun v1.3.2 or higher (or Node.js v24+).
No external accounts or API keys required.
Installation
Clone the repository:
git clone https://github.com/cyanheads/internet-archive-mcp-server.gitNavigate into the directory:
cd internet-archive-mcp-serverInstall dependencies:
bun installConfigure environment:
cp .env.example .env
# Optional: edit .env for custom User-Agent, timeouts, etc.Configuration
All configuration is validated at startup via Zod schemas in src/config/server-config.ts.
Variable | Description | Default |
| Transport: |
|
| HTTP server port |
|
| Auth mode: |
|
| Log level ( |
|
| Directory for log files (Node.js only) |
|
| Storage backend |
|
| Enable OpenTelemetry instrumentation |
|
| Custom User-Agent for IA API requests |
|
| HTTP request timeout in milliseconds |
|
| Default character cap for |
|
See .env.example for the full list of optional overrides.
Running the server
Local development
Build and run:
# One-time build bun run rebuild # Run the built server bun run start:stdio # or bun run start:httpRun checks and tests:
bun run devcheck # Lint, format, typecheck, security bun run test # Vitest test suite bun run lint:mcp # Validate MCP definitions against spec
Docker
docker build -t internet-archive-mcp-server .
docker run --rm -p 3010:3010 internet-archive-mcp-serverThe Dockerfile defaults to HTTP transport, stateless session mode, and logs to /var/log/internet-archive-mcp-server. OpenTelemetry peer dependencies are installed by default — build with --build-arg OTEL_ENABLED=false to omit them.
Project structure
Directory | Purpose |
|
|
| Server-specific environment variable parsing and validation with Zod. |
| Tool definitions ( |
| Resource definitions. |
|
|
|
|
|
|
| Unit and integration tests mirroring |
Development guide
See CLAUDE.md for development guidelines and architectural rules. The short version:
Handlers throw, framework catches — no
try/catchin tool logicUse
ctx.logfor request-scoped logging,ctx.statefor tenant-scoped storageRegister new tools and resources via the barrels in
src/mcp-server/*/definitions/index.tsWrap external API calls: validate raw → normalize to domain type → return output schema; never fabricate missing fields
Contributing
Issues and pull requests are welcome. Run checks and tests before submitting:
bun run devcheck
bun run testLicense
Apache-2.0 — see LICENSE for details.
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- Flicense-qualityCmaintenanceBridge the gap between your web crawl and AI language models. With mcp-server-webcrawl, your AI client filters and analyzes web content under your direction or autonomously, extracting insights from your web content. Supports WARC, wget, InterroBot, Katana, and SiteOne crawlers.Last updated44Python
- AlicenseBqualityDmaintenanceProvides access to the Internet Archive Wayback Machine to list snapshots, fetch archived web pages, and search archive.org items. Enables retrieval of historical website content and metadata through natural language queries.Last updated1234MIT
- AlicenseAqualityAmaintenanceMCP server for the Internet Archive's Wayback Machine. Search archived snapshots, extract page text from a specific date, track how a site has changed over time, check if broken links are recoverable, and perform research across Internet Archive collections.Last updated63MIT
- AlicenseAqualityBmaintenanceProvides tools to archive URLs, retrieve clean readable text from Wayback Machine snapshots, list snapshots, search Internet Archive items, and compare snapshots, designed to avoid context window blowup by returning stripped text.Last updated64MIT
Related MCP Connectors
Internet Archive (archive.org) item search & metadata MCP.
Search and fetch Wikidata entities, execute SPARQL queries, and resolve external identifiers.
Search MusicBrainz artists, releases, works, labels; resolve ISRC/ISWC/barcode; fetch cover art.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/cyanheads/internet-archive-mcp-server'
If you have feedback or need assistance with the MCP directory API, please join our Discord server