Skip to main content
Glama
cyanheads

@cyanheads/internet-archive-mcp-server

by cyanheads

Version License Docker MCP SDK npm TypeScript Bun

Install in Claude Desktop Install in Cursor Install in VS Code

Framework


Overview

The Wayback Machine and Internet Archive library (40M+ items). Find and fetch archived snapshots of any URL, search the library by keyword and metadata, and retrieve item metadata, file manifests, and OCR text from any MCP client. Runs as a stdio process or a local Streamable HTTP server.

Tools

Tool

Description

ia_find_snapshots

Find Wayback Machine snapshots of a URL, by closest timestamp or full capture history

ia_get_snapshot

Fetch archived page content at a specific Wayback timestamp

ia_search_items

Search the IA library (40M+ items) by keyword and metadata filters

ia_get_item

Retrieve full metadata and file manifest for an Archive item

ia_get_text

Retrieve readable OCR text from a text item, with paging

Resources

Resource

Description

ia://item/{identifier}

Metadata snapshot for an Archive item — title, creator, mediatype, description, subjects, collections, date, license, and file count

All resource data is also reachable via ia_get_item.

Related MCP server: MCP Wayback Machine Server

Capability reference

ia_find_snapshots tool

  • closest mode: returns the nearest capture to a given timestamp via the Availability API, falling back to one CDX closest-capture query (preferring a 200 capture, 25 s deadline) when the Availability API has no answer

  • history mode: full capture list via the CDX API; filter by date range (from/to), HTTP status (status_filter), and MIME type

  • Default collapse of timestamp:8 (one capture per day); adjustable to timestamp:N, N=1–14

  • Up to 10,000 records per call (limit, default 100); resume_key pagination for large histories

  • Replay URLs are always https://web.archive.org/web/…

  • Typed errors: missing_timestamp (closest mode without a timestamp, answered before any lookup), no_snapshots (no matches), no_snapshot_available (closest mode, no capture near timestamp from either source), cdx_unavailable (CDX 5xx, 429, or unreadable response, or the closest-mode CDX check did not complete), availability_unavailable (closest mode, Availability API 5xx, 429, or unreadable response); a 429 is answered after one request with a hint to wait


ia_get_snapshot tool

  • Resolves to the nearest available capture when the exact timestamp has no snapshot; exact 14-digit timestamps skip resolution

  • Reports the capture Wayback served — replay_url, resolved_timestamp, and resolved_status follow Wayback's redirect when the requested timestamp is not itself a capture

  • Decodes the page with its declared charset (Content-Type, then <meta>, else UTF-8 when the bytes are valid UTF-8, else Wayback's guessed charset or windows-1252), then removes scripts, styles, comments, and tags and decodes character references, returning readable plain text alongside the replay URL

  • Reads at most the first 4 MiB of a page (a notice says when a page is longer); output capped at IA_MAX_SNAPSHOT_CHARS (default 50,000 characters)

  • Typed errors: no_snapshot_available, content_fetch_failed (Wayback unreachable during lookup or fetch)


ia_search_items tool

  • Solr query syntax plus structured filters: mediatype, collection, creator, language, and date range (date_from/date_to)

  • mediatype takes the ten Internet Archive media types — texts, movies, audio, software, image, data, web, collection, etree, account — case-insensitively, and resolves common near-misses (text, book, books → texts; movie, video, videos → movies; images → image; collections → collection)

  • Sort by relevance, date, or downloads (sort, Solr syntax; default downloads desc)

  • Up to 200 results per page (rows, default 50), 1-indexed page

  • Output carries total_found, page, rows for pagination; an empty page returns a notice rather than an error — naming the mediatype applied when nothing matched, or the last page when page is past the end

  • Typed error invalid_mediatype for any other mediatype, listing the accepted values, answered before any search request


ia_get_item tool

  • Returns title, creator, description, subject, collection, licenseurl, rights, and language when present in upstream metadata; creator, description, subject, collection, and language may be a string or a list

  • files[] is one page of the manifest in upstream order — format, size, md5, and a direct download_url per file; file_count is always the full manifest size

  • max_files (1–500, default 50) and file_offset (default 0) page through large items; when files remain, the response sets truncated and names the next file_offset

  • format keeps one file type (exact match, case-insensitive — e.g. DjVuTXT, Text PDF, VBR MP3) before paging and reports the match count as totalCount; a format with no matches returns an empty page and a notice listing the formats the item has

  • Typed error item_not_found for unknown identifiers


ia_get_text tool

  • max_chars (defaults to IA_MAX_SNAPSHOT_CHARS) and char_offset page through long documents; has_more signals additional text remains

  • Locates the best available text file — DjVuTXT preferred, falls back to plain text; source_file names the file fetched

  • Typed errors: item_not_found, no_text_file, download_forbidden (restricted collections)


ia://item/{identifier} resource

  • Returns application/json — title, creator, mediatype, description, subject, collection, date, licenseurl, rights, language, and file_count

  • identifier comes from ia_search_items results

  • Typed error item_not_found for unknown identifiers

Features

Built on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.

Internet Archive-specific:

  • No credentials required — all four APIs are public

  • Three service layers: WaybackService (Availability + CDX), ArchiveSearchService (Solr), ArchiveMetadataService (Metadata + downloads)

  • CDX collapse-by-day default and configurable limit keep responses tractable for high-capture URLs

  • Identifies via a custom User-Agent on every request as required by IA's terms of use; configurable via IA_USER_AGENT

Agent-friendly output:

  • Pagination context on every list response — total_found, page, rows (search), resume_key (CDX history), and file_count plus the next file_offset (item files) so agents never have to guess whether results are complete

  • Typed error reasons (missing_timestamp, no_snapshots, no_snapshot_available, cdx_unavailable, availability_unavailable, content_fetch_failed, invalid_mediatype, item_not_found, no_text_file, download_forbidden) with recovery hints so callers can retry or explain to users without parsing text

  • Structured file manifests — ia_get_item returns file-level metadata (format, size, URL) and a format filter, so agents can pick the right file without paging through thumbnails

Getting started

No API key required — the Internet Archive's APIs are fully public.

Add the following to your MCP client configuration file:

{
  "mcpServers": {
    "internet-archive-mcp-server": {
      "type": "stdio",
      "command": "bunx",
      "args": ["@cyanheads/internet-archive-mcp-server@latest"],
      "env": {
        "MCP_TRANSPORT_TYPE": "stdio",
        "MCP_LOG_LEVEL": "info"
      }
    }
  }
}

Or with npx (no Bun required):

{
  "mcpServers": {
    "internet-archive-mcp-server": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "@cyanheads/internet-archive-mcp-server@latest"],
      "env": {
        "MCP_TRANSPORT_TYPE": "stdio",
        "MCP_LOG_LEVEL": "info"
      }
    }
  }
}

Or with Docker:

{
  "mcpServers": {
    "internet-archive-mcp-server": {
      "type": "stdio",
      "command": "docker",
      "args": [
        "run", "-i", "--rm",
        "-e", "MCP_TRANSPORT_TYPE=stdio",
        "ghcr.io/cyanheads/internet-archive-mcp-server:latest"
      ]
    }
  }
}

For Streamable HTTP, set the transport and start the server:

MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 bun run start:http
# Server listens at http://localhost:3010/mcp

Prerequisites

  • Bun v1.4.0 or higher (or Node.js v24+).

  • No external accounts or API keys required.

Installation

  1. Clone the repository:

git clone https://github.com/cyanheads/internet-archive-mcp-server.git
  1. Navigate into the directory:

cd internet-archive-mcp-server
  1. Install dependencies:

bun install
  1. Configure environment:

cp .env.example .env
# Optional: edit .env for custom User-Agent, timeouts, etc.

Configuration

All configuration is validated at startup via Zod schemas in src/config/server-config.ts.

Variable

Description

Default

MCP_TRANSPORT_TYPE

Transport: stdio or http

stdio

MCP_HTTP_PORT

HTTP server port

3010

MCP_AUTH_MODE

Auth mode: none, jwt, or oauth

none

MCP_LOG_LEVEL

Log level (debug, info, notice, warning, error)

info

LOGS_DIR

Directory for log files (Node.js only)

<project-root>/logs

STORAGE_PROVIDER_TYPE

Storage backend

in-memory

OTEL_ENABLED

Enable OpenTelemetry instrumentation

false

IA_USER_AGENT

Custom User-Agent for IA API requests

internet-archive-mcp-server/{version} (github.com/cyanheads/internet-archive-mcp-server)

IA_REQUEST_TIMEOUT_MS

HTTP request timeout in milliseconds

30000

IA_MAX_SNAPSHOT_CHARS

Default character cap for ia_get_text responses

50000

See .env.example for the full list of optional overrides.

Running the server

Local development

  • Build and run:

    # One-time build
    bun run rebuild
    
    # Run the built server
    bun run start:stdio
    # or
    bun run start:http
  • Run checks and tests:

    bun run devcheck   # Lint, format, typecheck, security
    bun run test       # Vitest test suite
    bun run lint:mcp   # Validate MCP definitions against spec

Docker

docker build -t internet-archive-mcp-server .
docker run --rm -p 3010:3010 internet-archive-mcp-server

The Dockerfile defaults to HTTP transport, stateless session mode, and logs to /var/log/internet-archive-mcp-server. OpenTelemetry peer dependencies are installed by default — build with --build-arg OTEL_ENABLED=false to omit them.

Project structure

Directory

Purpose

src/index.ts

createApp() entry point — registers tools, resource, and inits services.

src/config

Server-specific environment variable parsing and validation with Zod.

src/mcp-server/tools

Tool definitions (*.tool.ts). Five tools across Wayback and IA library.

src/mcp-server/resources

Resource definitions. ia://item/{identifier} item metadata resource.

src/services/wayback

WaybackService — Availability API + CDX API client, plus archived-page charset decoding and text extraction.

src/services/archive-search

ArchiveSearchService — Solr Advanced Search client.

src/services/archive-metadata

ArchiveMetadataService — Metadata API + file download client.

tests/

Unit and integration tests mirroring src/.

Development guide

See CLAUDE.md for development guidelines and architectural rules. The short version:

  • Handlers throw, framework catches — no try/catch in tool logic

  • Use ctx.log for request-scoped logging, ctx.state for tenant-scoped storage

  • Register new tools and resources via the barrels in src/mcp-server/*/definitions/index.ts

  • Wrap external API calls: validate raw → normalize to domain type → return output schema; never fabricate missing fields

Contributing

Issues are welcome. Run checks and tests before submitting:

bun run devcheck
bun run test

License

Apache-2.0 — see LICENSE for details.

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    C
    maintenance
    MCP server for the Internet Archive's Wayback Machine. Search archived snapshots, extract page text from a specific date, track how a site has changed over time, check if broken links are recoverable, and perform research across Internet Archive collections.
    6
    23 PyPI
    3
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    MCP server and CLI tool for interacting with the Internet Archive's Wayback Machine, supporting full CDX search, snapshot retrieval, screenshot listing, snapshot comparison, and optional authentication.
    8
    2,034 npm
    54
    Creative Commons Attribution Non Commercial Share Alike 4.0 International
  • A
    license
    Not graded
    quality
    B
    maintenance
    Full-coverage MCP server for Internet Archive, enabling search, metadata lookup, collection browsing, and Wayback Machine snapshot retrieval via 13 tools.
    1
    BSD Zero Clause