Skip to main content
Glama
README.md
# mcp-archive

> 6 Wayback Machine + Internet Archive tools for Claude. **The one that doesn't blow your context window.** Free, MIT.

Existing Wayback MCPs either dump raw HTML (blowing up the LLM context window) or are abandoned forks. **`mcp-archive` returns clean readable text** — stripped HTML, scripts, nav, Wayback toolbar — so you can actually fit archived pages in a Claude conversation.

---

## Install

```json
{
  "mcpServers": {
    "archive": {
      "command": "npx",
      "args": ["-y", "@dpm-tools/mcp-archive"]
    }
  }
}
```

Requires Node.js 22.18+.

---

## Tools (6)

- **`archive_url`** — Push a URL to Wayback right now. Returns snapshot URL.
- **`get_archived`** — Get nearest Wayback snapshot for a URL/date.
- **`list_snapshots`** — List Wayback snapshots in a date range (CDX API).
- **`fetch_archived_text`** ⭐ — Fetches archived page and returns CLEAN READABLE TEXT (strips HTML, scripts, nav, toolbar). The key differentiator.
- **`search_archive_items`** — Search Internet Archive items (books, audio, video, software).
- **`compare_snapshots`** — Compare two snapshots of the same URL by extracted text.

---

## Example prompts

- *"Archive this URL right now: https://example.com/important-page"*
- *"What did news.ycombinator.com look like on 2020-01-01?"*
- *"List all archived versions of openai.com from 2023."*
- *"Fetch the readable text of https://example.com archived in 2024."*
- *"Search the Internet Archive for 'public domain javascript books'."*
- *"Compare https://example.com between 2020 and 2024 — what changed?"*

See [`examples/prompts.md`](./examples/prompts.md) for more.

---

## Why this exists

Researchers, journalists, and link-rot fighters need Wayback access. Existing MCPs in this space:

- 5+ personal forks, mostly abandoned
- Most return raw HTML — instantly blows up Claude's context window
- None offer text extraction + truncation

**`mcp-archive` solves it:** regex-only HTML→text extraction (no jsdom/cheerio deps), Wayback toolbar stripping, configurable `max_chars`. The `id_` flag is used on snapshot URLs to fetch raw content without the toolbar at the source.

Zero API keys. Custom User-Agent. 200ms inter-call sleep on rate-limited operations to respect archive.org TOS.

---

## API sources (all free, all keyless)

- **Wayback Machine availability API** (`archive.org/wayback/available`)
- **Wayback CDX search** (`web.archive.org/cdx/search/cdx`)
- **Wayback save API** (`web.archive.org/save/`)
- **Internet Archive advanced search** (`archive.org/advancedsearch.php`)

---

## Sister servers from dpm

- **`@dpm-tools/mcp-devkit`** — 15 dev utilities (UUID, hash, JWT, JSON, cron, color)
- **`@dpm-tools/mcp-public-data`** — sun, moon, holidays, geocoding (keyless)
- **`@dpm-tools/mcp-diff`** — diff, patch, 3-way merge

---

## License

MIT © dpm (digital product mill)

TDQS

A3.7/5.0

Scored across 6 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: submitting URLs, retrieving single snapshots, listing multiple snapshots, comparing snapshots, fetching text from a snapshot, and searching non-Wayback items. No overlap.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern (e.g., archive_url, list_snapshots, fetch_archived_text), making the API predictable.

Tool Count5/5

6 tools cover the core operations of web archiving and searching the Internet Archive catalog, neither too few nor too many.

Completeness5/5

The toolset provides a complete lifecycle for Wayback snapshots (submit, retrieve, list, compare, get text) and adds search for other archive types, with no obvious gaps.

Maintenance

ActivityInactive
ResponsivenessNo issues