Skip to main content
Glama

web-archive-mcp

MCP server for persistent web fetch and search archiving. Every web_fetch and web_search result is saved as timestamped JSONL, indexed by fst-indexer, and searchable via unified-history-mcp.

Part of the Palimpsest investigative toolkit.

Why

web_fetch and web_search results normally evaporate when a session ends. Pages change, get deleted, or get memory-holed. This closes that gap — every result is persisted, content-addressed for dedup, and fed into the same search pipeline as your session logs and transcripts. Three months later, a search(domain="all", query="target name") still finds the page that's been 404'd since July.

Related MCP server: scout

Architecture

web_fetch / web_search          playwright-archive-mcp (browser capture)
        │                                │
        ▼                                ▼
  web-archive-store  (shared JSONL write-path + SSRF URL validation)
        │
        ▼
  ~/.local/share/web-archive/*.jsonl
        │                                │
        │                        fst-indexer (Jsonl extractor)
        │                                │
        ▼                                ▼
  returns content              index.fst + manifest.json
                                       │
                                       ▼
                              unified-history-mcp
                              domain: "web-archive"

Tools

Tool

Description

web_fetch(url, timeout, token, preview, redact_html, method, body, content_type)

Fetch a URL, convert to markdown, persist, return

web_fetch supports the standard HTTP verbs via method (GET, HEAD, POST, PUT, PATCH, DELETE, OPTIONS; default GET). Pass body for the request body and content_type for the Content-Type header (e.g. a JSON POST). The method is stamped on each archived entry and included in the response. | web_search(query) | Search the web (DuckDuckGo), persist results | | archive_list(date_from, date_to, max) | List archived entries with metadata | | archive_read(id, max_entries) | Read entries from an archive file | | rebuild | Rebuild FST index for the web-archive domain |

Installation

web-archive-mcp depends on the shared web-archive-store package (archive write-path + SSRF URL validation). Install it first:

git clone https://github.com/palimpsest-labs/web-archive-store
cd web-archive-store
python3 -m venv .venv
source .venv/bin/activate
pip install -e .

cd ..
git clone https://github.com/palimpsest-labs/web-archive-mcp
cd web-archive-mcp
python3 -m venv .venv
source .venv/bin/activate
pip install -e .

Once web-archive-store is published to PyPI this becomes a single pip install -e ..

Browser-driven traffic capture (the playwright_* tools) moved to the separate playwright-archive-mcp server, which records HTTP traffic into the same store.

Integration with unified-history-mcp

Add to your unified-history TOML config:

[domains.web-archive]
dir = "~/.local/share/web-archive"
pattern = "*.jsonl"
extractor = "jsonl"
label = "web-archive entry"
filters = []

Then rebuild: search(domain="web-archive", query="rebuild") or call rebuild directly.

Once indexed, a search(domain="all", query="your search") scans your sessions, transcripts, notifications, and every web page you've ever fetched — in a single query.

Entry format

{
  "type": "fetch",
  "source": "https://example.com/page",
  "title": "Example Page",
  "content": "# Example\n\nMarkdown content...",
  "timestamp": "2026-07-30T21:15:00Z",
  "content_hash": "abc123..."
}

For searches, source holds the query string and type is "search".

Content-addressed dedup prevents storing identical entries. Same source + same content hash = skipped.

License

MIT

Install Server
F
license - not found
A
quality
B
maintenance

Maintenance

Maintainers
Response time
Release cycle
Releases (12mo)
Commit activity

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Servers

  • A
    license
    Not graded
    quality
    F
    maintenance
    A local MCP server that records completed tasks to daily JSONL files and promotes substantial work to a cumulative weekly Markdown worklog, providing persistent, searchable logs of AI-assisted productivity.
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    MCP server that fetches web pages, extracts clean markdown (reducing token count), caches results, and provides searchable reading history.
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    An MCP server that fetches web pages and extracts clean, AI-usable context from them, enabling tools for link discovery, content search, and integrated fetch-and-search operations.
    5
    7
    1
    MIT

View all related MCP servers

Related MCP Connectors

  • Agent-native MCP server over the public saagarpatel.dev corpus. Read-only, stateless.

  • Cloud-hosted MCP server for durable AI memory

  • Official remote MCP server for Archivist AI TTRPG campaign memory: characters, sessions, and more.

View all MCP Connectors

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/palimpsest-labs/web-archive-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server