web-archive-mcp
Web Fetching: Fetch and archive any HTTP/HTTPS URL, converting content to markdown with SSRF protection.
Web Search: Perform DuckDuckGo searches and archive results.
Browser Automation: Drive a headless browser with Playwright: record traffic, start interactive sessions, navigate, click, fill forms, extract text/HTML, screenshot, and manage history.
Archive Management: List and read archived entries with metadata and date filtering.
Index Rebuilding: Rebuild full-text search index for integration with unified-history-mcp, enabling cross-domain search.
Persistent Storage: All data saved as timestamped JSONL in
~/.local/share/web-archive/with content-addressed deduplication.
Search the web (DuckDuckGo), persist results, and archive them for later retrieval.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@web-archive-mcpfetch https://en.wikipedia.org/wiki/Palimpsest"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
web-archive-mcp
MCP server for persistent web fetch and search archiving. Every web_fetch and web_search result is saved as timestamped JSONL, indexed by fst-indexer, and searchable via unified-history-mcp.
Part of the Palimpsest investigative toolkit.
Why
web_fetch and web_search results normally evaporate when a session ends. Pages change, get deleted, or get memory-holed. This closes that gap — every result is persisted, content-addressed for dedup, and fed into the same search pipeline as your session logs and transcripts. Three months later, a search(domain="all", query="target name") still finds the page that's been 404'd since July.
Related MCP server: Memory MCP
Architecture
web_fetch / web_search playwright-archive-mcp (browser capture)
│ │
▼ ▼
web-archive-store (shared JSONL write-path + SSRF URL validation)
│
▼
~/.local/share/web-archive/*.jsonl
│ │
│ fst-indexer (Jsonl extractor)
│ │
▼ ▼
returns content index.fst + manifest.json
│
▼
unified-history-mcp
domain: "web-archive"Tools
Tool | Description |
| Fetch a URL, convert to markdown, persist, return |
| Search the web (DuckDuckGo), persist results |
| List archived entries with metadata |
| Read entries from an archive file |
| Rebuild FST index for the web-archive domain |
Installation
web-archive-mcp depends on the shared web-archive-store package (archive write-path + SSRF URL validation). Install it first:
git clone https://github.com/palimpsest-labs/web-archive-store
cd web-archive-store
python3 -m venv .venv
source .venv/bin/activate
pip install -e .
cd ..
git clone https://github.com/palimpsest-labs/web-archive-mcp
cd web-archive-mcp
python3 -m venv .venv
source .venv/bin/activate
pip install -e .Once web-archive-store is published to PyPI this becomes a single pip install -e ..
Browser-driven traffic capture (the
playwright_*tools) moved to the separateplaywright-archive-mcpserver, which records HTTP traffic into the same store.
Integration with unified-history-mcp
Add to your unified-history TOML config:
[domains.web-archive]
dir = "~/.local/share/web-archive"
pattern = "*.jsonl"
extractor = "jsonl"
label = "web-archive entry"
filters = []Then rebuild: search(domain="web-archive", query="rebuild") or call rebuild directly.
Once indexed, a search(domain="all", query="your search") scans your sessions, transcripts, notifications, and every web page you've ever fetched — in a single query.
Entry format
{
"type": "fetch",
"source": "https://example.com/page",
"title": "Example Page",
"content": "# Example\n\nMarkdown content...",
"timestamp": "2026-07-30T21:15:00Z",
"content_hash": "abc123..."
}For searches, source holds the query string and type is "search".
Content-addressed dedup prevents storing identical entries. Same source + same content hash = skipped.
License
MIT
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- FlicenseNot gradedqualityNot gradedmaintenanceProvides cross-session memory for AI agents by maintaining a rolling 24-hour event stream and searchable daily archives to prevent context loss. It enables agents to record, query, and retrieve historical events and decisions through a structured markdown-based workspace.1
- AlicenseAqualityBmaintenanceProvides persistent cross-session memory and full-text search for AI coding assistants, storing project context, decisions, and preferences while enabling searchable access to conversation history via local SQLite.81MIT
- AlicenseNot gradedqualityCmaintenancePersistent activity journal for AI agents - enables logging and querying decisions, changes, errors, and observations across sessions.71MIT
- AlicenseAqualityAmaintenanceProvides persistent, searchable memory for AI agents, enabling them to retain, recall, and reflect on information across conversations.191MIT
Related MCP Connectors
Persistent memory for AI agents. Search, store, and recall across sessions.
Persistent memory for AI agents — verbatim conversations, searchable by meaning.
LLM-ready web search + instant answers + URL-to-clean-text fetch for agents and RAG.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/jmars/web-archive-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server