Muninn
Supports GenAI semantic convention tracing for monitoring, auditing, and observability of memory retrieval and storage operations.
Enables the discovery and ingestion of legacy memory data and prior assistant session history stored in SQLite databases.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Muninnrecall our discussion about the database schema changes from last week"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Muninn
"Muninn flies each day over the world to bring Odin knowledge of what happens." โ Prose Edda
Local-first persistent memory infrastructure for coding agents and MCP-compatible tools.
Muninn provides deterministic, explainable memory retrieval with robust transport behavior and production-grade operational controls. Designed for long-running development workflows where continuity, auditability, and measurable quality matter โ across sessions, across assistants, and across projects.
๐ฉ Status
Current Version: v3.24.0 (Phase 26 COMPLETE) Stability: Production Beta Test Suite: 1422+ passing, 0 failing
What's New in v3.24.0
Cognitive Architecture (CoALA): Integration of a proactive reasoning loop bridging memory with active decision-making.
Knowledge Distillation: Background synthesis of episodic memories into structured semantic manuals for long-term wisdom.
Epistemic Foraging: Active inference-driven search to resolve ambiguities and fill information gaps autonomously.
Omission Filtering: Automated detection of missing context required for successful task execution.
Elo-Rated SNIPS Governance: Dynamic memory retention system mapping retrieval success to Elo ratings for usage-driven decay.
Previous Milestones
Version | Phase | Key Feature |
v3.24.0 | 26 | Cognitive Architecture Complete |
v3.23.0 | 23 | Elo-Rated SNIPS Governance |
v3.22.0 | 22 | Temporal Knowledge Graph |
v3.19.0 | 20 | Multimodal Hive Mind Operations |
v3.18.3 | 19 | Bulk legacy import, NLI conflict detection, uncapped discovery |
v3.18.1 | 19 | Scout synthesis, hunt mode |
Related MCP server: @contextable/mcp
๐ Features
Core Memory Engine
Local-First: Zero cloud dependency โ all data stays on your machine
Multimodal: Native support for Text, Image, Audio, Video, and Sensor data
5-Signal Hybrid Retrieval: Dense vector ยท BM25 lexical ยท Graph traversal ยท Temporal relevance ยท Goal relevance
Explainable Recall Traces: Per-signal score attribution on every search result
Bi-Temporal Reasoning: Support for "Valid Time" vs "Transaction Time" via Temporal Knowledge Graph
Project Isolation:
scope="project"memories never cross repo boundaries;scope="global"memories are always availableCross-Session Continuity: Memories survive session ends, assistant switches, and tool restarts
Bi-Temporal Records:
created_at(real-world event time) vsingested_at(system intake time)
Memory Lifecycle
Elo-Rated Governance: Dynamic retention driven by retrieval feedback (SNIPS) and usage statistics
Consolidation Daemon: Background process for decay, deduplication, promotion, and shadowing โ inspired by sleep consolidation
Zero-Trust Ingestion: Isolated subprocess parsing for PDF/DOCX to neutralize document-based exploits
ColBERT Multi-Vector: Native Qdrant multi-vector storage for MaxSim scoring
NL Temporal Query Expansion: Natural-language time phrases ("last week", "before the refactor") parsed into structured time ranges
Goal Compass: Retrieval signal for project objectives and constraint drift
NLI Conflict Detection: Transformer-based contradiction detection (
cross-encoder/nli-deberta-v3-small) for memory integrityBulk Legacy Import: One-click ingestion of all discovered legacy sources (batched, error-isolated) via dashboard or API
Operational Controls
MCP Transport Hardening: Framed + line JSON-RPC, timeout-window guardrails, protocol negotiation
Runtime Profile Control:
get_model_profiles/set_model_profilesfor dynamic model routingProfile Audit Log: Immutable event ledger for profile policy mutations
Browser Control Center: Web UI for search, ingestion, consolidation, and admin at
http://localhost:42069OpenTelemetry: GenAI semantic convention tracing (feature-gated via
MUNINN_OTEL_ENABLED)
Multi-Assistant Interop
Handoff Bundles: Export/import memory checkpoints with checksum verification and idempotent replay
Legacy Migration: Discover and import memories from prior assistant sessions (JSONL chat history, SQLite state) โ uncapped provider limits
Bulk Import:
POST /ingest/legacy/import-allingests all discovered sources in batches of 50 with per-batch error isolationHive Mind Federation: Push-based low-latency memory synchronization across assistant runtimes
MCP 2025-11 Compliant: Full protocol negotiation, lifecycle gating, schema annotations
Quick Start
git clone https://github.com/wjohns989/Muninn.git
cd Muninn
pip install -e .Set the auth token (shared between server and MCP wrapper):
# Windows (persists across sessions)
setx MUNINN_AUTH_TOKEN "your-token-here"
# Linux/macOS
export MUNINN_AUTH_TOKEN="your-token-here"Start the backend:
python server.pyVerify it's running:
curl http://localhost:42069/health
# {"status":"ok","memory_count":0,...,"backend":"muninn-native"}Runtime Modes
Mode | Command | Description |
Muninn MCP | shared | Streamable HTTP MCP on the one machine-wide server |
Huginn Standalone |
| Browser-first UX for direct ingestion/search/admin |
REST API |
| FastAPI backend at |
Packaged App |
| PyInstaller executable (Huginn Control Center) |
All modes use the same memory engine and data directory.
MCP Client Configuration
Per-client setup (ChatGPT desktop app with Codex, Claude Desktop and Claude
Code, Gemini CLI, Cursor, VS Code, LM Studio, Open WebUI, web ChatGPT): see
docs/CLIENTS.md. All clients share one store; the server
instructions teach every agent to load get_project_context at the start and to
pass work on with create_handoff / resume_handoff.
Clients with tool limits or small models can load a smaller profile with
?toolset=core (or readonly, chatgpt) on the URL.
The preferred machine-wide topology is one verified server.py process and
HTTP clients connected to http://127.0.0.1:42069/mcp. Clients must not
auto-start private stdio copies when the shared endpoint is configured.
Generic Streamable HTTP client configuration:
{
"mcpServers": {
"muninn": {
"type": "http",
"url": "http://127.0.0.1:42069/mcp",
"headers": {
"Authorization": "Bearer ${MUNINN_AUTH_TOKEN}"
}
}
}
}The legacy stdio wrapper remains available for clients without Streamable HTTP support, but it connects to the existing backend and is not a second store owner.
The endpoint is dual-era: clients on MCP 2026-07-28 send stateless requests
(protocol version, client info and capabilities in params._meta, mirrored in the
MCP-Protocol-Version, Mcp-Method and Mcp-Name headers) and can call
server/discover; clients on 2025-11-25 and earlier keep using initialize and
Mcp-Session-Id. Stateless clients that want session inhibition pass a
session_id argument to search_memory.
Legacy wrapper registration:
claude mcp add -s user muninn \
-e MUNINN_AUTH_TOKEN="your-token-here" \
-- python /absolute/path/to/mcp_wrapper.pyGeneric MCP client (claude_desktop_config.json or equivalent):
{
"mcpServers": {
"muninn": {
"command": "python",
"args": ["/absolute/path/to/mcp_wrapper.py"],
"env": {
"MUNINN_AUTH_TOKEN": "your-token-here"
}
}
}
}Important: Both
server.pyandmcp_wrapper.pymust share the sameMUNINN_AUTH_TOKEN. If either process generates a random token (when the env var is unset), all MCP tool calls fail with 401.
MCP Tools
Tool | Description |
| Session-start briefing: goal, open handoffs, project rules, recent memories by agent, global preferences |
| Leave work for another agent: summary, next steps, decisions, open questions, files, branch |
| Claim the newest open handoff for a project (or a given id) |
| Mark a resumed handoff done or cancelled, or release it |
| Re-read an imported conversation in order, or list a project's threads |
| Import local Claude/Codex/Gemini history and chat exports as memories (dry run by default) |
| Store a memory with optional |
| Store a local image plus a searchable description and optional memory links |
| Hybrid 5-signal search with |
| Paginated memory listing with filters |
| Update content or metadata of an existing memory |
| Remove a memory by ID |
| Set the current project's objective and constraints |
| Retrieve the active project goal |
| Store a project-scoped rule ( |
| Get active model routing profiles |
| Update model routing profiles |
| Audit log for profile policy changes |
| Export a memory handoff bundle |
| Import a handoff bundle (idempotent) |
| Ingest files/folders into memory |
| Find prior assistant session files for migration |
| Import discovered legacy memories |
| Submit outcome signal for adaptive calibration |
| Agentic multi-hop search with a synthesized summary |
| Rewrite a wrong memory from a user correction |
| List missing details (paths, credentials, hosts) a task needs |
| Follow graph links when a search is ambiguous |
| Condense clusters of episodic memories into semantic notes |
| ChatGPT connector tools ( |
The full list (43 tools) is returned by tools/list; federation, periodic
ingestion, model-profile and mimir_relay tools are omitted above.
Python SDK
from muninn import Memory
# Sync client
client = Memory(base_url="http://127.0.0.1:42069", auth_token="your-token-here")
client.add(
content="Always use typed Pydantic models for API payloads",
metadata={"project": "muninn", "scope": "project"}
)
results = client.search("API payload patterns", limit=5)
for r in results:
print(r.content, r.recall_trace)Async client:
from muninn import AsyncMemory
async def main():
async with AsyncMemory(base_url="http://127.0.0.1:42069", auth_token="your-token-here") as client:
await client.add(content="...", metadata={})
results = await client.search("...", limit=5)REST API
Method | Path | Description |
|
| Counts plus content-free resource, cache, and MCP transport utilization |
|
| Add a memory (supports |
|
| Copy an image into managed local storage and add its description |
|
| Authenticated access to a managed image |
|
| Hybrid search (supports |
|
| MCP Streamable HTTP transport |
|
| Paginated memory listing |
|
| Update a memory |
|
| Delete a memory |
|
| Restore a memory that consolidation archived (merge, decay, temporal shadow) |
|
| Rebuild vectors/BM25 from metadata (dry run by default) |
|
| Import exported memories, keeping original timestamps (dry run by default) |
|
| Ingest files/folders |
|
| Discover legacy session files |
|
| Import selected legacy memories |
|
| Discover and import ALL legacy sources (batched) |
|
| Legacy discovery scheduler status |
|
| Paginated cached catalog of discovered sources |
|
| Get model routing profiles |
|
| Set model routing profiles |
|
| Profile audit log |
|
| Get user profile |
|
| Update user profile |
|
| Export handoff bundle |
|
| Import handoff bundle |
|
| Submit retrieval feedback |
|
| Get project goal |
|
| Set project goal |
Auth: Authorization: Bearer <MUNINN_AUTH_TOKEN> required on all non-health endpoints.
Configuration
Key environment variables:
Variable | Default | Description |
| random | Shared secret between server and MCP wrapper |
|
| Backend URL for MCP wrapper |
| off |
|
| off |
|
|
| Default model routing profile |
| off |
|
|
| OTLP HTTP endpoint for trace export |
| off |
|
| off |
|
| off |
|
| - | Comma-separated list of peer base URLs |
| off |
|
| off |
|
| off |
|
|
| Memories visited per phase per cycle; a persisted cursor pages through the whole store |
|
| Self-supervised importance: |
|
| Window in which a memory must be re-retrieved by a new session to count as needed |
|
| Predictions recorded per consolidation cycle for later self-labelling |
|
| Resolved outcomes required before the learned model may take over |
| on | Demote memories already returned in the same agent session (requires |
|
| Positions a repeated memory moves down in the final ranked pool |
|
| How long a returned memory stays inhibited within a session |
| on |
|
|
| FastEmbed cross-encoder model; set the prior |
|
| Load NLI integrity resources only for a consolidation cycle; |
|
| Hard bound for adaptive retrieval-feedback cache entries |
|
| Process-worker bound for multi-source ingestion ( |
|
| Maximum source image size copied into managed storage |
|
| Streamable HTTP session capacity |
|
| Idle Streamable HTTP session expiry |
|
| JSON-RPC batch bound |
|
| Streamable HTTP request-body bound |
|
| Process-wide Streamable HTTP dispatch bound |
|
| Per-session Streamable HTTP dispatch bound |
|
| Legacy SSE session capacity |
|
| Per-session legacy SSE response queue bound |
|
| Per-session legacy SSE dispatch-task bound |
| client name | Agent label recorded on memories and handoffs for a stdio client (HTTP clients use |
| git repo | Project for a stdio client started outside a repository (e.g. by Claude Desktop) |
| on | Keep a private copy of Claude Code/Desktop, Codex and Gemini CLI transcripts (the apps delete theirs); see |
|
| How often new conversation history is copied (and, after the first import, imported) |
| after first import |
|
| - | Enables |
| auto |
|
|
| Primary model for thread analysis; falls back to DeepSeek V4 Flash, then Gemini 3.5 Flash-Lite (all zero data retention), including when a model refuses. A |
|
| Conversation per analysis call; larger threads are split and merged |
| off |
|
| - | Extra home folders to scan for app history (e.g. the Windows home from WSL) |
|
| Tool profile for stdio clients: |
| - | Extra browser origins allowed besides localhost (comma-separated; |
config.template.yaml contains conservative, relative-path defaults. Keep real
tokens and machine-specific data paths in private environment/configuration files.
Importing your existing AI conversations
python -m muninn.cli history import (dry run, then --apply) turns the
conversations already on this machine (Claude Code and Claude Desktop, Codex
CLI and the ChatGPT desktop app, Gemini CLI, and ChatGPT/Claude data exports)
into memories. Each turn is filed under its original project, directory, branch,
agent and time, so it can be searched and re-read in order with get_thread.
The server also keeps a vault copy of these transcripts so nothing the apps
clean up is lost. python -m muninn.cli hooks install --apply adds Claude Code and
Codex hooks that brief every new session and capture transcripts before
compaction, and history analyze (OpenRouter with zero data retention, or local
Ollama) extracts decisions, preferences, fixes and open items per thread.
Details: docs/CLIENTS.md.
Upgrading and migrating memories
metadata.db is the source of truth; vectors and the keyword index are derived from it.
Every command below talks to the running server and is a dry run unless --apply is given.
# Rebuild vectors and BM25 from metadata.db (after an embedding-model change add
# --recreate-vectors; also use after restoring metadata.db into a fresh install)
python -m muninn.cli reindex --apply
# Import memories exported from another system, including the pre-3.0 Mem0-based
# Muninn: JSONL, a JSON array, or a Mem0 GET /memories response. Original
# timestamps are kept, exact duplicates skipped, and the original user id is
# stored as metadata.legacy_user_id.
python -m muninn.cli import export.json --source mem0
python -m muninn.cli import export.json --source mem0 --apply
# Rows of an older Muninn metadata.db `memories` table, exported as JSONL, keep
# their project, metadata (JSON text is parsed), memory type and archived state.
python -m muninn.cli import old-muninn.jsonl --source muninn-legacy/health reports legacy_stores (booleans only) when an older ~/.muninn/data or Mem0
store exists on the machine. Back up the data directory before any --apply.
Reproducible memory profiling
The benchmark harness refuses port 42069, strips credentials, disables
external model calls, and creates every store beneath a temporary directory:
python -m eval.memory_profile_benchmark \
--output eval/reports/memory/local-report.json \
--soak-seconds 1800 \
--idle-soak-seconds 1800Generated reports are intentionally ignored because they can contain local temporary paths. Publish only reviewed aggregate measurements.
Evaluation & Quality Gates
Muninn includes an evaluation toolchain for measurable quality enforcement:
# Run full benchmark dev-cycle
python -m eval.ollama_local_benchmark dev-cycle
# Check phase hygiene gates
python -m eval.phase_hygiene
# Emit SOTA+ signed verdict artifact
python -m eval.ollama_local_benchmark sota-verdict \
--longmemeval-report path/to/lme_report.json \
--min-longmemeval-ndcg 0.60 \
--min-longmemeval-recall 0.65 \
--signing-key "$SOTA_SIGNING_KEY"
# Run LongMemEval adapter selftest (no server needed)
python eval/longmemeval_adapter.py --selftest
# Run StructMemEval adapter selftest (no server needed)
python eval/structmemeval_adapter.py --selftest
# Run StructMemEval against a live server
python eval/structmemeval_adapter.py \
--dataset path/to/structmemeval.jsonl \
--server-url http://localhost:42069 \
--auth-token "$MUNINN_AUTH_TOKEN"Metrics tracked: nDCG@k, Recall@k, MRR@k, Exact Match, token-F1, p50/p95 latency, significance testing (Bonferroni/BH correction), effect-size analysis.
The sota-verdict command emits a signed JSON artifact with commit_sha, SHA256 file hashes, and HMAC-SHA256 promotion_signature โ enabling auditable, commit-pinned SOTA+ evidence.
Data & Security
Default data dir:
~/.local/share/AntigravityLabs/muninn/(Linux/macOS) ยท%LOCALAPPDATA%\AntigravityLabs\muninn\(Windows)Storage: SQLite (metadata) + Qdrant (vectors) + KuzuDB (memory chains graph)
No cloud dependency: All data local by default
Auth: when
MUNINN_AUTH_TOKENorMUNINN_API_KEYis set, every API and MCP call needs it as a Bearer token; without one, only local callers are expectedBrowser origins: requests from web pages other than
localhostare rejected (blocks cross-site access and DNS rebinding); extend withMUNINN_ALLOWED_ORIGINSNamespace isolation:
user_id+namespace+projectboundaries enforced at every retrieval layer
Documentation Index
Document | Description |
| Active development phases and roadmap |
| Current plan and status |
| Historical handoffs and remediation reports (including |
| System architecture deep-dive |
| Full feature roadmap (v3.1โv3.3+) |
| How to resume development across sessions |
| Python SDK reference |
| Connecting Claude, ChatGPT, Codex, Gemini, Cursor, VS Code and local-model clients |
| Ingestion pipeline internals |
| OpenTelemetry integration guide |
| Gap analysis against SOTA memory systems |
Licensing
Code: Apache License 2.0 (
LICENSE)Third-party dependency licenses remain with their respective owners
Attribution: See
NOTICE
This server cannot be deployed
Maintenance
Related MCP Connectors
Persistent memory and knowledge graphs for AI agents. Hybrid search, context checkpoints, and more.
Persistent memory for AI agents. Semantic search, memory graph, W3C DID identity.
Persistent memory and knowledge management for AI agents with semantic search and 50+ tools.
Persistent memory for AI agents. EU-hosted, privacy-first, hybrid recall, contradiction detection.
Related MCP Servers
- -licenseNot gradedqualityNot gradedmaintenancePersistent development memory server that automatically captures and organizes development context, code changes, and user interactions across projects.3-
- AlicenseNot gradedqualityDmaintenanceA persistent AI memory server that enables storage and retrieval of context and project artifacts across conversations. It features full-text search, version history, and automatic content chunking using local SQLite or hosted cloud storage.11 npmApache 2.0
- AlicenseNot gradedqualityCmaintenanceA persistent memory server that stores and retrieves atomic coding insights like architectural decisions and debugging patterns for AI agents. It enables agents to maintain institutional knowledge across sessions using semantic search and local SQLite storage.9 npm11MIT
- AlicenseAqualityCmaintenanceA persistent long-term memory server for AI assistants that enables storing and recalling solutions, facts, and decisions with intelligent confidence tracking and relationship mapping. It allows developers to build a cross-platform knowledge base that integrates seamlessly with IDEs and CLI agents.1724 PyPI2MIT