mnemosis-mcp
The Mnemosis MCP Server provides a human-inspired memory layer for AI agents, enabling storage, retrieval, consolidation, self-reflection, planning, reasoning, and continuous learning. Core memory operations allow storing episodic/semantic memories with cues and attributes (remember, remember_turn, remember_many), flexible retrieval (recall, search, search_batch, context_pack), and lifecycle management (update, forget, restore, suppress, tag, export/import). Memory consolidation happens via sleep cycles that deduplicate, promote, prune, and resolve conflicts, with optional calibration of forgetting curves. Spaced repetition and practice are supported through review scheduling (review_due, review_batch, next_interval), practice sessions with testing-effect reinforcement, and learning loops that generate self-test questions and track knowledge. Planning and reasoning tools create ordered step plans from goals (plan, replan, plan_tracker), support mental rehearsal, score plan quality, and assemble premise packs for math, comparison, and transitive tasks (reason, numeric_reasoning, physics_simulate). Prospective memory handles future intentions with deadlines, detects clashes, and prioritizes an action queue. Analysis and visualization includes metacognitive gap/contradiction checks, memory health scores, timeline reports, schema clustering, network graphs, analogy bridges, interference reports, and retrieval quality metrics. Utilities offer project briefs, effort estimates, lessons learned, working-set management, deep memory inspection, and batch tagging/import/export. Together, these capabilities transform the server into a dynamic, self-aware cognitive system that mimics human memory processes for robust AI agent performance.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@mnemosis-mcpremember that I prefer Chinese for technical discussions"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Mnemosis
把 AI 的记忆,从“无限存储 + 搜索”改造成“会记住、会遗忘、会整理、会自我怀疑”的系统。
Mnemosis is a human-inspired memory layer for AI agents. Most "AI memory" systems are just storage with semantic search bolted on: they save everything and recall by similarity. Mnemosis instead treats memory as a lifecycle — remembering, reinforcing, consolidating, forgetting, and reconciling — the way human memory actually works.
🍼 New here? Start with the 奶龙级入门教程——基本功能介绍版 — a teacher-style, zero-background tour of every feature (Chinese).
🗄️ Want the storage layer explained? See 存储功能介绍(技术人员版) (Chinese).
中文说明:README.zh-CN.md · English: README.md
Install
pip install git+https://github.com/liyexiaoyi/Mnemosis.gitZero runtime dependencies (pure Python stdlib + SQLite). No server, no cloud embeddings required — optional embedder hooks only.
PyPI 版(
pip install mnemosis)发布后会在这里同步更新。
SQLite 存储适合单进程/低并发场景;多个 Agent 并发写入时建议串行访问或接入 外部数据库适配层。
Related MCP server: aimemory
Quick start
from mnemosis import MemoryEngine
from mnemosis.types import MemoryKind, SourceRecord, SourceType
engine = MemoryEngine("memory.db") # pass a path for persistence
engine.remember(
"The user prefers Chinese for technical discussions.",
kind=MemoryKind.SEMANTIC,
source=SourceRecord(origin=SourceType.USER),
cues=["user", "language", "preference"],
importance=0.9,
)
for r in engine.recall("what language does the user prefer?", top_k=3):
print(f"[{r.item.kind.value}] {r.score:.2f} {r.item.content}")
check = engine.check("what is the user's favorite movie?")
print("knowledge gaps:", check.gaps or "none")
engine.sleep() # offline consolidation: dedupe, promote, detect contradictions中文快速开始
pip install git+https://github.com/liyexiaoyi/Mnemosis.gitfrom mnemosis import MemoryEngine
from mnemosis.types import MemoryKind, SourceRecord, SourceType
engine = MemoryEngine("memory.db")
user = SourceRecord(origin=SourceType.USER)
engine.remember(
"用户喜欢用中文讨论技术问题。",
kind=MemoryKind.SEMANTIC,
source=user,
cues=["语言", "偏好"],
importance=0.9,
)
engine.remember(
"昨天一起修了 SQLite 锁死的问题。",
kind=MemoryKind.EPISODIC,
source=user,
cues=["SQLite", "锁死"],
)
for r in engine.recall("用户用什么语言聊天?", top_k=3):
print(f"[{r.item.kind.value}] 相关度 {r.score:.2f} {r.item.content}")一分钟完整演示(记住 → 检索 → 新旧矛盾 → 睡眠整合 → 元认知 → 遗忘回收):
pip install git+https://github.com/liyexiaoyi/Mnemosis.git
python examples/demo.py # 仓库内不想安装?可以在线体验:
Google Colab 打开演示笔记本 (部分地区访问不了 Colab,可改用下面的方式)
GitHub Codespaces 一键打开:云端环境, 打开终端执行
python examples/demo.py即可下载
examples/Mnemosis_demo.ipynb后,用本地 Jupyter 或百度 AI Studio 打开
How it works
Mnemosis is a memory layer, not a recorder: your agent decides what to
save (via remember / remember_turn) and what to ask (via recall /
check). Inside, memories are stored, decay, consolidate and self-check the
way human memory does.
flowchart LR
A[Agent / MCP client] -->|remember / remember_turn| M[MCP server]
A -->|recall / check| M
M --> E[MemoryEngine]
E -->|store| DB[(SQLite)]
E -->|keyword + n-gram + optional embeddings| R[Retrieval]
E -->|forgetting curve| F[Forgetting & spaced review]
E -->|sleep consolidation| S[Consolidation: dedupe / link / resolve]
E -->|metacognition| C[Check: gaps / contradictions]
R --> DB
F --> DB
S --> DB原理:Mnemosis 照着人脑记忆机制给 agent 做长期记忆——remember
写入、recall 联想检索、遗忘曲线自动衰减、sleep 睡眠式整理(去重/建联/消矛盾)、
check 知道自己不知道(这里的意思是,比如你问了他一个他没有记住的事情,他就会回答不知道,而不是乱编一个答案)。它不是录音机,谁调用、存什么由你和 agent 决定。
Use with your AI client (MCP)
One-line MCP integration for Claude Desktop, Cursor, Codex and any MCP client:
{
"mcpServers": {
"mnemosis": {
"command": "mnemosis-mcp",
"args": ["--db", "/path/to/memory.db"]
}
}
}Windows users: if the client reports that mnemosis-mcp is not found,
add Python's Scripts directory to PATH, or use an absolute path, e.g.
"command": "C:\\Users\\you\\AppData\\Local\\Programs\\Python\\Python312\\Scripts\\mnemosis-mcp.exe".
Full guide (including Cursor and Codex configs): docs/mcp-quickstart.md.
Remote deployment (HTTP)
The MCP server also speaks Streamable HTTP (POST), so it can run on a VPS or NAS and be reached from other machines by URL:
mnemosis-mcp --transport http --host 0.0.0.0 --port 8000 --db /data/memory.dbPoint an MCP client at http://your-server:8000/ (Claude Desktop, Cursor and
Cherry Studio accept remote MCP URLs). Put it behind a reverse proxy
(Caddy/nginx) with TLS and basic auth before exposing it to the internet.
Or run it in Docker:
docker build -t mnemosis .
docker run -d -p 8000:8000 -v mnemosis-data:/data mnemosisMemory is a single SQLite file in /data, so backups are just one file.
Command line
mnemosis --db memory.db remember "用户喜欢用中文讨论技术问题。" --kind semantic
mnemosis --db memory.db recall "用户喜欢什么语言?"
mnemosis --db memory.db sleep
mnemosis --db memory.db check "用户最喜欢的电影是什么?"
mnemosis mcp --db memory.db # or: mnemosis-mcp --db memory.dbBy default the MCP server hides experimental tools from tools/list
(they stay callable). --expose core shows only the 16 everyday tools
(remember/recall/check/review/stats/...), --expose advanced (default)
adds the rest except experimental, and --expose experimental shows all
100+ tools.
Semantic embeddings (optional)
Without an embedder, recall uses zero-dependency keyword + n-gram matching. For real semantic recall, point the MCP server at an embedding API:
# local Ollama (e.g. nomic-embed-text)
mnemosis-mcp --db memory.db --embedder ollama
# any OpenAI-compatible endpoint (DashScope, OpenAI, ...)
export MNEMOSIS_EMBEDDING_API_KEY=sk-...
mnemosis-mcp --db memory.db --embedder openai --embedding-model text-embedding-v3Vectors are cached next to the DB (memory.db.cache) and indexed in
memory.db.vec, so repeated recalls skip embedding calls.
The in-memory embedding cache is bounded by an estimated memory budget
(default 512 MB, configurable via embed_cache_memory_limit_mb): Python
lists of floats count ~32 bytes per element and compact arrays
(numpy.ndarray/array.array) count by their real nbytes, so the cap
reflects actual process memory. An entry-count floor (100k) remains as a
guard for extreme vector dimensions.
For the strongest retrieval quality, enable a dense embedder (Ollama or any OpenAI-compatible endpoint): on the independent LongMemEval benchmark this mode reaches parity with mem0 on turn recall while keeping ingestion ~13x faster. Lexical-only mode stays available and is several times faster per query (dense re-ranking only embeds the top-64 lexical candidates).
After enough usage, call the calibrate_decay MCP tool to fit the
forgetting curve to your real retrieval history (median survival span ->
per-user decay rate). The fitted rate is persisted in the SQLite database
and automatically reloaded the next time you open it.
To see what the memory actually holds, call memory_map (topics with
counts/retrievability plus a weak/ok/strong histogram), or render a Chinese
chart locally:
python benchmarks/render_memory_map.py --db memory.db --out memory_map.svgAutomatic memory saving
Mnemosis never eavesdrops: whoever owns the agent decides what is worth remembering. The easiest automatic pattern is one tool call per turn:
mnemosis-mcp --db memory.dbThen tell your agent in its system prompt:
After every user/assistant exchange, call
remember_turnwith the raw text of the turn. It splits sentences, extracts cues and stores them. Callrecall(orcheck) before answering when the user references earlier topics.
remember_turn is a single call that stores the whole turn as segmented
memories with automatic cues, so agents do not need to hand-write
remember calls for each fact.
For temporal and multi-session questions ("how much did I spend in total", "what changed since last month"), LongMemEval bad-case analysis shows the biggest accuracy gains come from prompting, not retrieval: ask the agent to build a chronological timeline of the recalled memories, sum amounts when the question asks for a total, and treat the most recently dated fact as authoritative when facts conflict (knowledge updates). Mnemosis returns recalled items with dates and confidence, so the agent can apply exactly those rules.
Batch ingestion (importing large histories)
remember_many imports a batch of memories with the same per-item
semantics as remember (semantic dedupe, cues, term index, association
graph), but commits storage, term rows and links in bulk:
engine.remember_many([
{"content": "用户喜欢喝咖啡。", "kind": MemoryKind.SEMANTIC, "cues": ["用户"]},
{"content": "上周修了空调。", "kind": MemoryKind.EPISODIC},
{"content": "阿丽在2026年3月1日买了笔记本。", "kind": MemoryKind.EPISODIC},
])10,000 memories take about 13 seconds via remember_many vs ~70 seconds in
a single-remember loop. When a vector index + Ollama/OpenAI embedder is
enabled, embedding is also batched (100 texts per call, chunked by size,
with automatic retry on 429/5xx).
If a batch embedding still fails halfway (the memories are stored but some have no vector), lexical recall keeps working and the gap can be repaired in one pass:
rebuilt = engine.rebuild_missing_vectors() # or the MCP tool rebuild_vectorsThreading
MemoryEngine is designed for one agent loop per instance: SQLite access and
the shared intent / suppression state are internally locked, so light
concurrent read/write calls are safe. If you fan one engine out across many
threads, serialize the high-level calls yourself.
Testing
python -m unittest discover -s tests -q # 395 unit tests
python benchmarks/locomo_bench.py --mode keyword # LoCoMo-style long dialoguePerformance (10,000 memories, local machine)
Operation | Before | After |
Keyword recall (zero-hit) | 149 ms | 34 ms |
Fused (keyword + n-gram + RRF) recall | 260 ms | 13 ms |
N-gram recall | ~200 ms | 5.5 ms |
Sleep consolidation | 6.0 s | 0.41 s |
Batch ingestion ( | ~70 s | 13 s |
On the independent LongMemEval benchmark (20 questions, cloud Qwen judging), the dense mode matches mem0 on every retrieval metric (@1 0.5 / @5 0.8 / @10 0.8), scores slightly higher on answer accuracy (0.45 vs 0.40) and ingests 1.4x faster (44s vs 62s per question).
A larger retrieval-only run (50 questions) confirms exact parity on every metric (@1 0.42 / @5 0.78 / @10 0.86 / gold-answer-token hit@5 0.38) with 1.37x faster ingestion (44.6s vs 61.2s per question).
The end-to-end run with cloud Qwen judging (30 questions, mem0 vs dense) again shows identical retrieval (@1 0.433 / @5 0.767 / @10 0.867) and equal answer accuracy (0.433 both), with 2.1x faster ingestion (28.0s vs 58.8s per question).
A second cloud-judged sample (different seed, 20 questions, dense mode) scored 0.6 answer accuracy, so the accuracy range across samples is ~0.43-0.60 (50 combined questions ≈ 0.5), consistent with mem0-tier retrieval and answer quality.
A 2026-08-13 retrieval-only rerun (20 questions, seed 42) confirms exact dense-mode parity with mem0 on every retrieval metric (@1 0.5 / @5 0.8 / @10 0.8 / gold-token hit@5 0.35) while ingesting ~5.8x faster (10.8s vs 62.4s per question).
The sentence-segmented ingestion pattern (remember_turn, one call per
conversation turn) is even stronger on the same 20-question set: turn
recall @1 0.60 / @5 0.85 / @10 0.90, ahead of mem0 on every retrieval
metric while ingesting ~4.8x faster. For long multi-session conversations,
store turns with remember_turn and recall with recall / recall_fused.
Cloud-judged end-to-end (qwen3.7-plus answers + judging, same 20 questions) confirms the pattern: mnemosis seg answer accuracy 0.65 vs mem0 0.45, with retrieval also ahead and ingestion ~5x faster. (The benchmark initially capped answer generation at 300 tokens, truncating the mandated timeline + ANSWER line for seg contexts; raising it to 1200 removed that artifact.)
At a 50k-input / 25k-active real-Chinese store, recall is 26-85ms and sleep consolidation ~2.3s after the large-store scaling work (batched term df, batched REM links, generic-cue skipping).
At 100k input (50k active, real Chinese): batch ingestion is ~6.8 minutes (was >2h before the bulk upsert + generic-cue skip work), hit recall ~4ms, zero-hit recall ~5ms (dual fallback: recent 150 + strongest 50, so old high-importance core facts still surface), sleep ~1.3s after the trusted row fast path (SQLite reads skip per-row validation/normalization).
Independent public benchmark (LongMemEval)
Mnemosis is also validated against an external, non-self-built benchmark:
LongMemEval (Wu et al., ICLR 2025),
which uses ~115k tokens of real chat history per question and heavy
interference. The comparison with the official mem0 package is in
benchmarks/longmemeval_bench.py.
First fetch the official dataset (cached copies are detected and reused):
python benchmarks/fetch_longmemeval.py # oracle + S set (~290 MB)
python benchmarks/fetch_longmemeval.py --all # also the 2.7 GB M setIf Hugging Face is unreachable from your network, download the files from
huggingface.co/datasets/xiaowu0162/longmemeval-cleaned and point the
script at the local folder:
python benchmarks/fetch_longmemeval.py --source-dir /path/to/downloadsThen run a head-to-head against mem0 (requires a Python with mem0
installed, plus Ollama or a cloud embedder/LLM for the dense mode):
python benchmarks/longmemeval_bench.py --data work/longmemeval_s_cleaned.json --questions 20License
MIT. See LICENSE.
Contributing
PRs, issues and new benchmark scenarios are welcome — see CONTRIBUTING.md. Every change should come with a test or a measured benchmark result.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceAn MCP-native, local-first memory server that gives AI agents persistent, structured memory across sessions and tools, enabling them to maintain identity and context without reconfiguration.3MIT
- AlicenseNot gradedqualityDmaintenanceGives AI agents persistent memory with semantic search, automatic extraction, and memory decay, accessible via MCP protocol.12MIT
- AlicenseNot gradedqualityDmaintenanceLocal-first AI memory layer with hybrid retrieval and brain-inspired namespaces. Enables agents to save, search, and manage memories directly via MCP tools.5MIT
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to have persistent, self-managing memory with bi-temporal supersession, timely forgetting, and recall under a limited context window, using MCP protocol.MIT
Related MCP Connectors
Shared, governed long-term memory for AI agents across tools and sessions via MCP and REST.
Your memory, everywhere AI goes. Build knowledge once, access it via MCP anywhere.
Person-owned AI memory that learns, not just stores — portable context for any MCP client.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/liyexiaoyi/Mnemosis'
If you have feedback or need assistance with the MCP directory API, please join our Discord server