engram
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@engramremember that I prefer dark mode and recall it next session"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
engram
Long-term memory for LLM agents that gets better the more it is used.
Most agent memory is a pile of notes with a search box. engram tracks how often each
memory is actually recalled: memories that keep proving useful rise in the ranking, and
ones nobody asks for fade and are pruned first.
Hybrid retrieval – BM25 keyword search blended with vector similarity, so a misspelt or differently inflected query still finds the right memory.
Learns from use – every recall reinforces the memories it returns.
Measured – a labelled benchmark reports recall@k and MRR for each search mode.
MCP server – plug the memory into Claude Code or any other Model Context Protocol client.
Zero dependencies for the store, the embeddings and the MCP server: Python standard library only. The optional chat demo uses the Anthropic SDK.
Deduplicates and forgets – near-identical facts strengthen the existing memory, and
prunekeeps only the strongest N.
Benchmark
benchmarks/retrieval.json holds 30 memories and 40 queries, each with one correct answer, in four kinds: exact keywords, different word endings, typos, and paraphrases that share no words with the memory.
python -m engram.evaluate benchmarks/retrieval.jsonQuery kind | Queries | Mode | Recall@1 | Recall@3 | MRR |
exact | 10 | bm25 | 100% | 100% | 1.00 |
exact | 10 | vector | 100% | 100% | 1.00 |
exact | 10 | hybrid | 100% | 100% | 1.00 |
inflected | 10 | bm25 | 80% | 80% | 0.80 |
inflected | 10 | vector | 80% | 80% | 0.80 |
inflected | 10 | hybrid | 100% | 100% | 1.00 |
typo | 10 | bm25 | 20% | 20% | 0.20 |
typo | 10 | vector | 60% | 60% | 0.60 |
typo | 10 | hybrid | 70% | 70% | 0.70 |
paraphrase | 10 | bm25 | 10% | 20% | 0.15 |
paraphrase | 10 | vector | 10% | 20% | 0.16 |
paraphrase | 10 | hybrid | 10% | 30% | 0.18 |
all | 40 | bm25 | 52% | 55% | 0.54 |
all | 40 | vector | 62% | 65% | 0.64 |
all | 40 | hybrid | 70% | 75% | 0.72 |
What this shows:
Hybrid is the best mode overall and is never worse than keywords alone on any query kind.
Paraphrases are not solved. The built-in embeddings compare spelling, not meaning, so "name of the pet" does not find "the user's dog is called Biscuit". Fixing that needs a learned embedding model, which you can plug in (see below).
The benchmark is small and written by the author of the code, so treat the numbers as a comparison between modes, not as an absolute score.
Related MCP server: ebbingflow-mcp
Use the store
from engram import MemoryStore
store = MemoryStore("engram.db")
store.add("The user prefers dark mode", kind="preference")
store.add("The project deploys to Fly.io from main", importance=2.0)
for hit in store.search("where do we deploy?"):
print(hit.score, hit.memory.text)search takes mode="hybrid" (default), "bm25" or "vector".
Plugging in a real embedding model
Pass any function that turns text into a list of numbers:
store = MemoryStore("engram.db", embedder=my_model.encode)Use one embedder per database file: stored vectors are not recomputed when it changes.
Use it from Claude Code (MCP)
claude mcp add engram -- python -m engram.mcp_server --db ~/engram.dbRun that from this folder, or install the package first. The server exposes three tools,
remember, recall and forget, over stdio. It is a from-scratch implementation of the
protocol's tool subset (initialize, ping, tools/list, tools/call) in
engram/mcp_server.py.
Command line
python -m engram add "The project deploys to Fly.io from main"
python -m engram search "deploy" --mode hybrid
python -m engram list
python -m engram prune 100Chat demo
pip install -r requirements.txt
python -m engram chatNeeds ANTHROPIC_API_KEY. Claude decides when to call remember, recall and forget.
Tell it something in one session, quit, start a new one and ask about it.
How ranking works
keyword = BM25(query, memory) / best BM25 score for this query
vector = cosine(query, memory) / best cosine for this query (ignored below 0.25)
relevance = 0.5 × keyword + 0.5 × vector
score = relevance × importance × (1 + ln(1 + uses)) × (0.5 + 0.5 × recency)recency halves every 30 days since the memory was last used, so that factor never drops
below 0.5: old memories get weaker, but a strong match can still surface them.
The default embeddings (engram/embed.py) hash each word's character 3- to 5-grams into a 512-dimensional vector, the idea behind fastText's subword vectors. Words that share most of their letters end up close together.
Tests
python -m unittest discover -s tests30 tests: the store, hybrid search, upgrading an older database file, the MCP server (in process and as a real subprocess over stdio), and a guard that the benchmark numbers above do not regress.
Not verified
The chat demo (engram/agent.py) follows the Anthropic SDK's documented tool-runner pattern but has not been run against the live API.
This server cannot be deployed
Maintenance
Related MCP Connectors
Persistent memory, hybrid search and a goal graph for AI agents, over stdio or remote HTTP.
Persistent memory for AI agents across Claude, ChatGPT and any MCP client.
- mcpOAuthai.butlerbrain
Persistent memory for AI assistants. Save once; recall from Claude, ChatGPT, or any MCP client.
Cross-tool persistent memory and context for AI assistants over MCP.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceLocal-first memory daemon for AI coding agents that captures session transcripts, distills typed memories (decisions, facts, lessons, commands, todos), and serves them via hybrid search through MCP tools.26 npmMIT
- AlicenseNot gradedqualityCmaintenanceEnables MCP clients like OpenCode or Claude to use EbbingFlow long-term memory across sessions through stdio tools for writing, retrieving, and chatting with remembered context.MIT
- FlicenseNot gradedqualityBmaintenanceEnables MCP clients to persist, recall, update, forget, and search agent memories over stdio, backed by Postgres with pgvector for vector search. Supports cross-session memory validation with per-operation latency and outcome logging.1-
- AlicenseNot gradedqualityBmaintenanceProvides persistent memory for AI coding agents, enabling them to save, search (hybrid vector/BM25/concept graph), list sessions, and forget memories via MCP tools.11 npmApache 2.0