Skip to main content
Glama

engram

tests

Long-term memory for LLM agents that gets better the more it is used.

Most agent memory is a pile of notes with a search box. engram tracks how often each memory is actually recalled: memories that keep proving useful rise in the ranking, and ones nobody asks for fade and are pruned first.

  • Hybrid retrieval – BM25 keyword search blended with vector similarity, so a misspelt or differently inflected query still finds the right memory.

  • Learns from use – every recall reinforces the memories it returns.

  • Measured – a labelled benchmark reports recall@k and MRR for each search mode.

  • MCP server – plug the memory into Claude Code or any other Model Context Protocol client.

  • Zero dependencies for the store, the embeddings and the MCP server: Python standard library only. The optional chat demo uses the Anthropic SDK.

  • Deduplicates and forgets – near-identical facts strengthen the existing memory, and prune keeps only the strongest N.

Benchmark

benchmarks/retrieval.json holds 30 memories and 40 queries, each with one correct answer, in four kinds: exact keywords, different word endings, typos, and paraphrases that share no words with the memory.

python -m engram.evaluate benchmarks/retrieval.json

Query kind

Queries

Mode

Recall@1

Recall@3

MRR

exact

10

bm25

100%

100%

1.00

exact

10

vector

100%

100%

1.00

exact

10

hybrid

100%

100%

1.00

inflected

10

bm25

80%

80%

0.80

inflected

10

vector

80%

80%

0.80

inflected

10

hybrid

100%

100%

1.00

typo

10

bm25

20%

20%

0.20

typo

10

vector

60%

60%

0.60

typo

10

hybrid

70%

70%

0.70

paraphrase

10

bm25

10%

20%

0.15

paraphrase

10

vector

10%

20%

0.16

paraphrase

10

hybrid

10%

30%

0.18

all

40

bm25

52%

55%

0.54

all

40

vector

62%

65%

0.64

all

40

hybrid

70%

75%

0.72

What this shows:

  • Hybrid is the best mode overall and is never worse than keywords alone on any query kind.

  • Paraphrases are not solved. The built-in embeddings compare spelling, not meaning, so "name of the pet" does not find "the user's dog is called Biscuit". Fixing that needs a learned embedding model, which you can plug in (see below).

  • The benchmark is small and written by the author of the code, so treat the numbers as a comparison between modes, not as an absolute score.

Related MCP server: ebbingflow-mcp

Use the store

from engram import MemoryStore

store = MemoryStore("engram.db")
store.add("The user prefers dark mode", kind="preference")
store.add("The project deploys to Fly.io from main", importance=2.0)

for hit in store.search("where do we deploy?"):
    print(hit.score, hit.memory.text)

search takes mode="hybrid" (default), "bm25" or "vector".

Plugging in a real embedding model

Pass any function that turns text into a list of numbers:

store = MemoryStore("engram.db", embedder=my_model.encode)

Use one embedder per database file: stored vectors are not recomputed when it changes.

Use it from Claude Code (MCP)

claude mcp add engram -- python -m engram.mcp_server --db ~/engram.db

Run that from this folder, or install the package first. The server exposes three tools, remember, recall and forget, over stdio. It is a from-scratch implementation of the protocol's tool subset (initialize, ping, tools/list, tools/call) in engram/mcp_server.py.

Command line

python -m engram add "The project deploys to Fly.io from main"
python -m engram search "deploy" --mode hybrid
python -m engram list
python -m engram prune 100

Chat demo

pip install -r requirements.txt
python -m engram chat

Needs ANTHROPIC_API_KEY. Claude decides when to call remember, recall and forget. Tell it something in one session, quit, start a new one and ask about it.

How ranking works

keyword   = BM25(query, memory) / best BM25 score for this query
vector    = cosine(query, memory) / best cosine for this query     (ignored below 0.25)
relevance = 0.5 × keyword + 0.5 × vector
score     = relevance × importance × (1 + ln(1 + uses)) × (0.5 + 0.5 × recency)

recency halves every 30 days since the memory was last used, so that factor never drops below 0.5: old memories get weaker, but a strong match can still surface them.

The default embeddings (engram/embed.py) hash each word's character 3- to 5-grams into a 512-dimensional vector, the idea behind fastText's subword vectors. Words that share most of their letters end up close together.

Tests

python -m unittest discover -s tests

30 tests: the store, hybrid search, upgrading an older database file, the MCP server (in process and as a real subprocess over stdio), and a guard that the benchmark numbers above do not regress.

Not verified

The chat demo (engram/agent.py) follows the Anthropic SDK's documented tool-runner pattern but has not been run against the live API.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    Local-first memory daemon for AI coding agents that captures session transcripts, distills typed memories (decisions, facts, lessons, commands, todos), and serves them via hybrid search through MCP tools.
    26 npm
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables MCP clients like OpenCode or Claude to use EbbingFlow long-term memory across sessions through stdio tools for writing, retrieving, and chatting with remembered context.
    MIT
  • F
    license
    Not graded
    quality
    B
    maintenance
    Enables MCP clients to persist, recall, update, forget, and search agent memories over stdio, backed by Postgres with pgvector for vector search. Supports cross-session memory validation with per-operation latency and outcome logging.
    1
    -
  • A
    license
    Not graded
    quality
    B
    maintenance
    Provides persistent memory for AI coding agents, enabling them to save, search (hybrid vector/BM25/concept graph), list sessions, and forget memories via MCP tools.
    11 npm
    Apache 2.0