Engram
Allows ChatGPT and Codex CLI to write memories to and recall from Engram, including through tools like remember, recall, note, teach, expand, timeline, and forget.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Engramwhat do you remember about my job search?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Every assistant starts from zero. Tell Claude you're job hunting, and ChatGPT still doesn't know. Ask any of them "how is my job search going?" and the best a vector store can do is return the five sentences that look most similar — it can't tell you that you've applied 14 times in 16 days, that 3 became interviews, or that system design came up weak in both onsites.
Engram is a single, self-hosted memory service that every assistant shares. Agents call
remember when something happens and recall(situation) before they act. Engram turns free text
into typed memories in the background and answers with one ranked, token-budgeted brief:
An illustrative brief:
recall("I'm applying for jobs")
## You
- Backend/AI engineer, prefers concise answers, targets remote roles
## Current task
- deadline: CV to Acme by Friday
## Job search so far
14 events since 2026-09-14 (16d), last 2026-09-30 · applied 10 · interview 3 · rejection 1 · outcomes: rejected 1
Momentum is good; DSA rounds go well, system design is the recurring weak spot.
## Lessons
- [r:7] System design was flagged in 2/2 onsite interviews — prep it before the next one
## Relevant history
- [e:41] 2026-09-29 Acme onsite: system design round went poorly
- [e:38] 2026-09-24 Beta Labs rejection after recruiter screen
## How you do it
- [f:2c…] How to apply: 1. cv send <company> → 2. log it → 3. follow up in 7 daysHow it works
Engram uses every kind of memory where it fits, and one retrieval pipeline ranks them together.
Memory type | Lives in | Example |
Working / context | the brief | — |
Short-term |
|
|
Episodic |
| "applied 10 · interview 3" |
Semantic | mem0 on pgvector, with supersession | "Prefers FastAPI over NestJS" |
Procedural | mem0 ( | "How to deploy: test → build → push → verify" |
User profile | an always-included block built from your most important facts | — |
Shared | one store for every agent, each write tagged with the agent | — |
External | your documents (resume, projects…) with hybrid vector + full-text search | — |
Reflective |
| "System design is the weak spot" |
Everything is joined by entities (topic:job-search, company:acme, project:cityfix):
mention one in a situation and every layer that touches it comes along.
flowchart LR
A["Any assistant<br/>(MCP / REST / hooks)"] -- remember --> Q[(ingest queue)]
Q --> X["extract<br/>1 LLM call → typed JSON"]
X --> R["resolve entities<br/>split · dedupe · redact"]
R --> E[(episodes)]
R --> F[(mem0 facts &<br/>procedures)]
R --> S[(session state)]
F --> C{"reconcile<br/>duplicate / supersedes"}
E & F --> Z["consolidate (sleep pass)<br/>digests · reflections · profile"]
A -- "recall(situation)" --> P["plan<br/>entities · weights"]
P --> CH["channels: profile · session · aggregates ·<br/>episodes · facts · procedures · lessons · docs"]
CH --> K["rank: RRF × importance × recency × confidence"]
K --> B["token-budgeted brief"] --> AWrite path. remember returns immediately; a worker extracts entities, events, facts,
procedures and session notes in one LLM call, resolves dates in your timezone, redacts secrets,
splits multi-company career events so counts stay right, de-duplicates events that were mentioned
twice, and reconciles new facts against old ones (a changed preference supersedes the old one —
hidden from recall, kept in history).
Read path. recall matches entities (aliases, embeddings, one hop of co-occurrence), runs
the channels, fuses them with reciprocal-rank fusion, applies importance, recency half-lives and
confidence, and packs a sectioned brief under your token budget. fast=true uses no LLM at all,
and recall keeps working (profile, counts, full-text) even if the embedding model is down.
Sleep pass. Every 6 hours (and after sessions end) Engram rebuilds topic digests, derives lessons from ≥2 related events (with the events as evidence), refreshes your profile and expires old session notes. Nothing is deleted.
Full design: docs/design.md.
Related MCP server: mnemosis-mcp
Quick start
Requirements: Python 3.11+ with uv, Postgres 16 with pgvector, and Ollama for local embeddings.
git clone https://github.com/mohammed-usmani/engram && cd engram
uv sync
ollama pull nomic-embed-text # embeddings run locally and free
cat > .env <<'ENV'
CONNECTION_STRING=postgresql+asyncpg://postgres:postgres@localhost:5432/engram
DATA_DIR=data/examples # text files imported once on first start (optional)
USER_NAME=Alex # how prompts refer to you
USER_TZ=Asia/Kolkata # "yesterday" means *your* yesterday
TOGETHER_API_KEY=... # or GEMINI_/GROQ_/MISTRAL_/OPENAI_/ANTHROPIC_API_KEY
MEMORY_LLM_CHAIN=together,ollama # tried in order; local Ollama as the fallback
MEMORY_MODEL_TOGETHER=deepseek-ai/DeepSeek-V4-Flash-0731
ENV
uv run alembic upgrade head
uv run uvicorn main:app --host 127.0.0.1 --port 8001On macOS, scripts/launchd/install.sh installs it as a login service (plus nightly backups).
On Linux, scripts/systemd/install.sh does the same with systemd user units (sudo loginctl enable-linger $USER keeps it running while logged out).
Try it:
curl -X POST localhost:8001/api/memory/remember -H 'Content-Type: application/json' \
-d '{"text": "Applied to Zomato and Swiggy for SDE-2 roles last Monday", "agent": "curl"}'
# a few seconds later
curl -X POST localhost:8001/api/memory/recall -H 'Content-Type: application/json' \
-d '{"situation": "how is my job search going?", "fast": true}'Connect your assistants
Assistant | Setup |
Claude Code |
|
Codex CLI |
|
Gemini CLI |
|
Cursor / Windsurf / VS Code |
|
Claude Desktop |
|
claude.ai | set |
ChatGPT / other OAuth-only apps | also set |
Anything else | REST under |
Conversations save themselves. scripts/transcript_sync.py (installed by the
launchd/systemd scripts, every 10 minutes) reads Claude Code, Codex CLI and Gemini CLI transcripts and sends
new user/assistant prose to Engram in chunks, each stamped with the agent and the time of its last message.
Long and still-open sessions are captured in full; one offset file shared with the hook means nothing is sent
twice. Tools whose transcripts aren't readable (Antigravity, web apps) save through remember.
Every write says who and when. remember requires agent (claude-code, claude-web, chatgpt, gemini,
codex…) and occurred_at (when it happened, ISO 8601); teach and the document tools require agent, and
documents record it as updated_by. Writes without them are rejected.
Local tools on http://localhost:8001/mcp (Claude Code, Codex, Gemini CLI, Antigravity, Cursor) need no token
or sign-in; the token and OAuth only apply to requests that arrive through a tunnel.
For assistants without hooks, add one custom instruction: "Before answering anything personal or
starting a task, call recall with my words. When I share a fact, preference, decision or outcome,
call remember."
Tools: recall · remember · note · teach · expand · timeline · forget, and for
documents list_documents · get_document · search_context · edit_document · save_document ·
delete_document
Your documents
Resume, projects, skills, experience, education, achievements, certifications and profiles live
in the database and are edited in the browser at http://localhost:8001/admin, or by any assistant over MCP (edit_document for a
small change, save_document to create or replace) — no files, no reseeding. Write plain text: Key: value
lines at the top become fields, a line like Summary: starts a section; both are re-derived and the
document is re-embedded on every save, so recall sees the change immediately. Documents appear in
briefs under Documents (e.g. [d:resume/master_resume]).
A profile document (profile/linkedin, profile/indeed, profile/github…) holds the exact text
a public profile currently shows, so an assistant asked "what does my LinkedIn say?" quotes it instead
of guessing. Note anything you haven't captured yet in the document itself, so nothing gets invented.
The admin pages follow the same rule as everything else: no token from your own browser on
localhost, the token for anything arriving through a tunnel.
Text files are only an import path: on an empty database Engram imports DATA_DIR
(data/examples/ ships a fictional sample), and Import new files adds files that aren't in the
database yet. Importing never overwrites a document you've edited.
See what it knows
Open http://localhost:8001/memory for a browser view of the whole memory:
Tab | Shows |
Overview | counts per layer, topics with their summaries, recent events, last backup — and a recall box: type what you'd say to an assistant and see the exact brief it would receive |
Topics | every entity (project, company, topic…) with event counts and its "so far" summary; click through to its timeline |
Events | everything that happened, filterable by text, kind and topic |
Facts | facts, preferences and procedures, searchable (superseded ones are hidden but kept) |
Lessons | reflections with confidence and how many events support them |
Notes | short-term notes for tasks in progress |
Queue | every |
Anything wrong? Forget removes it everywhere at once — including the profile at the top of every brief, which is built live from your current facts.
Choosing the extraction model
Every memory write costs one LLM call (plus a small one when a new fact may contradict an old one). Any OpenAI-compatible provider works: Together, Groq, Mistral, Gemini, Cerebras, OpenAI, Anthropic, or local Ollama. Pick one with the included eval — 14 labelled cases (multi-company applications, relative dates, interview + offer in one sentence, procedures, session notes, noisy coding transcripts, small talk that must be ignored), repeated to catch flakiness:
uv run python scripts/eval_extraction.py --chain together --together-model <model> --repeats 5
uv run python scripts/eval_extraction.py --chain ollama --ollama-model qwen2.5-coder:7b --concurrency 1Results at the time of writing:
Model | Score | Median latency | Est. cost / month* |
| 96%, before the multi-company split fix that targets its only miss | 21 s | ≈ $1.30 |
| 92% | 20 s | free |
| 83% | 20 s | free |
*≈ 15 coding sessions, 20 notes and 30 recalls a day. Writes are asynchronous, so latency never blocks your assistant.
Batch mode for big backlogs
When the queue piles up (an import, a long day of sessions), the Queue tab on the Memory page has a
Switch queue to batch mode button. The live workers then stand aside and, about once a minute, every
waiting job goes to Together's Batch API as one batch; results come back in minutes to hours and are saved
exactly like live extractions. Batches already sent are still collected after you switch back. Batches cost
about half and use their own rate limits, but Together only batches serverless models (DeepSeek-V4-Flash is
refused), so they use MEMORY_BATCH_MODEL, default openai/gpt-oss-120b (126/126 on the eval above).
Your data stays yours
Local-first. One process on
127.0.0.1:8001; embeddings run on your machine.Secrets are redacted before anything is stored — API keys, tokens, private keys, URL credentials, card numbers (Luhn-checked), "my password is …".
Auth. With
ADMIN_TOKENset, every path except/api/health(and the OAuth endpoints below) requires the bearer token for remote or tunnelled requests; direct local use keeps working.OAuth for apps that can't send a header (ChatGPT, Gemini). Set
PUBLIC_URLto the HTTPS address the app reaches (e.g. your Tailscale Funnel URL) alongsideADMIN_TOKEN. Engram then serves MCP OAuth discovery, client registration,/authorize,/tokenand/revoke. Connecting sends you to/oauth/consent, where you approve by typingADMIN_TOKEN, so OAuth never grants more than the token does. Clients and token hashes are kept inOAUTH_STORE(default~/.config/engram/oauth.json); access tokens last an hour and refresh silently, refresh tokens 180 days. Delete that file to sign every app out.Encrypted backups.
scripts/backup.shdumps Postgres + your documents + fact history into one age-encrypted archive (nightly via launchd, keeps 14). PointENGRAM_BACKUP_DIRat a private git repo and every backup is pushed off-machine.scripts/restore.shrebuilds everything on a new machine and refuses to overwrite a database that already holds memories unless you pass--force.Several devices. Run Engram on one machine and connect the others over Tailscale with
ADMIN_TOKEN— one memory, no sync conflicts.
API
|
|
|
|
| short-term note for a session |
| save a procedure |
| expand / forget an item ( |
| chronological events for |
| extraction queue status / retry |
| run the sleep pass soon |
Engram also ships a document manager (/admin), a document store with hybrid search, and a RAG chat
(/chat) over your documents.
Development
createdb engram_test # TEST_DATABASE_URL defaults to <CONNECTION_STRING db>_test
uv run pytestsrc/memory/ models · llm (provider chain) · facts (mem0) · entities · extract · ingest
recall · consolidate · worker · api · dashboard (/memory) · backfill
src/mcp_server.py MCP tools (also served at /mcp by main.py)
scripts/ claude_memory_hook.py · eval_extraction.py · backup.sh · restore.sh · launchd/
data/examples/ fictional sample documents
docs/design.md design notesLicense
MIT © 2026 Mohammed Usmani
This server cannot be deployed
Maintenance
Related MCP Connectors
- memnodeOAuthdev.memnode
Persistent, inspectable memory for AI agents with lineage, correction, and a hosted MCP endpoint.
An MCP memory server. One memory your agents share — across models, devices and apps.
Persistent, portable memory for AI assistants — your private memory graph, from any MCP client.
Persistent personal memory for AI assistants — save, search, and recall across every MCP client.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceProvides persistent semantic memory for AI agents via MCP, enabling them to remember, recall, list, update, and forget memories with vector-based similarity search.ISC
- AlicenseCqualityAmaintenanceProvides AI agents with a human-inspired memory layer via MCP, enabling episodic and semantic memory recall, forgetting curves, consolidation, and contradiction detection. It integrates with MCP clients to offer local-first, dependency-free memory management.981MIT
- AlicenseNot gradedqualityAmaintenanceProvides AI agents with persistent, human-like memory infrastructure via MCP, enabling them to store, search, summarize, and forget episodic, semantic, procedural, and working memories across sessions.815 npmMIT
- FlicenseNot gradedqualityBmaintenanceProvides a self-hosted shared memory service that lets AI agents capture and recall durable facts, decisions, and context across multiple tools and MCP-capable clients.3-