engram-global
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@engram-globalRemember that I want all API responses to include a request ID."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
engram — Human-like memory for AI agents (MCP server)
Persistent memory shared by Claude Code, Codex, and Antigravity (Gemini CLI). The more a memory is used, the easier it is to recall; unused memories sink but never disappear — the same dynamics as human memory.
This is the international edition: all documentation, the setup
wizard, CLI messages, and agent-facing instructions are in English, and the
default embedding model is intfloat/multilingual-e5-small (100+ languages,
~470 MB download) — so your memories themselves can be written in English or
almost any other language. A fully Japanese edition tuned for Japanese
(Ruri-v3 embeddings) lives at
engram.
Quick start
Option 1: one-line install
Windows (PowerShell):
irm https://raw.githubusercontent.com/ricoaiproject-cmd/engram-global/main/install.ps1 | iexmacOS / Linux:
curl -LsSf https://raw.githubusercontent.com/ricoaiproject-cmd/engram-global/main/install.sh | shThis single line installs uv, installs engram, and runs the setup wizard.
(macOS: git is required — run xcode-select --install first if you don't
have it.)
Option 2: manual install in three commands
Windows (PowerShell):
# 1. Install uv (skip if you already have it)
irm https://astral.sh/uv/install.ps1 | iex
# 2. Install engram
uv tool install --python 3.12 git+https://github.com/ricoaiproject-cmd/engram-global.git
# 3. Run the setup wizard
engram setupmacOS / Linux:
# 1. Install uv (skip if you already have it)
curl -LsSf https://astral.sh/uv/install.sh | sh
# 2. Install engram (force a uv-managed Python — see note below)
UV_PYTHON_PREFERENCE=only-managed uv tool install --python 3.12 git+https://github.com/ricoaiproject-cmd/engram-global.git
# 3. Run the setup wizard
engram setupWhy a uv-managed Python? engram needs a Python whose SQLite supports loadable extensions (for sqlite-vec). uv-managed Python provides this; system / python.org builds on macOS do not, and uv would otherwise prefer them when present — hence
UV_PYTHON_PREFERENCE=only-managed(install.sh sets it for you).engram doctorhas a check row for this.
The setup wizard automatically:
creates the config file (
~/.engram/config.toml)initializes the memory folder
downloads the embedding model (first run only, ~470 MB)
registers engram with Claude Code / Codex / Antigravity
registers the hooks for auto-encoding and proactive recall (Claude Code)
Related MCP server: cell-mem
After installation
Just talk to your agent
Your agent performs every engram operation on its own. You simply have normal conversations, and memories accumulate and get used automatically.
Take the onboarding interview first (recommended)
Ask your agent:
Read ~/.engram/ONBOARDING.md and interview me.Seeding your working style, preferences, and background makes the memory useful from day one.
Check your environment
engram doctorShows Python version, config file, model cache, embedding backend
(ONNX / torch), install health (detects leftover ~ngram-style remnants of a
failed pip reinstall that break import engram), per-agent registration
status as [OK] / [NG] / [--], FTS5 availability (whether SQLite's
full-text search extension is loaded, since keyword search silently degrades
without it), and a perf summary section that surfaces recent MCP tool-call
and startup timings recorded to data_dir/perf/perf_log.jsonl (see below) so
a "something feels slow" complaint can be diagnosed from data rather than
guesswork.
Re-run setup (e.g. after installing a new agent)
engram setupSafe to run any number of times (idempotent). Only unregistered agents are added.
Choose which agents to register
Even with multiple agents installed, you can connect engram to just the ones you want.
# Register with Claude Code only
engram setup --agents claude
# Register with both Claude Code and Codex
engram setup --agents claude,codexValid names: claude / codex / gemini (antigravity is an alias for
gemini). Without --agents, interactive mode lists the detected agents and
lets you pick by number (Enter selects all); --non-interactive registers all
detected agents as before.
Choosing a different embedding model
The engine is model-agnostic. Set embed_model (and, for prefix-style
retrieval models, query_prefix / doc_prefix) in ~/.engram/config.toml,
then rebuild the index:
# Example: a small English-only model (no prefixes)
embed_model = "sentence-transformers/all-MiniLM-L6-v2"
query_prefix = ""
doc_prefix = ""engram reindex # re-embed all memories with the new model
engram export-onnx # optional: regenerate the ONNX fast pathNote: the built-in mean pooling matches models such as MiniLM, E5, and Ruri; CLS-pooling models (e.g. bge) are not supported by the ONNX path.
Faster startup with ONNX
Since v0.11, engram setup (which the one-line installer runs for you)
exports the model to ONNX automatically, so no manual step is needed —
it is skipped when already exported (idempotent). Use the manual command
only when you want to re-export, e.g. after switching embedding models:
engram export-onnxThis converts the embedding model to ONNX (no extra dependencies — the
already-installed torch does the one-time conversion) and the server starts
using it automatically (embed_backend=auto). Startup drops from 12–24 s
(torch import) to ~2 s, and no MCP client timeout tuning is needed
anymore.
Safety: the export verifies that the ONNX embeddings match the torch path on
a set of sample texts (min cosine ≥ 0.999, including a long text that crosses
the sliding-window boundary of ModernBERT-based models) and refuses to
install a drifted model — a silently drifted embedding space would corrupt
recall against your existing index.db.
embed_backend in config.toml (or ENGRAM_EMBED_BACKEND) selects the
runtime: auto (default; ONNX if exported, else torch), onnx (forced;
errors if not exported), torch (forced fallback).
Startup mode (ENGRAM_PRELOAD)
The default is auto (v0.10.0+): if the ONNX model has been exported, engram
picks background (handshake responds immediately); on the torch fallback it
picks blocking. No tuning is normally needed, and clients whose MCP startup
timeout settings have no effect (seen in the wild: Codex Desktop 26.707)
connect out of the box.
Value | Behavior |
|
|
| Load the model on the main thread before answering the handshake. Every |
| Respond to the handshake immediately and load the model in a background thread. Safe on the ONNX path (the first tool call just waits a few seconds; measured ~5 s even off the main thread). Not recommended on the torch path: on Windows, importing torch on a non-main thread while the asyncio event loop is running is pathologically slow (measured ~184 s vs ~20 s on the main thread), so the first |
| No preload; the model loads lazily on the first tool call. |
If engram fails to connect at startup on the torch fallback, raise the client's
MCP startup timeout (for Claude Code: MCP_TIMEOUT=120000) or run
engram export-onnx, rather than forcing background — on torch that only
converts a visible startup timeout into a 3-minute first recall.
More like real memory (new in v0.3)
Auto-encoding — sessions become memories by themselves
When a Claude Code session ends, a hook summarizes the conversation and saves
it as an episode memory (engram setup registers the hook for you). Even if
the agent forgets to call remember, "what we did yesterday" is preserved.
Disable with auto_encode = false in config.toml.
Proactive recall — memory speaks up on its own
Every time you say something, a hook runs a lightweight search (a fast path
that never loads the embedding model) for related memories. The mode is set
by surface_mode in config.toml:
Mode | Behavior |
| Injects nothing; logs "this is what I would have surfaced" (for observation and tuning) |
| Actually injects strongly related memories into the agent's context |
| Does nothing |
The log lives at ~/.engram/surface/surface_log.jsonl. We recommend watching
shadow mode for a while and switching to active once the surfaced candidates
look right. Use engram surface "some text" to check manually what would
surface.
Tuning parameters: surface_threshold (score threshold, default 0.45) /
surface_min_relevance (relevance floor, default 0.25 — a gate that keeps
even important memories from surfacing when they are unrelated to what you
said) / surface_max_items (max items per prompt, default 2).
Memory rooms — separating work and personal contexts
Every memory carries a room label. Map folders to rooms in config.toml
and the room is resolved automatically from the working directory:
[room_paths]
'C:/Users/you/work-projects' = 'work'
'C:/Users/you/personal' = 'personal'Unmapped folders and pre-existing memories are all
commonrecall searches only "current room + common" (
room="*"searches across all rooms)Auto-encoding and proactive recall respect rooms too, so work memories never leak into personal contexts (and vice versa)
Sharing memories across machines
The Markdown store can live on a synced folder (e.g. a cloud drive) shared by
several machines, but index.db is per-machine and local — so a memory written
on one machine isn't searchable on another until that machine indexes it. The
MCP server checks this at startup via startup_index_check in config.toml:
auto (default) reindexes when it detects a markdown/index mismatch, warn
logs a notice, off disables it. You can also run engram reindex any time.
How the memory works
Core design
Embeddings (where a memory sits in meaning-space) stay fixed; a separate axis — activation — modulates search ranking. Meaning = where it is, activation = how easily it comes to mind.
The same properties as human memory
The more you use it, the easier it is to recall — ACT-R activation model; every use by the agent reinforces it automatically
Unused memories sink but never disappear — power-law decay; deep recall can always reach them through associative links
Memories from striking contexts are engraved deeply — initial encoding boost by importance + slower decay = flashbulb memory
Corrected mistakes are engraved deepest of all — the correct tool records the error together with the fix = hypercorrection effect
Memory dynamics in brief
Activation:
B = ln(Σ w_j·(now−t_j)^(−d_i)), normalized to 0..1 with a sigmoid; computed on the fly from the access logd_i = clamp(0.5 − 0.2·(imp−5)/5, 0.3, 0.6)— higher importance forgets slowercreate event weight
1 + 2·(imp/10)— critical memories start strongmerely recalled: weight 0.3 / actually useful (reinforce): weight 1.0×strength
Memories reinforced together get co_recall links (Hebbian learning), growing an associative network that deep recall's spreading activation can traverse
Search
Vector neighbors (sentence-transformers embeddings) + BM25 full-text search
merged with RRF, then re-ranked by
0.6·relevance + 0.25·activation + 0.15·importance.
Hybrid recall: exact tokens no longer sink
Candidate relevance blends the two search paths instead of collapsing FTS
hits onto the vector similarity scale: vector hits keep their cosine
similarity, and FTS hits get a lexical relevance 1 - exp(bm25) derived
directly from BM25 (0 for bm25 >= 0). When an id is hit by both, the higher
of the two wins. Rare, decisive lexical matches — memory IDs, file paths,
error codes, other exact tokens — push bm25 deep negative and surface
lex near 1.0, clearing the narrow band into which dense-retriever cosine
similarities tend to compress (with the Japanese model Ruri-v3, for example,
0.8–0.87). Previously, FTS-only hits were assigned the minimum vector
similarity among the candidate pool, which buried exact-match results at the
bottom of the ranking even when they were the obviously correct answer.
Short queries are covered too (v0.7.1). The FTS5 trigram tokenizer cannot
index terms shorter than 3 characters, so 2-character terms (common in CJK
text) used to be invisible to lexical search, and mixed queries containing
any short token returned zero rows because of the implicit AND. The MATCH
expression is now built from tokens of 3+ characters only, and when no such
token exists the search falls back to LIKE substring matching with an
IDF-based pseudo score (lex = N/(N+df): rarer terms score higher).
Memory types
type | Contents |
knowledge | Insights, solutions to problems, how to use tools |
preference | The user's preferences, style, patterns in instructions |
project | Goals, constraints, history, and background of the work |
episode | A summary of what happened in a session |
File layout
~/.engram/
config.toml Config file (generated by engram setup)
index.db SQLite index (rebuildable from Markdown via reindex)
MEMORY_PROTOCOL.md Agent operating instructions (imported into each agent's instruction file)
ONBOARDING.md Initial interview script
surface/ Proactive recall log and session state
hooks.log Hook activity log
consolidation_state.json Candidate-cluster count + last-nudge timestamp (consolidation nudge)
perf/perf_log.jsonl Timing log for MCP tool calls and startup (when perf_log = true)
<memories_dir>/ Source of truth: Markdown (opens and edits fine in Obsidian)
knowledge/
preferences/
projects/
episodes/YYYY/MM/
_trash/memories_dir defaults to ~/.engram/memories, but pointing it at a Google
Drive or OneDrive synced folder gives you backup and multi-device sharing.
The SQLite index always stays local, so there is no sync-conflict risk.
MCP tools
Tool | When to use |
| At task start. fast = normal / deep = explores associative links, cold tier, and episodes / exhaustive = relevance-only full scan, ignoring activation, to dig up sunk memories |
| When you learn an insight, preference, context, or event. importance 1–10 scores how critical the context is |
| At task end, report which memories actually helped (the nutrient for consolidation) |
| When a memory was wrong. Use this, not forget (engraves the mistake itself deeply) |
| Auxiliary operations |
| Consolidation (below) |
Consolidation (the sleep of the system)
Clusters old episode memories and distills them into knowledge. The server only returns candidates; the summarization is done by the LLM (your agent). Example nightly run:
claude -p "Call engram's consolidation_candidates, summarize each cluster with remember (type=knowledge or project, related_ids=the source episodes), finish with mark_consolidated, then report stats."Automatic nudge cycle
On top of the cron-style nightly run above, engram nudges the agent to consolidate on its own, without any scheduled job:
SessionEnd counts how many consolidation-candidate clusters currently exist (via
consolidation_candidates) and stores the count indata_dir/consolidation_state.json.UserPromptSubmit (the same lightweight hook that powers proactive recall) checks that state on the next session and, if enough clusters have piled up and enough time has passed since the last nudge, injects an
additionalContextmessage asking the agent to runconsolidation_candidates→remember→mark_consolidatedat a natural pause in the conversation. The nudge fires even whensurface_mode = "off"— it is independent of proactive recall.
Controlled by three settings (in config.toml or as ENGRAM_* environment
variables):
Setting | Default | Meaning |
|
| Master switch for the nudge cycle |
|
| Minimum candidate clusters before nudging |
|
| Minimum time between nudges |
For developers
Setting up to develop in this repository (run at the repo root):
# Virtual env (for development; distribution uses uv tool install)
python -m venv "$env:USERPROFILE\.engram\venv"
& "$env:USERPROFILE\.engram\venv\Scripts\python.exe" -m pip install -e ".[dev]"Tests and verification
$py = "$env:USERPROFILE\.engram\venv\Scripts\python.exe"
& $py -m pytest # all tests
& $py -m pytest tests\test_setup.py -q # setup logic only
& $py scripts\simulate.py # simulate access patterns (30 days)
& $py scripts\check_mcp_e2e.py # MCP end-to-end checkDiagnostics / CLI (for manual checks)
$engram = "$env:USERPROFILE\.engram\venv\Scripts\engram.exe"
& $engram doctor
& $engram remember "content" --type knowledge --importance 7
& $engram recall "query" --deep
& $engram surface "utterance text"
& $engram statsProject structure
src/engram/
config.py Settings (defaults < config.toml < env vars) + room resolution
engine.py The memory engine
store.py Markdown source-of-truth store
db.py SQLite index (sqlite-vec + FTS5)
dynamics.py ACT-R activation model
embedder.py RuriEmbedder / FakeEmbedder
server.py MCP server (stdio)
cli.py CLI entry point
setup.py Setup wizard & doctor & hook registration
hooks.py Hook entry points (auto-encoding / proactive recall)
transcript.py Deterministic transcript summarization (auto-encoding)
surface.py Lightweight search path for proactive recall (no model)
templates/ MEMORY_PROTOCOL.md / ONBOARDING.md
tests/
test_setup.py Setup pure-logic tests
test_config.py Settings precedence tests
test_store.py Markdown store tests
test_db.py DB operation tests
test_engine.py Engine tests
test_room.py Memory room tests
test_surface.py Proactive recall tests
test_transcript.py Transcript summarization tests
test_hooks.py Hook and hook-registration tests
test_integration.py Integration testsTroubleshooting
Codex says engram is enabled but the connection times out on startup
On the torch fallback path, engram loads the embedding model (plus checks the
memories folder) on every startup, which can take longer than Codex's default
30-second MCP startup timeout — especially right after a reboot, during
antivirus scans, or when the memories folder lives on a cloud-synced drive
(Google Drive, OneDrive, etc.). Your memories are fine; only the initial
connection is timing out. (Running engram export-onnx once largely
eliminates this — startup drops to ~2 s.)
Newer versions of engram setup write a longer startup timeout automatically.
If you registered with an older version, either re-run
engram setup --agents codex, or add one line to the engram block in
~/.codex/config.toml yourself:
[mcp_servers.engram]
command = "..." # leave as is
startup_timeout_sec = 120.0 # add this lineThen fully restart Codex (quit and relaunch, not just close the window).
Update
Re-run the same one-line installer to overwrite with the latest version:
irm https://raw.githubusercontent.com/ricoaiproject-cmd/engram-global/main/install.ps1 | iexmacOS / Linux:
curl -LsSf https://raw.githubusercontent.com/ricoaiproject-cmd/engram-global/main/install.sh | shuv tool upgrade engram does the same (package update only).
The one-line installer also re-runs setup afterwards, so environments that never exported the ONNX model get converted automatically and startup drops to ~2 s (skipped when already exported)
Your memories and config (
~/.engramand the memories folder) are kept — nothing is deletedAfter updating, restart each agent that uses engram (Claude Code, etc.) so the MCP server reconnects
Uninstall
# 1. Remove the registration from each agent
claude mcp remove engram
# 1b. Manually remove the engram entries from hooks in ~/.claude/settings.json
# (the "engram hook ..." commands under SessionEnd / UserPromptSubmit)
# 2. Manually remove the engram block from ~/.claude/CLAUDE.md
# 3. Manually remove the [mcp_servers.engram] block from ~/.codex/config.toml
# 4. Manually remove the engram entry from ~/.gemini/config/mcp_config.json
# 5. Uninstall engram itself
uv tool uninstall engram
# 6. To delete the data as well (memories, config, model cache)
Remove-Item -Recurse -Force "$env:USERPROFILE\.engram"
Remove-Item -Recurse -Force "$env:USERPROFILE\.cache\huggingface\hub\models--intfloat--multilingual-e5-small*"On macOS / Linux, steps 5–6 are:
uv tool uninstall engram
rm -rf ~/.engram
rm -rf ~/.cache/huggingface/hub/models--intfloat--multilingual-e5-small*This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- Alicense-qualityBmaintenanceSelf-hosted MCP server giving AI agents persistent memory for personalization and context across conversations.Last updated195Apache 2.0
- Alicense-qualityBmaintenanceAn MCP server that provides AI agents with persistent, multi-layered memory inspired by the human brain, including consolidation, self-reflection, and generative replay.Last updated1MIT
- AlicenseAqualityDmaintenanceMCP server for long-term agent memory, providing persistent memory, searchable knowledge, and evolving identity for AI agents.Last updated53Apache 2.0
- Flicense-qualityCmaintenanceA self-hosted MCP server that gives AI agents persistent, searchable memory with importance scoring, knowledge graphs, and autonomous memory consolidation.Last updated1
Related MCP Connectors
Cloud-hosted MCP server for durable AI memory
Person-owned, portable AI memory as a remote MCP server, readable and writable by any MCP client.
An MCP server that integrates with Discord to provide AI-powered features.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/ricoaiproject-cmd/engram-global'
If you have feedback or need assistance with the MCP directory API, please join our Discord server