hipercampo
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@hipercamporemember that I have a meeting at 3pm tomorrow"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
🧠 hipercampo
🌍 Español: README.es.md · You are reading the English version.
A living memory for AI agents — Claude, Codex, and whatever comes next — built on hypervectors, not embeddings.
Most LLM memories are the same thing: chunk text, turn it into dense vectors, and retrieve the closest ones (ANN / top-k). That measures similarity, but not relevance, not importance, and it never forgets. It's a landfill with a search box.
hipercampo tries something else. It's an MCP server that gives AI agents a memory
modeled on the hippocampus, with four ideas integrated into a cycle:
Idea | What it does | Inspiration |
VSA / hypervectors | Memories as 10,000-bit binary vectors with real algebra ( | Kanerva (SDM), Plate (HRR) |
Surprise-gated writing | Double veto: it won't store the redundant (something similar exists) nor the predictable (an internal incremental language model already predicted it, measured in bits — compression/MDL). That's where token savings point (not yet measured end-to-end). | Hippocampal prediction error; compression-as-intelligence (Hutter) |
Consolidation ("sleep") | An offline process groups similar episodes into a semantic memory (structural grouping: fewer nodes, text is joined; with an optional | Hippocampus→cortex replay |
Active forgetting | Strength decays with disuse; the weak becomes dormant (not deleted, like the human mind) and can later resurface via | Adaptive forgetting |
Engineering honesty. Surprise combines two signals: lexical novelty (
1 − max similarity to what's stored) and real prediction error, estimated by an in-house incremental language model in bits/token (compression/MDL, no neural net, no GPU). The base encoder is lexical; for synonyms there's an optional semantic hook (below). Everything is swappable without touching the rest.
Install (full guide: INSTALL.md)
Quick path — from PyPI:
pip install hipercampo # or: pip install "hipercampo[semantic]"
# Codex (CLI, IDE extension, and desktop app share this MCP configuration):
codex mcp add hipercampo --env HIPERCAMPO_NAMESPACE=proj-hipercampo -- python -m hipercampo.server
# Claude Code:
claude mcp add --scope user hipercampo -- python -m hipercampo.serverFrom source (contributors):
git clone https://github.com/armandojaleo/hipercampo.git
cd hipercampo && pip install -e .
python scripts/demo.py # watch the cycle runRestart your MCP client and you'll have 18 memory tools (hc_remember, hc_recall,
hc_muse, hc_dream, hc_accept_bridge, hc_reject_bridge, hc_update, hc_remember_fact, hc_ask_role,
hc_assist, hc_sleep, hc_consolidate, hc_forget, hc_health, hc_stats). For Docker, Claude Desktop,
.mcp.json, verification and troubleshooting → INSTALL.md.
Related MCP server: MCP Memory Server
30-second try (no agent client)
pip install numpy
python scripts/demo.pyYou'll see the algebra distinguishing word order and the full cycle (surprise → recall → sleep → forget) working.
Real use cases in examples/ — 7 runnable scripts: a personal
assistant that remembers across sessions, a project knowledge base with role
queries, creative brainstorming where forgotten memories resurface, temporal facts
with history, linked projects that read without writing, and the full dream →
confirm loop.
See your memory without leaving the editor: the VS Code viewer in
editor/ has four tabs — a list of cards, the associative map
(force-directed graph), a timeline with the forgetting curve, and an
importance×reliability scatter — plus project filters and a muse "eureka"
search. Read-only for browsing; it shells out to the CLI and never touches the
database directly.
Test battery — it does what it says
python -m pytest tests/ # complete regression suite
python -m ruff check hipercampo/ tests/ scripts/ examples/
python -m mypy hipercampo/
python scripts/nav_real.py --check # navigation fidelity, latency and RAM
python scripts/ablations.py --check # isolate the four cognitive mechanisms
python scripts/context_efficiency.py --check # quality, context tokens and latency
python scripts/jargon_bench.py --check # retrieval on a shared-jargon corpusjargon_bench.py exists because the other benchmarks describe their corpus but
never its regime, and that blind spot hid a real defect: spreading activation
earned nothing in recall, and no bench sat anywhere near a saturated association
graph to show it. What saturates a graph is narrowness, not scale — measured on
real memories, 240 notes spanning one project link 2.2% of pairs, while 8 notes
about a single sub-topic link 75%. This bench reproduces both and reports the
regime next to the score, so a number is never read without knowing which world it
came from.
Its numbers are deliberately harsher than the others: on shared jargon, keyword retrieval holds (MRR 0.90) while paraphrase (0.16) and synonym (0.03) collapse. That gap is the honest state of lexical encoding, and it is the ceiling the roadmap points at.
Coverage, both halves
Measured 2026-08-10. Reproduce it yourself with the two commands below — the numbers are printed by the tools, not typed here by hand.
tests | coverage | how | |
Python core ( | 344 | 83% |
|
Viewer ( | 8 end-to-end | 72% |
|
CI enforces a floor of 78% on the Python side. Viewer coverage comes from Chromium's own V8 instrumentation during the Playwright run — no extra dependency, and measured while the tests actually assert, rather than by a script driving the UI to inflate a number.
The badges above cannot rot. Their numbers live in .github/badges/*.json (read
straight from this repository; no third-party service holds the data) and CI runs
python scripts/badges.py --check, which fails if a badge claims more than the
measured coverage. The check is one-sided on purpose: a badge is never allowed to
overstate, and if coverage improves it simply understates until someone regenerates
it with --write.
Two details learned the hard way, on the very commit that introduced the badges.
Coverage is not identical everywhere — CI measures the viewer at 72.1% where a
Windows machine measures 74.3% — so the badge tracks CI, which is the gate, and
--write will lower a badge from any machine but only raise one when CI=true.
And the percentage is truncated, never rounded: rounding 72.6 up to 73 would make
the badge claim more than was measured, which is the one thing it exists to prevent.
Where the coverage is thin, stated plainly: the extension host
(editor/src/*.ts, which shells out to the CLI) has no tests at all, and the
Python CLI sits at 62%. Both are on the roadmap. Three bugs shipped through the
viewer in one week, all of them the same shape — a JSON key renamed on the Python
side that the JavaScript side kept reading, breaking a panel with no error
anywhere. tests/contracts/test_viewer_json.py now pins those key names for all
six payloads, and the end-to-end tests check the panels actually render the value.
The complete suite runs in CI on Windows, macOS and Ubuntu with Python 3.11–3.13. Example invariants checked: a duplicate never creates a second memory, a needle is retrieved among 25 distractors, forgetting never deletes something with importance ≥ 0.8, one context can neither see nor modify another's data, a failed transaction leaves no trace. The context-efficiency runner also accepts official LongMemEval JSON; it reports evidence-session recall separately from answer quality, so retrieval is not confused with an LLM judge.
Baseline comparison (Phase 2)
python scripts/baselines.py [--semantic] pits hipercampo against the standard
methods on the same corpus (10 facts + 10 confusable distractors). MRR per category
false-recall rate (unrelated queries that still return something):
method | keyword | typo | synonym | global | falseRec |
BM25 (exact lexical) | 1.00 | 0.77 | 0.33 | 0.70 | 1.00 |
embeddings + cosine | 0.95 | 0.88 | 0.77 | 0.87 | 0.20 |
hipercampo (lexical) | 1.00 | 0.85 | 0.29 | 0.71 | 0.00 |
hipercampo + semantic | 1.00 | 0.95 | 0.75 | 0.90 | 0.00 |
Honest reading:
On ranking (MRR), hipercampo+semantic wins (0.90 vs 0.87 for embeddings): it fuses lexical precision (keyword/typo) with semantic reach (synonyms). In pure-lexical mode it already beats BM25, especially on typos (character trigrams).
Abstention now works: the threshold (
ANSWER_MIN_SCORE) was mis-set below the noise floor and never filtered. Measuring and re-calibrating it (seescripts/calibrate.py) brings false-recall to 0.00 here — better than embeddings' cosine cutoff (0.20).Honest about scale: 0.00 is at N=20. On the N=500 sweep the rate settles at 0.17 (lexical) / ~0.10 (semantic) — still on par with embeddings, but not zero. Small, synthetic corpus: a signal, not proof at scale. See ROADMAP.md.
Scale & latency (measured)
Corpus | Quality | p95 | Visited |
655 real Python stdlib documents | navigation-vs-scan fidelity 1.000 | CI-gated | 42.6% |
10,000 structured memories | precision@5 1.000 | ~2.2 ms | 1.751% |
100,000 structured memories | group precision@5 1.000 | 6.94 ms | 1.094% |
Recall uses a persistent, topology-adaptive graph and hierarchical VSA landmarks,
with bounded full scan as a safe fallback. At 100k, the resident index is 141.5 MB,
cold construction takes 7.46 s and warm reuse 0.073 ms. Structured scale results show
navigation cost; the real-corpus gate separately prevents speed from hiding retrieval
quality regressions. Reproduce them with scripts/nav_scale.py and
scripts/nav_real.py --check.
Tools every agent gains
Tool | For |
| Store something (if novel/surprising). |
| Retrieve by similarity. Can abstain (return |
| Creative recall: surfaces indirect connections and dormant memories that can resurface and tie ideas together. For insight/brainstorming. |
| Creative sleep: proposes bridges between memories sharing a common associate. Hypotheses don't contaminate memory: they never propagate until confirmed. |
| Confirm a dream hypothesis (it becomes a real association) or discard it. |
| Update a fact that changed (safe supersession; the old one stays as history). |
| Sleep phase: group episodes into semantic knowledge. |
| Active forgetting. |
| Store a structured fact (compositional VSA). If it updates a current fact, the old one isn't deleted — its validity is closed and it becomes history. |
| Ask for a field knowing others: "who bites the man?" → unbinding. Answers what's currently true; |
| Memory state (includes the DB path). |
Guardrails (env): HIPERCAMPO_MAX_MEMORIES caps memories per context (evicts the
lowest-retention, never the protected); HIPERCAMPO_REDACT_SECRETS=1 masks detected
secrets before storing instead of only warning.
What it costs you (measured)
A memory that eats your context window is in the way, so the bill is measurable:
python scripts/tokens.py.
Source | Cost | When |
Announced tools (7, default) | ~810 tok | every request |
…with | ~2,070 tok | every request |
Hook injection | ≤350 tok | only on turns that fire |
The expensive part is not the memory — it's the tool descriptions, which travel in
every request even if you never call them. So only the six daily tools are
announced; the other twelve are activated hot by hc_tools, which registers
them, notifies the client (tools/list_changed) and runs the requested one in the
same call — so the capability holds even if the client ignores the notification.
Memories are injected whole or not at all: one cut in half looks like
information and isn't, so what doesn't fit is omitted and said, with a pointer to
hc_recall. And when nobody asked a question, hipercampo only interrupts if the
memory's direct activation clears a stricter bar — measured, because both the
final score and z-contrast failed to separate signal from noise. Measured end to
end: 87k → 26k tokens over a 30-turn session. Counts are character-based
estimates; install tiktoken for exactness.
The four axes of a memory (novelty ≠ importance ≠ reliability ≠ utility)
Axis | What it measures | Who sets it | Used for |
novelty / surprise | new or predictable? (MDL) | derived | decide whether to write |
importance | how much it matters | the caller ( | protect from forgetting |
reliability | how true/credible | the caller ( | ranking at retrieval |
utility | how much it's actually used | derived ( | protect from forgetting by use |
Forgetting combines the last three into a transparent retention
(0.4·importance + 0.3·reliability + 0.3·utility): time only flags candidates, but
value decides.
Compositional memory with roles (the differentiator)
The thing embeddings can't do: ask who did what to whom and get the right
answer by role. A fact is encoded by binding each value to its ROLE and bundling —
then you recover any field by unbinding (hipercampo/cycle/roles.py):
from hipercampo.roles import ItemMemory, encode_fact, query_role
im = ItemMemory()
fact = encode_fact({"subject": "dog", "predicate": "bites", "object": "man"}, im)
query_role(fact, "subject", im) # -> [("dog", 0.74)]
query_role(fact, "object", im) # -> [("man", 0.76)]python scripts/roles_demo.py shows the punchline: "dog bites man" and "man bites
dog" have the same values but the recovered subject/object are swapped — a
dense embedding places them at nearly the same point; VSA keeps them distinct.
Measured: correct filler recovered per role with a clear margin (0.74 vs 0.54),
capacity up to 5 roles. Wiring these role-records into the live MCP cycle is next
(see ROADMAP.md).
Contexts, Docker, security
Contexts: all memory lives in a single file; namespaces
(HIPERCAMPO_NAMESPACE) are drawers inside it. You write to your own and read from
the ones you link:
~/.hipercampo/hipercampo.db
├── __self__ the agent's working identity
├── personal who you are
├── proj-webshop ══> you write here while working on the shop
└── proj-blog ──> you read it, but never touch itWhat is linked is read, never touched: storing, reinforcing, forgetting and consolidating operate on your own drawer alone, and a project that is not linked is invisible. The full map, with read and write arrows, is in INSTALL.md.
You can also isolate by separate files (
HIPERCAMPO_DB) instead of namespaces. Local isolation, not multi-user security — hipercampo is local-first. See SECURITY.md.Docker:
docker compose build && docker compose run --rm hipercampo.Security: retrieved text is data, not instructions. Built-in safeguards (
hipercampo/support/safety.py):hc_rememberwarns on likely secrets (plaintext DB),hc_recallflags memories that look like injected instructions asuntrusted. They warn, not block. Details in SECURITY.md.
Architecture
text ──▶ encoder.py ──▶ hypervector (10,000 bits) semantic.py optional dense
│ bridge (SimHash)
vsa.py (bind / bundle / permute / vectorized popcount)
│
roles.py ── compositional facts (role-filler binding, temporal validity)
│
memory.py ── surprise · recall+spreading · sleep · forget · 4 axes
│ └── safety.py (secrets / injection), config.py (env)
store.py ── SQLite WAL (memories + graph, namespace-isolated, transactional)
│ └── backup.py (consistent copy), audit.py (decision log)
policy.py ── what to do at THIS turn (reads run, writes only suggested)
│
server.py ── MCP (stdio) ──▶ Claude · Codex · others cli.py ── terminal + hooksEvery public operation is wrapped in @resiliente: if SQLite fails it logs the
error, reconnects and retries once; if it still fails it returns a readable error
instead of crashing. hc_health (or hipercampo doctor) reports integrity, schema,
readability and write permission.
Related work & honest positioning
hipercampo did not invent hyperdimensional computing (HDC/VSA dates to the 90s: Kanerva, Plate), nor is it the first attempt at agent memory (Mem0, Letta, Graphiti, MemGPT; MnemoCore uses HDC). What's original is the specific combination: VSA + surprise (MDL) + consolidation + forgetting + four axes, exposed as an MCP server, treating memory as a cycle. We don't claim to beat embedding-based hybrid memories; we explore a different paradigm, with its limits measured.
License & attribution
MIT (see LICENSE). Original code; dependencies and ideas credited in ATTRIBUTION.md. House rule: if we use others' work, especially copyrighted, we say so.
Could this be a product? (spin-off, stated openly)
hipercampo stays local-first by design — that's its identity, not a budget constraint. But the core (binary hypervectors, surprise-gated writing, auditable retention) would translate to a multi-tenant memory backend: per-tenant namespaces already exist, the state machines are tested, and the algebra is CPU-cheap at scale. That would be a separate project with funding behind it — serious infrastructure, security audits, SLAs. If that's your conversation, open an issue. The core here will remain free, local and MIT either way.
Feedback (this is a beta)
hipercampo is local-first with no telemetry — nothing phones home, by design. That also means we can't see how it runs for you, so your feedback is the only signal we get, and it's what decides when the beta becomes stable.
Share your experience: open a Beta feedback issue — what you use it for, what worked, what rubbed. Praise and friction both help.
Something broken? File a bug report. The VS Code viewer has a 🐛 button that opens one for you.
Want to help? Pull requests are welcome; the quality gates run in CI.
No account-tracking, no analytics — just issues, PRs, and stars. That's the whole adoption signal, and we'd rather earn it honestly.
Acknowledgments
This project is built with Claude and Codex, for every agent that needs memory — the systems that use it help improve it, with a human making it more rigorous at every step. Armando Jaleo put the judgment, the patience and the house rule: measure before believing, and tell the truth about the limits. Thanks to Pentti Kanerva and Tony Plate, whose decades-old ideas are still alive here. And to whoever audits with rigor: honest criticism made this project better on every pass.
And yes — congratulations, Spain! 🇪🇸⚽ Some memories deserve confidence=1.0.
A memory is not a store: it's a cycle that saves, relates, consolidates, and forgets. If one day this helps machines remember with judgment — and lets the people who use them audit it — it will have been worth it. — made with care. 🧠
This server cannot be installed
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Servers
- AlicenseAqualityAmaintenanceAn MCP server that gives Claude persistent, locally-stored memory with verification-gated lessons and a Pattern Oracle to prevent unverified guesses from becoming ingrained facts.13731MIT
- Alicense-qualityDmaintenanceAn MCP server that gives Claude persistent memory by storing conversation context, entities, and enabling semantic search across sessions.231MIT
- Alicense-qualityFmaintenanceAn MCP server that gives Claude Desktop persistent memory, self-awareness, epistemic hygiene, and genuine agency across conversations with a typed memory system, embedding-based semantic search, a 3-judge memory jury, and a real-time dashboard.7MIT
- Alicense-qualityBmaintenanceA MCP server that gives Claude Code and other AI assistants long-term memory by automatically extracting technical knowledge from conversations and retrieving relevant experiences in future sessions.14MIT
Related MCP Connectors
Cloud-hosted MCP server for durable AI memory
Hosted MCP server connecting claude.ai, ChatGPT and other AI apps to your own computer
Person-owned, portable AI memory as a remote MCP server, readable and writable by any MCP client.
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/armandojaleo/hipercampo'
If you have feedback or need assistance with the MCP directory API, please join our Discord server