ZenBrain
OfficialThis MCP server gives AI agents a persistent 7-layer long-term memory: it can store memories (auto-routed to the right layer), recall them by relevance, run sleep-style consolidation, and report memory-layer fill levels.
Store memories (
zenbrain_store): write content in plain language so it survives the conversation; supportscontext,source,outcome, and returns the new memory'sid.Automatic layer routing: procedures (steps/instructions), episodes (emotional weight > 0.5), core memories (confidence > 0.9), or semantic facts — overridable with
typeand tunable viaemotionalWeight/confidence(0–1).Procedure capture: supply ordered
stepsand thetoolsa procedure uses; steps are otherwise parsed from numbered/bulleted lines.Deduplication caveat: every call adds a new memory (duplicates possible) except re-storing the same core memory, which updates it; recall first if unsure.
Recall memories (
zenbrain_recall): relevance-ranked search across episodic, semantic, procedural and core layers, optionally includingworking; results are tagged with layer, score, confidence and emotional weight.Filtering and limits: cap results with
limit(1–100, default 10) and drop low-confidence hits withminConfidence; core blocks, episodes and working-memory items always count as confidence 1.Consolidate memory (
zenbrain_consolidate): promotes up to 100 recent highly emotional episodes into semantic facts (once each), decays aging working-memory slots, and prunes slots whose relevance falls to ≤0.01 — nothing in long-term memory is deleted.Inspect memory state (
zenbrain_health): read-only, cheap counts of working-memory slots used/max, short-term interactions, episodes, facts (withdueForReview), procedures and core blocks — good for comparing before/after consolidation.Recommended workflow: check with recall before storing, run consolidation between sessions or after batches of stores, and use health to gauge how much is stored versus recall to read it.
arXiv preprint (cs.AI): arxiv.org/abs/2604.23878
Open-access archive (Zenodo / CERN): doi.org/10.5281/zenodo.19353663
ORCID: 0009-0001-1793-012X
License: CC BY 4.0 (paper) · Apache-2.0 (code)
@misc{bering2026zenbrain,
title = {ZenBrain: A Neuroscience-Inspired 7-Layer Memory Architecture for Autonomous AI Systems},
author = {Bering, Alexander},
year = {2026},
eprint = {2604.23878},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
doi = {10.5281/zenodo.19353663},
url = {https://arxiv.org/abs/2604.23878}
}Feedback, replications, and counter-results are explicitly welcome — please open an issue or reach out via research@zensation.ai.
Your AI forgets everything after every conversation. ZenBrain fixes that — with the same mechanisms your brain uses: spaced repetition, emotional consolidation, Hebbian strengthening, and exponential forgetting curves. Not a vector database with a wrapper. Actual neuroscience.
Architecture vs. this package. ZenBrain's architecture is 15 neuroscience-inspired mechanisms — 9 foundational algorithms + 6 Predictive Memory Architecture (PMA) components (paper). The 6 PMA components are proprietary and run in the production system. This open-source package ships the algorithm library: 10 core algorithms + 10 advanced research modules (20 modules), zero-dependency.
Benchmark: LongMemEval-500
On LongMemEval-500, three of nine head-to-head answer-quality comparisons hold against Letta, Mem0 and A-Mem — all three against A-Mem, the remaining six are ties, none lost. Three competitors x three LLM judges, Bonferroni-corrected (alpha = 0.05/18) and version-matched. It reaches 91.3% of a full-context oracle's binary-judge accuracy at 1/109.6 of the per-query token cost (47.7% vs. 52.2%).
Works out of the box without an embedding provider — lexical ranking, zero
dependencies. With nomic-embed-text as the embedding provider you get the
configuration those figures were measured in.
The paper prints where ZenBrain loses as well: on LoCoMo, substring-based aggregate F1 favours lexical retrieval (BM25) by metric design, and we do not contest that. The advantage is most pronounced on judge-graded answer quality and cross-session reasoning.
The mechanism comparison further down re-runs from this repository in under a minute —
bash scripts/compare-mechanisms.sh, no API keys and nothing to install. It prints a positive
and a negative control before the result, so the instrument can be checked before its output
is trusted. The method, the effect sizes and the ablations behind the numbers above are in the
paper; this repository ships no runner for them. The archived packages below run the significance
tests and effect sizes in full, and the mechanism ablation behind paper Tables 7–9.
Method, effect sizes and ablations: arXiv:2604.23878
Open-access archive: 10.5281/zenodo.19353663
Related MCP server: brainlayer
Reproduction packages
The raw material behind the numbers above is deposited on Zenodo, open access and citable. Both links are concept DOIs and resolve to the latest version, the same convention this README uses for the paper archive; each description was measured against the version named after it.
Mechanism ablation, paper Tables 7–9 — 10.5281/zenodo.22162063 (described here: v1.0.0). Four experiment suites (95 tests), the reference JSON the paper's tables were generated from, and
verify-against-reference.mjs, which diffs a fresh run against that reference and exits non-zero on drift.npm install && npm run experiments; the run itself needs no API keys and no network, and finishes in under a minute on a laptop. Two of the paper's other ablation tables need data this package does not carry: Table 11 the LoCoMo corpus, Table 13 a different pipeline. The package says so itself.Measurement package, LongMemEval-500 and the real-pipeline flag ablation — 10.5281/zenodo.22161977 (described here: v1). Per-(system, judge, seed) judged outputs, the flag manifests as recorded at run time, a
SHA256SUMS.txtcovering every file in the package, and the analysis scripts. Three of those scripts are stdlib-only and self-checking — the oracle comparison behind the 91.3% figure, the judge-agreement figures, and the real-pipeline flag-ablation table: each prints every re-derived value next to the reference it has to match, and exits non-zero on mismatch. The significance tests behind the nine head-to-head comparisons against Letta, Mem0 and A-Mem sit in a separate script that needs numpy and scipy; it recomputes all eighteen pairwise tests and rewrites the deposited significance JSON byte-identically, so what catches a mismatch there is the checksum, not an exit code.
Both packages name what they do not cover. Replications and counter-results are welcome: research@zensation.ai.
How ZenBrain differs from Mem0, Letta and Zep
ZenBrain implements fifteen mechanisms taken from human memory research. No system among those surveyed in the paper integrates more than two of them. The table below records which of the mechanisms appear in the public source of three widely used memory systems, at pinned versions, on a fixed date.
Mechanism | ZenBrain | Mem0 | Letta | Zep |
FSRS spaced repetition | yes | — | — | — |
Hebbian learning | yes | — | — | — |
Ebbinghaus forgetting curves | yes | — | — | — |
Sleep consolidation | yes | — | — | — |
Emotional tagging | yes | — | — | — |
Zero runtime dependencies | yes | — | — | — |
How this was measured, 27 August 2026. Full-text search over the checked-out public source of mem0ai/mem0 (npm mem0ai 3.1.7, PyPI mem0ai 2.0.19), letta-ai/letta-code (npm @letta-ai/letta-code 0.31.2) and getzep/zep, lockfiles excluded. A dash means the term does not occur in that snapshot — not that the system cannot do something comparable under another name. Dependency counts are declared direct dependencies: @zensation/core resolves to two packages, both our own; mem0ai declares four, @letta-ai/letta-code eighteen. Re-run the whole check yourself with scripts/compare-mechanisms.sh; it prints its own positive and negative controls so you can see the instrument works before you trust the result.
Human memory does not work like a key-value store. The brain keeps specialised systems for different kinds of memory, forgets actively, modulates by emotion and retrieves by context. ZenBrain brings those mechanisms to AI agents.
Advanced algorithms (since v0.3.0, May 2026)
On top of the 10 core algorithms above, @zensation/algorithms ships 10 advanced algorithms grounded in recent neuroscience and ML research. Each is exposed as its own sub-path (@zensation/algorithms/<name>) and remains zero-dependency:
fsrs-vmPFC— Prediction-Error coupled FSRShebbian-two-factor— Two-Factor synaptic consolidationsleep-simulation-selection— RL-based replay selectionspectral-health— Fiedler-value KG health monitorib-budget— Information-Bottleneck retention budgetdopamine-routing·hopfield-stm·personalized-pagerank·surprise-gradient-memory·temporal-multi-route
See CHANGELOG.md for details.
Quick Start
Requires Node.js 22 or newer (since
0.4.0). On Node 20 or older, npm silently installs the last compatible release (@zensation/algorithms@0.3.4,@zensation/core@0.2.2) instead of the current one — which looks like a broken package but is a platform mismatch. See CHANGELOG.
npm install @zensation/algorithmsimport {
initFromDecayClass,
getRetrievability,
updateAfterRecall,
tagEmotion,
computeEmotionalWeight,
computeHebbianStrengthening,
propagateForRelation,
} from '@zensation/algorithms';
// 1. Schedule a memory with FSRS
const memory = initFromDecayClass('normal_decay');
// 2. A week later, check recall probability (Ebbinghaus curve)
const aWeekLater = new Date(Date.now() + 7 * 24 * 60 * 60 * 1000);
const retention = getRetrievability(memory, aWeekLater);
console.log(`Recall probability: ${(retention * 100).toFixed(1)}%`);
// ~36.8% — retrievability has decayed over the week
// 3. User recalled it anyway — update scheduling
const updated = updateAfterRecall(memory, 4, retention, aWeekLater);
// stability 7 -> 8.19: recalling at low retrievability gives a bigger boost
// 4. Tag emotional significance
const emotion = tagEmotion('I am absolutely thrilled — I got the promotion!');
const weight = computeEmotionalWeight(emotion);
console.log(`Decay multiplier: ${weight.decayMultiplier}x`);
// 2.7x — emotional memories decay nearly 3x slower
// 5. Strengthen knowledge connections (Hebbian)
const stronger = computeHebbianStrengthening(1.0);
// 1.09 — "neurons that fire together wire together"
// 6. Propagate confidence through your knowledge graph
const confidence = propagateForRelation(0.5, 0.8, 1.0, 'supports');
// 0.9 — supporting evidence increases confidenceWant the advanced algorithms?
import {
computeKGPredictionError,
computeAdaptiveFSRSInterval,
} from '@zensation/algorithms/fsrs-vmPFC';
// Couple FSRS scheduling with the prediction-error signal from your
// knowledge graph: when the embedding has shifted a lot since the last
// review (high cosine distance), shrink the next interval; otherwise push
// it out. Both arrays must have the same length.
const lastEmbedding = [0.1, 0.2, 0.3, 0.4];
const currentEmbedding = [0.5, 0.4, 0.1, 0.2];
const pe = computeKGPredictionError(lastEmbedding, currentEmbedding);
const nextInterval = computeAdaptiveFSRSInterval(14, pe);Each advanced algorithm has its own sub-path (@zensation/algorithms/spectral-health, @zensation/algorithms/ib-budget, …). All zero dependencies.
Runnable examples
Seven self-contained examples live in examples/:
Example | Shows |
Working Memory + Short-Term Memory for conversation context — no SDK needed | |
An Anthropic Claude assistant that remembers across conversations | |
ZenBrain as the memory backend of a LangChain agent | |
Multiple agents sharing Working Memory, with Hebbian strengthening | |
A memory-aware system prompt for the Vercel AI SDK | |
ZenBrain as long-term memory for a LlamaIndex.TS agent — LlamaIndex keeps the short-term window, ZenBrain decides what outlives it | |
ZenBrain behind a Mastra agent: a Processor records each turn, and the instructions are rebuilt from what is currently recallable |
npx tsx examples/basic-chatbot.tsThe integration examples additionally need their respective SDK installed; examples/README.md lists which. Open good first issues are the best way in if you want to contribute.
The Science Behind It
7-Layer Memory Architecture
Layer 7: Cross-Context Memory ← Shared knowledge across domains
Layer 6: Core Memory ← Pinned facts (Letta-style)
Layer 5: Procedural Memory ← "How to do X" (skills & workflows)
Layer 4: Long-Term Semantic ← Facts with FSRS scheduling
Layer 3: Episodic Memory ← Concrete experiences & events
Layer 2: Short-Term / Session ← Current conversation context
Layer 1: Working Memory ← Active task focus (7±2 items)Each layer has different retention characteristics, consolidation rules, and retrieval mechanisms — just like the human brain.
FSRS Spaced Repetition
FSRS (Free Spaced Repetition Scheduler) outperforms SM-2 by 30%. It uses the desirable difficulty principle: reviewing when retention is low gives a bigger stability boost. Your AI reviews important facts at optimal intervals — never too early (wasteful), never too late (forgotten).
Emotional Memory
The amygdala modulates memory consolidation — emotional events are remembered more vividly (flashbulb memory). ZenBrain's emotional tagger assigns arousal, valence, and significance scores using a 400+ keyword lexicon (English & German). Emotional memories get up to 3x longer decay half-lives.
Hebbian Learning
"Neurons that fire together wire together" (Hebb, 1949). Knowledge graph edges that are frequently co-activated grow stronger. Unused edges decay and eventually get pruned. The result: a self-organizing knowledge structure that reflects actual usage patterns, with homeostatic normalization to prevent runaway growth.
Ebbinghaus Forgetting Curves
Ebbinghaus (1885) showed that memory decays exponentially: R = e^(-t/S). ZenBrain implements personalized decay profiles that adapt to individual learning patterns, with SM-2 compatibility for existing spaced repetition systems.
Context-Dependent Retrieval
Tulving's Encoding Specificity Principle (1973): memories are recalled better when the retrieval context matches the encoding context. ZenBrain captures temporal context (time of day, day of week) and task type at encoding time, providing up to a 30% retrieval boost when contexts match.
Bayesian Confidence Propagation
Knowledge isn't isolated — facts support or contradict each other. ZenBrain propagates confidence through your knowledge graph using Bayesian belief updates: supporting evidence increases confidence, contradictions decrease it, with damping for numerical stability.
Sleep Consolidation
During sleep, the hippocampus replays recent experiences, strengthening important memories and pruning weak connections (Stickgold & Walker, 2013). ZenBrain simulates this process: selectForReplay() prioritizes emotional and recently-accessed memories, simulateReplay() boosts their stability by 50%, and pruneWeakConnections() removes weak Hebbian edges — implementing the Synaptic Homeostasis Hypothesis (Tononi & Cirelli, 2006).
import { selectForReplay, simulateReplay } from '@zensation/algorithms/sleep-consolidation';
// Select memories for overnight consolidation
const toReplay = selectForReplay(allMemories);
// Simulate sleep replay — stability ↑, weak edges pruned
const result = simulateReplay(toReplay);
console.log(`Replayed ${result.summary.totalReplayed} memories, avg stability +${result.summary.avgStabilityIncrease.toFixed(1)} days`);Memory Coordinator
The MemoryCoordinator orchestrates all 7 layers into a single cohesive system — inspired by Global Workspace Theory (Baars, 1988):
import { MemoryCoordinator } from '@zensation/core';
const memory = new MemoryCoordinator({ storage: adapter, embedding: embedder });
// Auto-routes to the right layer (semantic, episodic, procedural, or core)
await memory.store('User prefers TypeScript', { type: 'auto' });
// Cross-layer search with ranked, deduplicated results
const results = await memory.recall('programming preferences');
// Consolidate: promote episodic → semantic, apply decay
await memory.consolidate();
// FSRS review queue across all layers
const dueItems = await memory.getReviewQueue();Packages
Package | Description | Status |
20 algorithm modules — 10 core (FSRS, Hebbian, Ebbinghaus, emotional, Bayesian, sleep consolidation, intervals, visualization) + 10 advanced (vmPFC-FSRS, two-factor Hebbian, IB budget, Hopfield STM, …) | :white_check_mark: Published | |
Memory layers, coordinator, adapter interfaces | :white_check_mark: Published | |
PostgreSQL + pgvector storage adapter | :white_check_mark: Published | |
SQLite storage adapter (zero-config) | :white_check_mark: Published | |
MCP server — gives any MCP client (Claude Desktop, Claude Code, Cursor) the seven layers as four tools. Carries the protocol SDK, so the core stays dependency-free | :white_check_mark: Published | |
Vercel AI SDK middleware — recall before the model call, store after it. Works with any provider, zero runtime dependencies | :white_check_mark: Published |
Tree-Shakeable Imports
Every algorithm is available as a subpath export:
// Import everything
import { tagEmotion, updateAfterRecall } from '@zensation/algorithms';
// Or just what you need (better tree-shaking)
import { updateAfterRecall } from '@zensation/algorithms/fsrs';
import { tagEmotion } from '@zensation/algorithms/emotional';
import { computeHebbianStrengthening } from '@zensation/algorithms/hebbian';
import { propagateForRelation } from '@zensation/algorithms/bayesian';
import { selectForReplay } from '@zensation/algorithms/sleep-consolidation';
import { getRetrievabilityWithCI } from '@zensation/algorithms/intervals';
import { generateRetentionCurve } from '@zensation/algorithms/visualization';Use Cases
AI Chatbots with Long-Term Memory
import { updateAfterRecall, getRetrievability, scheduleNextReview } from '@zensation/algorithms/fsrs';
import { tagEmotion, computeEmotionalWeight } from '@zensation/algorithms/emotional';
// When your AI learns a fact about the user:
function rememberFact(fact: string) {
const memory = initFromDecayClass('normal_decay');
const emotion = tagEmotion(fact);
const weight = computeEmotionalWeight(emotion);
// Emotional facts get longer retention
return {
...memory,
emotionalWeight: weight.consolidationWeight,
decayMultiplier: weight.decayMultiplier,
};
}
// Before each conversation, check what needs reinforcement:
function getFactsDueForReview(facts: MemoryState[]) {
return facts.filter(f => getRetrievability(f) < 0.7);
}Knowledge Graph with Self-Organizing Edges
import { computeHebbianStrengthening, computeHebbianDecay } from '@zensation/algorithms/hebbian';
import { propagateForRelation } from '@zensation/algorithms/bayesian';
// When two concepts are mentioned together:
function coActivate(edge: { weight: number }) {
edge.weight = computeHebbianStrengthening(edge.weight);
}
// Periodic maintenance — decay unused edges:
function decayEdges(edges: { weight: number; lastUsed: Date }[]) {
for (const edge of edges) {
edge.weight = computeHebbianDecay(edge.weight);
// Edges below MIN_WEIGHT (0.1) can be pruned
}
}RAG with Confidence Scoring
import { propagateForRelation, isSignificantChange } from '@zensation/algorithms/bayesian';
// After retrieval, propagate confidence through related facts:
function updateConfidenceGraph(facts: Fact[], relations: Relation[]) {
for (const rel of relations) {
const newConf = propagateForRelation(
rel.target.confidence,
rel.source.confidence,
rel.weight,
rel.type // 'supports' | 'contradicts' | 'related_to'
);
if (isSignificantChange(newConf, rel.target.confidence)) {
rel.target.confidence = newConf;
}
}
}Extracted From Production
These aren't toy implementations — ZenBrain's algorithms are extracted from ZenAI, a production AI platform. Everything claimed here is verifiable in this repository:
801 tests (489 algorithms · 137 core · 54 adapter-sqlite · 48 adapter-postgres · 55 mcp · 18 ai-sdk), all passing
Zero runtime dependencies — pure TypeScript, dual ESM + CJS, tree-shakeable subpath exports
Reproducible — building from this source produces the same 153-file
@zensation/algorithms@0.5.0tarball published on npm7-layer memory architecture grounded in published neuroscience
Contributing
We welcome contributions! See CONTRIBUTING.md for guidelines. Issues and pull requests are answered; replications and counter-results are especially welcome.
Resources: API Reference | Recipes | Architecture | Benchmarks | FAQ | Roadmap
# Clone the repo
git clone https://github.com/zensation-ai/zenbrain.git
cd zenbrain
# Install dependencies
npm install
# Run tests
npm test
# Build all packages
npm run buildResearch
ZenBrain's architecture and algorithms are documented in an open-access technical disclosure:
arXiv preprint (cs.AI): arxiv.org/abs/2604.23878
Open-access archive: ZenBrain: A Neuroscience-Inspired 7-Layer Memory Architecture for Autonomous AI Systems (Zenodo, DOI: 10.5281/zenodo.19353663 — resolves to the latest version)
TDCommons: Technical Disclosure (CC BY 4.0)
HuggingFace: Model Card & Benchmarks
If you use ZenBrain in academic work, please cite:
@misc{bering2026zenbrain,
title = {ZenBrain: A Neuroscience-Inspired 7-Layer Memory Architecture for Autonomous AI Systems},
author = {Bering, Alexander},
year = {2026},
eprint = {2604.23878},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
doi = {10.5281/zenodo.19353663},
url = {https://arxiv.org/abs/2604.23878}
}Cited by
Five works cite ZenBrain (as of 28 September 2026): a survey by the Sico team at Microsoft Research, MindMemOS by Huawei's Noah's Ark Lab, the EVD preprint, the HindsightTag manuscript and the Lamperouge manifesto.
Agentic Evolution: From Self-Improving Agents to Co-Evolving Human-AI Systems, a survey by the Sico team at Microsoft Research (2026), calls forgetting "critically understudied": very few of the roughly 100 Memory & Sense papers it surveys implement explicit forgetting. It lists ZenBrain as one of four works in its forgetting/lifecycle group and as its only source for sleep-based consolidation.
MindMemOS (Huawei's Noah's Ark Lab, arXiv:2608.12428, August 2026) relies on exactly three references for the foundational claim of its introduction: two field surveys and a single system paper — ZenBrain.
HindsightTag (Vivek Govindbhai Dudhat, manuscript, July 2026) discusses ZenBrain as its closest prior system and uses its
MemoryCoordinatorinterface as the integration example.EVD: An Emotional Valence Dimension for Persistent Agent Memory (R. J. Vandelinder and I. Vandelinder, Exile Research, Zenodo, June 2026) places ZenBrain in its related work.
The field measures the wrong thing (Lamperouge, manifesto, Zenodo, August 2026) selects ZenBrain among six works on persistent internal state and notes that its "days" are simulation steps.
Each work with the sentence that cites ZenBrain: zensation.ai/en/publikationen.
Citing ZenBrain or building on it? Open an issue and we will add you here.
Community
GitHub Discussions: Ask a question, show what you built — help, show-and-tell, feature requests
GitHub Issues: Bug reports & feature requests
Email: open-source@zensation.ai
License
Apache 2.0 — use it in production, modify it, distribute it. Just keep the attribution.
Available Tools
4 toolszenbrain_consolidateConsolidate memoryADestructive
Run one consolidation pass: among the 100 most recent episodes, each one with an emotional weight above 0.5 becomes a semantic fact — once, however often the pass runs. Working-memory slots lose relevance with age, and slots whose relevance has dropped to 0.01 or below are removed; that is the only deletion, and nothing in long-term memory is deleted. Run it between sessions or after a batch of zenbrain_store calls, not on every turn; compare zenbrain_health before and after to see the effect.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| pruned | Yes | Always 0: consolidation deletes nothing from long-term memory. |
| decayed | Yes | Working-memory slots decayed. |
| promoted | Yes | Episodes promoted to semantic facts in this pass. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only say destructiveHint=true and idempotentHint=false; the description goes well beyond by scoping the deletion precisely (working-memory slots with relevance <= 0.01 are removed, that is the ONLY deletion, long-term memory is untouched) and clarifying dedup behavior ('once, however often the pass runs'). This is exactly the context needed to trust a destructive-hinted tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the action and mechanism, and every clause carries information (thresholds, deletion scope, timing). Slightly run-on across three long sentences, but there is no filler to cut.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be explained, and annotations cover the safety profile. The description supplies everything else an agent needs: trigger timing, thresholds, exact deletion scope, and a verification workflow via zenbrain_health.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline of 4 applies. The description instead documents implicit inputs (the 100 most recent episodes, the 0.5 weight threshold, the 0.01 relevance cutoff) that would otherwise be invisible to the caller.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The first clause gives a specific verb (run a consolidation pass) and resource (memory), then immediately defines the mechanism: recent episodes with emotional weight > 0.5 become semantic facts. An agent can distinguish this from zenbrain_store/recall/health without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use ('between sessions or after a batch of zenbrain_store calls'), explicit when-not ('not on every turn'), and it recommends comparing zenbrain_health before and after. Both the trigger condition and the sibling to use for verification are named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zenbrain_healthInspect memory stateARead-only
Report how full the memory layers are: working-memory slots in use, interactions held, episodes, facts and how many are due for review, procedures, core blocks. Read-only and cheap, safe to call at any time. It returns counts only; to read the memories themselves, use zenbrain_recall.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| core | Yes | Core memory blocks. |
| working | Yes | Working-memory slots in use and the slot limit. |
| episodic | Yes | Episodes stored. |
| semantic | Yes | Facts stored, and how many are due for spaced-repetition review. |
| shortTerm | Yes | Interactions held in short-term memory. |
| procedural | Yes | Procedures stored. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is partly covered; the description adds cost/performance context ('cheap, safe to call at any time') and states the output shape ('counts only'). It does not add detail on latency budgets or pagination, but none is relevant for a zero-parameter count report.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences: the first front-loads what is measured, the second the cost/safety profile, the third the alternative. Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values needn't be described; the description nonetheless sets expectations ('counts only'). For a zero-parameter diagnostic tool with full annotation coverage, nothing an agent needs is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes no parameters, so the baseline is 4. There is nothing for the description to disambiguate, and it correctly spends no words on parameter semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Name and description pair a specific verb ('report how full') with the exact resources measured: working-memory slots, interactions, episodes, due-for-review facts, procedures, core blocks. It also distinguishes itself from the sibling zenbrain_recall by stating it returns counts, not contents.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly frames the tool as a cheap, read-only, always-safe call and names the alternative (zenbrain_recall) for the adjacent need of reading actual memories. Both the when-to-use and the when-to-use-something-else conditions are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zenbrain_recallRecall memoriesARead-only
Search long-term memory for anything relevant to a query. By default it searches the episodic, semantic, procedural and core layers; working memory (a handful of recently stored items, held only while the server runs) is searched only when named in layers. Returns results ranked by relevance, each tagged with the layer it came from. Use this before answering when the user refers to something from an earlier session. To see how much is stored rather than what, use zenbrain_health; to save something, use zenbrain_store.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | How many results to return at most, after ranking (default 10). | |
| query | Yes | What to look for, in plain language. | |
| layers | No | Restrict the search to these layers. Defaults to episodic, semantic, procedural and core; add 'working' to include recently stored items held in memory. | |
| minConfidence | No | Leave out results whose confidence is below this value. Applied before the final ranking, so fewer than `limit` results may come back. Facts use their stored confidence, procedures their success rate; core blocks, episodes and working-memory items count as 1 and are kept. |
Output Schema
| Name | Required | Description |
|---|---|---|
| count | Yes | How many memories were returned. |
| results | Yes | Matching memories, most relevant first. |
| skipped | Yes | Rows that carried no readable content and were left out of `results`. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds real context beyond that: working memory is volatile and only lives while the server runs, default layers searched, results are ranked and layer-tagged. It stops short of describing pagination or result shape, but the output schema handles that.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, each earning its place: capability, default layer behavior, return characteristics, then usage and alternatives. Front-loaded with the core verb and scope with no throat-clearing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only search tool with an output schema and full parameter coverage, the description covers purpose, default behavior, volatility caveat, and sibling routing. Nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so all four parameters are documented in the schema itself, and the description largely echoes the layers default. Baseline 3 is appropriate; the description adds minimal meaning beyond what the schema already states about layers and limit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (search) and resource (long-term memory) with clear scope: relevance-ranked recall across named layers. An agent can tell it apart from zenbrain_store and zenbrain_health without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use it ('before answering when the user refers to something from an earlier session') and names both alternatives with their own purposes ('zenbrain_health' for how much is stored, 'zenbrain_store' to save). Routing is fully determined.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
zenbrain_storeStore a memoryA
Write something into long-term memory so it survives this conversation. Routing is automatic by default: steps or instructions become a procedure; content with an emotional weight above 0.5 (detected, or set via emotionalWeight) becomes an episode; a confidence above 0.9 makes it a pinned core memory; anything else becomes a semantic fact. Set type only when you want to override that. Every call adds a new memory, except that storing the same core memory again updates it; the content is also kept in working memory while the server runs. Returns the id of the stored memory. If it may already be stored, check with zenbrain_recall first: storing it again adds a duplicate.
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | Routing hint. 'auto' (default) decides from the content. | |
| steps | No | Ordered steps of a procedure. Optional: without them the steps are taken from the content (numbered or bulleted lines, otherwise every line). | |
| tools | No | Tools a procedure uses. | |
| source | No | Where this came from, e.g. 'user', 'ai', 'import'. | |
| content | Yes | The memory to store, in plain language. | |
| context | No | Context domain, e.g. 'work', 'personal', 'learning'. | |
| outcome | No | What the procedure achieves. | |
| confidence | No | How certain this is (0–1). Above 0.9 routes to core memory. | |
| emotionalWeight | No | Emotional significance (0–1). Detected from the content when omitted. Above 0.5 routes to episodic memory. |
Output Schema
| Name | Required | Description |
|---|---|---|
| id | Yes | Identifier of the stored memory. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare a non-read-only, non-idempotent write, and the description goes well beyond that: it discloses that every call creates a new memory except duplicate core memories (which update), that content is mirrored into working memory while the server runs, that duplicates are possible, and what the call returns. These are exactly the mutation/dedup traits annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core purpose, then routing, then the dedup caveat and return value. Dense and mostly waste-free, though the routing sentence is long and re-states thresholds the schema already carries.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter write tool with an output schema, the description covers the decision logic, mutation semantics, dedup risk, and return value. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds value by tying parameters into composite routing logic (the interaction of `type`, `confidence`, and `emotionalWeight`) and by clarifying that `type` is an override rather than a requirement — the schema documents the thresholds individually but not the routing precedence.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Write something into long-term memory') and immediately distinguishes itself from the read-side sibling by naming zenbrain_recall as the dedup check. An agent can tell what this tool does without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit routing rules ('steps or instructions become a procedure; emotional weight above 0.5 becomes an episode; confidence above 0.9 becomes core'), an explicit override condition for `type`, and a when-to-use-otherwise instruction ('If it may already be stored, check with zenbrain_recall first'). This is a near-complete decision procedure.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.4.10- Changed
zenbrain_recall1 field changed- changed
Input schema / properties / minConfidence / descriptionPrevious value: -"Leave out results whose stored confidence is below this value; results stored without a confidence count as 1 and are kept."New value: +"Leave out results whose confidence is below this value. Applied before the final ranking, so fewer than `limit` results may come back. Facts use their stored confidence, procedures their success rate; core blocks, episodes and working-memory items count as 1 and are kept."
2 tool updates
v0.4.9- Changed
zenbrain_recall4 fields changed- removed
Input schema / properties / includeContextRemoved value: -{ - "description": "Boost results whose stored context (time of day, weekday, task type) matches the current one.", - "type": "boolean" -} - changed
Input schema / properties / limit / descriptionPrevious value: -"Maximum results (default 10)."New value: +"How many results to return at most, after ranking (default 10)." - changed
Input schema / properties / minConfidence / descriptionPrevious value: -"Drop results below this confidence."New value: +"Leave out results whose stored confidence is below this value; results stored without a confidence count as 1 and are kept." - removed
Input schema / properties / taskTypeRemoved value: -{ - "description": "Current task, e.g. 'coding', 'writing'. Only has an effect together with `includeContext: true`.", - "type": "string" -}
- Changed
zenbrain_store1 field changed- changed
Input schema / properties / steps / descriptionPrevious value: -"Ordered steps. Required when type is 'procedure'."New value: +"Ordered steps of a procedure. Optional: without them the steps are taken from the content (numbered or bulleted lines, otherwise every line)."
4 tool updates
v0.1.7- Changed
zenbrain_consolidate2 fields changed- changed
Output schema / properties / promoted / descriptionPrevious value: -"Episodes promoted to semantic facts."New value: +"Episodes promoted to semantic facts in this pass." - changed
Output schema / properties / pruned / descriptionPrevious value: -"Items pruned below the retention threshold."New value: +"Always 0: consolidation deletes nothing from long-term memory."
- Changed
zenbrain_health1 field changed- changed
Output schema / (root)Previous value: -nullNew value: +{ + "$schema": "http://json-schema.org/draft-07/schema#", + "additionalProperties": false, + "properties": { + "core": { + "additionalProperties": false, + "description": "Core memory blocks.", + "properties": { + "blocks": { + "type": "number" + } + }, + "required": [ + "blocks" + ], + "type": "object" + }, + "episodic": { + "additionalProperties": false, + "description": "Episodes stored.", + "properties": { + "count": { + "type": "number" + } + }, + "required": [ + "count" + ], + "type": "object" + }, + "procedural": { + "additionalProperties": false, + "description": "Procedures stored.", + "properties": { + "count": { + "type": "number" + } + }, + "required": [ + "count" + ], + "type": "object" + }, + "semantic": { + "additionalProperties": false, + "description": "Facts stored, and how many are due for spaced-repetition review.", + "properties": { + "count": { + "type": "number" + }, + "dueForReview": { + "type": "number" + } + }, + "required": [ + "count", + "dueForReview" + ], + "type": "object" + }, + "shortTerm": { + "additionalProperties": false, + "description": "Interactions held in short-term memory.", + "properties": { + "interactions": { + "type": "number" + } + }, + "required": [ + "interactions" + ], + "type": "object" + }, + "working": { + "additionalProperties": false, + "description": "Working-memory slots in use and the slot limit.", + "properties": { + "max": { + "type": "number" + }, + "used": { + "type": "number" + } + }, + "required": [ + "used", + "max" + ], + "type": "object" + } + }, + "required": [ + "working", + "shortTerm", + "episodic", + "semantic", + "procedural", + "core" + ], + "type": "object" +}
- Changed
zenbrain_recall3 fields changed- changed
Input schema / properties / includeContext / descriptionPrevious value: -"Boost results matching the current context."New value: +"Boost results whose stored context (time of day, weekday, task type) matches the current one." - changed
Input schema / properties / layers / descriptionPrevious value: -"Restrict the search to these layers. Defaults to all but working."New value: +"Restrict the search to these layers. Defaults to episodic, semantic, procedural and core; add 'working' to include recently stored items held in memory." - changed
Input schema / properties / taskType / descriptionPrevious value: -"Current task, e.g. 'coding', 'writing' — used for context matching."New value: +"Current task, e.g. 'coding', 'writing'. Only has an effect together with `includeContext: true`."
- Changed
zenbrain_store1 field changed- changed
Input schema / properties / emotionalWeight / descriptionPrevious value: -"Emotional significance (0–1). Detected from the content when omitted."New value: +"Emotional significance (0–1). Detected from the content when omitted. Above 0.5 routes to episodic memory."
4 tool updates
v0.1.5- First observed
zenbrain_consolidate - First observed
zenbrain_health - First observed
zenbrain_recall - First observed
zenbrain_store
TDQS
Scored across 4 tools
Each tool has a clearly distinct role: recall reads, store writes, consolidate maintains, and health reports stats. Descriptions explicitly cross-reference to avoid confusion (e.g., 'to see how much is stored rather than what, use zenbrain_health'), and no overlap exists among the four operations.
All tools share the zenbrain_ prefix and use snake_case, forming a predictable pattern. However, the action words mix verbs (recall, store, consolidate) with a noun (health), a minor deviation from a pure verb-based convention.
Four tools cover the core memory lifecycle: write, read, maintain, and monitor. The count is well within the 3-15 well-scoped range, and each tool clearly earns its place without redundancy.
The surface covers storing, recalling, consolidating, and monitoring memory layers. Minor gaps exist: no explicit update or delete for non-core memories, though storing anew works around this and deletion is intentionally absent.
Maintenance
Related MCP Connectors
InfoLang semantic memory MCP — investigate, memorize, and recall compressed agent context.
- memnodeOAuthdev.memnode
Persistent, inspectable memory for AI agents with lineage, correction, and a hosted MCP endpoint.
- KogniteOAuthdev.kognite
Hosted agent memory: store, search, and recall facts across sessions from any MCP client.
Analytical memory for AI agents: a real Postgres queried in plain English over MCP. One command.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceA local memory engine for AI agents. Stores conversation episodes, consolidates knowledge through a neuroscience-inspired lifecycle, and builds a personal knowledge graph — all in a local SQLite database.17MIT
- AlicenseNot gradedqualityCmaintenanceLocal-first persistent memory layer for AI agents. Provides hybrid search (FTS5 keyword + vector embeddings) over 223K+ knowledge chunks via MCP. Tools: brain_search, brain_store, brain_entity, brain_subscribe. Features pub/sub with stable agent identity, delivery tracking, and Claude --channels integration. SQLite + BrainBar Swift daemon on Unix socket.2,361 PyPI9Apache 2.0
- AlicenseNot gradedqualityCmaintenanceLocal-first semantic memory layer for MCP agents. Recall, remember, forget, stats over stdio. ChromaDB plus sentence-transformers, all on-machine.69 PyPIMIT
- AlicenseNot gradedqualityAmaintenanceProvides persistent, cooperative memory for LLMs via MCP, with SQLite storage and tools for capturing, recalling, consolidating, crystallizing, and forgetting memories across sessions.3MIT