Skip to main content
Glama

@misc{bering2026zenbrain,
  title         = {ZenBrain: A Neuroscience-Inspired 7-Layer Memory Architecture for Autonomous AI Systems},
  author        = {Bering, Alexander},
  year          = {2026},
  eprint        = {2604.23878},
  archivePrefix = {arXiv},
  primaryClass  = {cs.AI},
  doi           = {10.5281/zenodo.19353663},
  url           = {https://arxiv.org/abs/2604.23878}
}

Feedback, replications, and counter-results are explicitly welcome — please open an issue or reach out via research@zensation.ai.

Your AI forgets everything after every conversation. ZenBrain fixes that — with the same mechanisms your brain uses: spaced repetition, emotional consolidation, Hebbian strengthening, and exponential forgetting curves. Not a vector database with a wrapper. Actual neuroscience.

Architecture vs. this package. ZenBrain's architecture is 15 neuroscience-inspired mechanisms — 9 foundational algorithms + 6 Predictive Memory Architecture (PMA) components (paper). The 6 PMA components are proprietary and run in the production system. This open-source package ships the algorithm library: 10 core algorithms + 10 advanced research modules (20 modules), zero-dependency.


Benchmark: LongMemEval-500

On LongMemEval-500, three of nine head-to-head answer-quality comparisons hold against Letta, Mem0 and A-Mem — all three against A-Mem, the remaining six are ties, none lost. Three competitors x three LLM judges, Bonferroni-corrected (alpha = 0.05/18) and version-matched. It reaches 91.3% of a full-context oracle's binary-judge accuracy at 1/109.6 of the per-query token cost (47.7% vs. 52.2%).

Works out of the box without an embedding provider — lexical ranking, zero dependencies. With nomic-embed-text as the embedding provider you get the configuration those figures were measured in.

The paper prints where ZenBrain loses as well: on LoCoMo, substring-based aggregate F1 favours lexical retrieval (BM25) by metric design, and we do not contest that. The advantage is most pronounced on judge-graded answer quality and cross-session reasoning.

The mechanism comparison further down re-runs from this repository in under a minute — bash scripts/compare-mechanisms.sh, no API keys and nothing to install. It prints a positive and a negative control before the result, so the instrument can be checked before its output is trusted. The method, the effect sizes and the ablations behind the numbers above are in the paper; this repository ships no runner for them. The archived packages below run the significance tests and effect sizes in full, and the mechanism ablation behind paper Tables 7–9.


Related MCP server: brainlayer

Reproduction packages

The raw material behind the numbers above is deposited on Zenodo, open access and citable. Both links are concept DOIs and resolve to the latest version, the same convention this README uses for the paper archive; each description was measured against the version named after it.

  • Mechanism ablation, paper Tables 7–9 — 10.5281/zenodo.22162063 (described here: v1.0.0). Four experiment suites (95 tests), the reference JSON the paper's tables were generated from, and verify-against-reference.mjs, which diffs a fresh run against that reference and exits non-zero on drift. npm install && npm run experiments; the run itself needs no API keys and no network, and finishes in under a minute on a laptop. Two of the paper's other ablation tables need data this package does not carry: Table 11 the LoCoMo corpus, Table 13 a different pipeline. The package says so itself.

  • Measurement package, LongMemEval-500 and the real-pipeline flag ablation — 10.5281/zenodo.22161977 (described here: v1). Per-(system, judge, seed) judged outputs, the flag manifests as recorded at run time, a SHA256SUMS.txt covering every file in the package, and the analysis scripts. Three of those scripts are stdlib-only and self-checking — the oracle comparison behind the 91.3% figure, the judge-agreement figures, and the real-pipeline flag-ablation table: each prints every re-derived value next to the reference it has to match, and exits non-zero on mismatch. The significance tests behind the nine head-to-head comparisons against Letta, Mem0 and A-Mem sit in a separate script that needs numpy and scipy; it recomputes all eighteen pairwise tests and rewrites the deposited significance JSON byte-identically, so what catches a mismatch there is the checksum, not an exit code.

Both packages name what they do not cover. Replications and counter-results are welcome: research@zensation.ai.


How ZenBrain differs from Mem0, Letta and Zep

ZenBrain implements fifteen mechanisms taken from human memory research. No system among those surveyed in the paper integrates more than two of them. The table below records which of the mechanisms appear in the public source of three widely used memory systems, at pinned versions, on a fixed date.

Mechanism

ZenBrain

Mem0

Letta

Zep

FSRS spaced repetition

yes

—

—

—

Hebbian learning

yes

—

—

—

Ebbinghaus forgetting curves

yes

—

—

—

Sleep consolidation

yes

—

—

—

Emotional tagging

yes

—

—

—

Zero runtime dependencies

yes

—

—

—

How this was measured, 27 August 2026. Full-text search over the checked-out public source of mem0ai/mem0 (npm mem0ai 3.1.7, PyPI mem0ai 2.0.19), letta-ai/letta-code (npm @letta-ai/letta-code 0.31.2) and getzep/zep, lockfiles excluded. A dash means the term does not occur in that snapshot — not that the system cannot do something comparable under another name. Dependency counts are declared direct dependencies: @zensation/core resolves to two packages, both our own; mem0ai declares four, @letta-ai/letta-code eighteen. Re-run the whole check yourself with scripts/compare-mechanisms.sh; it prints its own positive and negative controls so you can see the instrument works before you trust the result.

Human memory does not work like a key-value store. The brain keeps specialised systems for different kinds of memory, forgets actively, modulates by emotion and retrieves by context. ZenBrain brings those mechanisms to AI agents.

Advanced algorithms (since v0.3.0, May 2026)

On top of the 10 core algorithms above, @zensation/algorithms ships 10 advanced algorithms grounded in recent neuroscience and ML research. Each is exposed as its own sub-path (@zensation/algorithms/<name>) and remains zero-dependency:

  • fsrs-vmPFC — Prediction-Error coupled FSRS

  • hebbian-two-factor — Two-Factor synaptic consolidation

  • sleep-simulation-selection — RL-based replay selection

  • spectral-health — Fiedler-value KG health monitor

  • ib-budget — Information-Bottleneck retention budget

  • dopamine-routing · hopfield-stm · personalized-pagerank · surprise-gradient-memory · temporal-multi-route

See CHANGELOG.md for details.

Quick Start

Requires Node.js 22 or newer (since 0.4.0). On Node 20 or older, npm silently installs the last compatible release (@zensation/algorithms@0.3.4, @zensation/core@0.2.2) instead of the current one — which looks like a broken package but is a platform mismatch. See CHANGELOG.

npm install @zensation/algorithms
import {
  initFromDecayClass,
  getRetrievability,
  updateAfterRecall,
  tagEmotion,
  computeEmotionalWeight,
  computeHebbianStrengthening,
  propagateForRelation,
} from '@zensation/algorithms';

// 1. Schedule a memory with FSRS
const memory = initFromDecayClass('normal_decay');

// 2. A week later, check recall probability (Ebbinghaus curve)
const aWeekLater = new Date(Date.now() + 7 * 24 * 60 * 60 * 1000);
const retention = getRetrievability(memory, aWeekLater);
console.log(`Recall probability: ${(retention * 100).toFixed(1)}%`);
// ~36.8% — retrievability has decayed over the week

// 3. User recalled it anyway — update scheduling
const updated = updateAfterRecall(memory, 4, retention, aWeekLater);
// stability 7 -> 8.19: recalling at low retrievability gives a bigger boost

// 4. Tag emotional significance
const emotion = tagEmotion('I am absolutely thrilled — I got the promotion!');
const weight = computeEmotionalWeight(emotion);
console.log(`Decay multiplier: ${weight.decayMultiplier}x`);
// 2.7x — emotional memories decay nearly 3x slower

// 5. Strengthen knowledge connections (Hebbian)
const stronger = computeHebbianStrengthening(1.0);
// 1.09 — "neurons that fire together wire together"

// 6. Propagate confidence through your knowledge graph
const confidence = propagateForRelation(0.5, 0.8, 1.0, 'supports');
// 0.9 — supporting evidence increases confidence

Want the advanced algorithms?

import {
  computeKGPredictionError,
  computeAdaptiveFSRSInterval,
} from '@zensation/algorithms/fsrs-vmPFC';

// Couple FSRS scheduling with the prediction-error signal from your
// knowledge graph: when the embedding has shifted a lot since the last
// review (high cosine distance), shrink the next interval; otherwise push
// it out. Both arrays must have the same length.
const lastEmbedding = [0.1, 0.2, 0.3, 0.4];
const currentEmbedding = [0.5, 0.4, 0.1, 0.2];
const pe = computeKGPredictionError(lastEmbedding, currentEmbedding);
const nextInterval = computeAdaptiveFSRSInterval(14, pe);

Each advanced algorithm has its own sub-path (@zensation/algorithms/spectral-health, @zensation/algorithms/ib-budget, …). All zero dependencies.

Runnable examples

Seven self-contained examples live in examples/:

Example

Shows

basic-chatbot.ts

Working Memory + Short-Term Memory for conversation context — no SDK needed

with-claude.ts

An Anthropic Claude assistant that remembers across conversations

with-langchain.ts

ZenBrain as the memory backend of a LangChain agent

with-crewai.ts

Multiple agents sharing Working Memory, with Hebbian strengthening

with-vercel-ai.ts

A memory-aware system prompt for the Vercel AI SDK streamText pattern

with-llamaindex.ts

ZenBrain as long-term memory for a LlamaIndex.TS agent — LlamaIndex keeps the short-term window, ZenBrain decides what outlives it

with-mastra.ts

ZenBrain behind a Mastra agent: a Processor records each turn, and the instructions are rebuilt from what is currently recallable

npx tsx examples/basic-chatbot.ts

The integration examples additionally need their respective SDK installed; examples/README.md lists which. Open good first issues are the best way in if you want to contribute.

The Science Behind It

7-Layer Memory Architecture

Layer 7: Cross-Context Memory    ← Shared knowledge across domains
Layer 6: Core Memory             ← Pinned facts (Letta-style)
Layer 5: Procedural Memory       ← "How to do X" (skills & workflows)
Layer 4: Long-Term Semantic      ← Facts with FSRS scheduling
Layer 3: Episodic Memory         ← Concrete experiences & events
Layer 2: Short-Term / Session    ← Current conversation context
Layer 1: Working Memory          ← Active task focus (7±2 items)

Each layer has different retention characteristics, consolidation rules, and retrieval mechanisms — just like the human brain.

FSRS Spaced Repetition

FSRS (Free Spaced Repetition Scheduler) outperforms SM-2 by 30%. It uses the desirable difficulty principle: reviewing when retention is low gives a bigger stability boost. Your AI reviews important facts at optimal intervals — never too early (wasteful), never too late (forgotten).

Emotional Memory

The amygdala modulates memory consolidation — emotional events are remembered more vividly (flashbulb memory). ZenBrain's emotional tagger assigns arousal, valence, and significance scores using a 400+ keyword lexicon (English & German). Emotional memories get up to 3x longer decay half-lives.

Hebbian Learning

"Neurons that fire together wire together" (Hebb, 1949). Knowledge graph edges that are frequently co-activated grow stronger. Unused edges decay and eventually get pruned. The result: a self-organizing knowledge structure that reflects actual usage patterns, with homeostatic normalization to prevent runaway growth.

Ebbinghaus Forgetting Curves

Ebbinghaus (1885) showed that memory decays exponentially: R = e^(-t/S). ZenBrain implements personalized decay profiles that adapt to individual learning patterns, with SM-2 compatibility for existing spaced repetition systems.

Context-Dependent Retrieval

Tulving's Encoding Specificity Principle (1973): memories are recalled better when the retrieval context matches the encoding context. ZenBrain captures temporal context (time of day, day of week) and task type at encoding time, providing up to a 30% retrieval boost when contexts match.

Bayesian Confidence Propagation

Knowledge isn't isolated — facts support or contradict each other. ZenBrain propagates confidence through your knowledge graph using Bayesian belief updates: supporting evidence increases confidence, contradictions decrease it, with damping for numerical stability.

Sleep Consolidation

During sleep, the hippocampus replays recent experiences, strengthening important memories and pruning weak connections (Stickgold & Walker, 2013). ZenBrain simulates this process: selectForReplay() prioritizes emotional and recently-accessed memories, simulateReplay() boosts their stability by 50%, and pruneWeakConnections() removes weak Hebbian edges — implementing the Synaptic Homeostasis Hypothesis (Tononi & Cirelli, 2006).

import { selectForReplay, simulateReplay } from '@zensation/algorithms/sleep-consolidation';

// Select memories for overnight consolidation
const toReplay = selectForReplay(allMemories);
// Simulate sleep replay — stability ↑, weak edges pruned
const result = simulateReplay(toReplay);
console.log(`Replayed ${result.summary.totalReplayed} memories, avg stability +${result.summary.avgStabilityIncrease.toFixed(1)} days`);

Memory Coordinator

The MemoryCoordinator orchestrates all 7 layers into a single cohesive system — inspired by Global Workspace Theory (Baars, 1988):

import { MemoryCoordinator } from '@zensation/core';

const memory = new MemoryCoordinator({ storage: adapter, embedding: embedder });

// Auto-routes to the right layer (semantic, episodic, procedural, or core)
await memory.store('User prefers TypeScript', { type: 'auto' });

// Cross-layer search with ranked, deduplicated results
const results = await memory.recall('programming preferences');

// Consolidate: promote episodic → semantic, apply decay
await memory.consolidate();

// FSRS review queue across all layers
const dueItems = await memory.getReviewQueue();

Packages

Package

Description

Status

@zensation/algorithms · source

20 algorithm modules — 10 core (FSRS, Hebbian, Ebbinghaus, emotional, Bayesian, sleep consolidation, intervals, visualization) + 10 advanced (vmPFC-FSRS, two-factor Hebbian, IB budget, Hopfield STM, …)

:white_check_mark: Published

@zensation/core · source

Memory layers, coordinator, adapter interfaces

:white_check_mark: Published

@zensation/adapter-postgres · source

PostgreSQL + pgvector storage adapter

:white_check_mark: Published

@zensation/adapter-sqlite · source

SQLite storage adapter (zero-config)

:white_check_mark: Published

@zensation/mcp · source

MCP server — gives any MCP client (Claude Desktop, Claude Code, Cursor) the seven layers as four tools. Carries the protocol SDK, so the core stays dependency-free

:white_check_mark: Published

@zensation/ai-sdk · source

Vercel AI SDK middleware — recall before the model call, store after it. Works with any provider, zero runtime dependencies

:white_check_mark: Published

Tree-Shakeable Imports

Every algorithm is available as a subpath export:

// Import everything
import { tagEmotion, updateAfterRecall } from '@zensation/algorithms';

// Or just what you need (better tree-shaking)
import { updateAfterRecall } from '@zensation/algorithms/fsrs';
import { tagEmotion } from '@zensation/algorithms/emotional';
import { computeHebbianStrengthening } from '@zensation/algorithms/hebbian';
import { propagateForRelation } from '@zensation/algorithms/bayesian';
import { selectForReplay } from '@zensation/algorithms/sleep-consolidation';
import { getRetrievabilityWithCI } from '@zensation/algorithms/intervals';
import { generateRetentionCurve } from '@zensation/algorithms/visualization';

Use Cases

AI Chatbots with Long-Term Memory

import { updateAfterRecall, getRetrievability, scheduleNextReview } from '@zensation/algorithms/fsrs';
import { tagEmotion, computeEmotionalWeight } from '@zensation/algorithms/emotional';

// When your AI learns a fact about the user:
function rememberFact(fact: string) {
  const memory = initFromDecayClass('normal_decay');
  const emotion = tagEmotion(fact);
  const weight = computeEmotionalWeight(emotion);

  // Emotional facts get longer retention
  return {
    ...memory,
    emotionalWeight: weight.consolidationWeight,
    decayMultiplier: weight.decayMultiplier,
  };
}

// Before each conversation, check what needs reinforcement:
function getFactsDueForReview(facts: MemoryState[]) {
  return facts.filter(f => getRetrievability(f) < 0.7);
}

Knowledge Graph with Self-Organizing Edges

import { computeHebbianStrengthening, computeHebbianDecay } from '@zensation/algorithms/hebbian';
import { propagateForRelation } from '@zensation/algorithms/bayesian';

// When two concepts are mentioned together:
function coActivate(edge: { weight: number }) {
  edge.weight = computeHebbianStrengthening(edge.weight);
}

// Periodic maintenance — decay unused edges:
function decayEdges(edges: { weight: number; lastUsed: Date }[]) {
  for (const edge of edges) {
    edge.weight = computeHebbianDecay(edge.weight);
    // Edges below MIN_WEIGHT (0.1) can be pruned
  }
}

RAG with Confidence Scoring

import { propagateForRelation, isSignificantChange } from '@zensation/algorithms/bayesian';

// After retrieval, propagate confidence through related facts:
function updateConfidenceGraph(facts: Fact[], relations: Relation[]) {
  for (const rel of relations) {
    const newConf = propagateForRelation(
      rel.target.confidence,
      rel.source.confidence,
      rel.weight,
      rel.type // 'supports' | 'contradicts' | 'related_to'
    );
    if (isSignificantChange(newConf, rel.target.confidence)) {
      rel.target.confidence = newConf;
    }
  }
}

Extracted From Production

These aren't toy implementations — ZenBrain's algorithms are extracted from ZenAI, a production AI platform. Everything claimed here is verifiable in this repository:

  • 801 tests (489 algorithms · 137 core · 54 adapter-sqlite · 48 adapter-postgres · 55 mcp · 18 ai-sdk), all passing

  • Zero runtime dependencies — pure TypeScript, dual ESM + CJS, tree-shakeable subpath exports

  • Reproducible — building from this source produces the same 153-file @zensation/algorithms@0.5.0 tarball published on npm

  • 7-layer memory architecture grounded in published neuroscience

Contributing

We welcome contributions! See CONTRIBUTING.md for guidelines. Issues and pull requests are answered; replications and counter-results are especially welcome.

Resources: API Reference | Recipes | Architecture | Benchmarks | FAQ | Roadmap

# Clone the repo
git clone https://github.com/zensation-ai/zenbrain.git
cd zenbrain

# Install dependencies
npm install

# Run tests
npm test

# Build all packages
npm run build

Research

ZenBrain's architecture and algorithms are documented in an open-access technical disclosure:

If you use ZenBrain in academic work, please cite:

@misc{bering2026zenbrain,
  title         = {ZenBrain: A Neuroscience-Inspired 7-Layer Memory Architecture for Autonomous AI Systems},
  author        = {Bering, Alexander},
  year          = {2026},
  eprint        = {2604.23878},
  archivePrefix = {arXiv},
  primaryClass  = {cs.AI},
  doi           = {10.5281/zenodo.19353663},
  url           = {https://arxiv.org/abs/2604.23878}
}

Cited by

Five works cite ZenBrain (as of 28 September 2026): a survey by the Sico team at Microsoft Research, MindMemOS by Huawei's Noah's Ark Lab, the EVD preprint, the HindsightTag manuscript and the Lamperouge manifesto.

  • Agentic Evolution: From Self-Improving Agents to Co-Evolving Human-AI Systems, a survey by the Sico team at Microsoft Research (2026), calls forgetting "critically understudied": very few of the roughly 100 Memory & Sense papers it surveys implement explicit forgetting. It lists ZenBrain as one of four works in its forgetting/lifecycle group and as its only source for sleep-based consolidation.

  • MindMemOS (Huawei's Noah's Ark Lab, arXiv:2608.12428, August 2026) relies on exactly three references for the foundational claim of its introduction: two field surveys and a single system paper — ZenBrain.

  • HindsightTag (Vivek Govindbhai Dudhat, manuscript, July 2026) discusses ZenBrain as its closest prior system and uses its MemoryCoordinator interface as the integration example.

  • EVD: An Emotional Valence Dimension for Persistent Agent Memory (R. J. Vandelinder and I. Vandelinder, Exile Research, Zenodo, June 2026) places ZenBrain in its related work.

  • The field measures the wrong thing (Lamperouge, manifesto, Zenodo, August 2026) selects ZenBrain among six works on persistent internal state and notes that its "days" are simulation steps.

Each work with the sentence that cites ZenBrain: zensation.ai/en/publikationen.

Citing ZenBrain or building on it? Open an issue and we will add you here.

Community

License

Apache 2.0 — use it in production, modify it, distribute it. Just keep the attribution.


Available Tools

4 tools
zenbrain_consolidateConsolidate memoryA
Destructive

Run one consolidation pass: among the 100 most recent episodes, each one with an emotional weight above 0.5 becomes a semantic fact — once, however often the pass runs. Working-memory slots lose relevance with age, and slots whose relevance has dropped to 0.01 or below are removed; that is the only deletion, and nothing in long-term memory is deleted. Run it between sessions or after a batch of zenbrain_store calls, not on every turn; compare zenbrain_health before and after to see the effect.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
prunedYesAlways 0: consolidation deletes nothing from long-term memory.
decayedYesWorking-memory slots decayed.
promotedYesEpisodes promoted to semantic facts in this pass.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only say destructiveHint=true and idempotentHint=false; the description goes well beyond by scoping the deletion precisely (working-memory slots with relevance <= 0.01 are removed, that is the ONLY deletion, long-term memory is untouched) and clarifying dedup behavior ('once, however often the pass runs'). This is exactly the context needed to trust a destructive-hinted tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action and mechanism, and every clause carries information (thresholds, deletion scope, timing). Slightly run-on across three long sentences, but there is no filler to cut.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained, and annotations cover the safety profile. The description supplies everything else an agent needs: trigger timing, thresholds, exact deletion scope, and a verification workflow via zenbrain_health.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline of 4 applies. The description instead documents implicit inputs (the 100 most recent episodes, the 0.5 weight threshold, the 0.01 relevance cutoff) that would otherwise be invisible to the caller.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first clause gives a specific verb (run a consolidation pass) and resource (memory), then immediately defines the mechanism: recent episodes with emotional weight > 0.5 become semantic facts. An agent can distinguish this from zenbrain_store/recall/health without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use ('between sessions or after a batch of zenbrain_store calls'), explicit when-not ('not on every turn'), and it recommends comparing zenbrain_health before and after. Both the trigger condition and the sibling to use for verification are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

zenbrain_healthInspect memory stateA
Read-only

Report how full the memory layers are: working-memory slots in use, interactions held, episodes, facts and how many are due for review, procedures, core blocks. Read-only and cheap, safe to call at any time. It returns counts only; to read the memories themselves, use zenbrain_recall.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
coreYesCore memory blocks.
workingYesWorking-memory slots in use and the slot limit.
episodicYesEpisodes stored.
semanticYesFacts stored, and how many are due for spaced-repetition review.
shortTermYesInteractions held in short-term memory.
proceduralYesProcedures stored.

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is partly covered; the description adds cost/performance context ('cheap, safe to call at any time') and states the output shape ('counts only'). It does not add detail on latency budgets or pagination, but none is relevant for a zero-parameter count report.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences: the first front-loads what is measured, the second the cost/safety profile, the third the alternative. Every sentence earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values needn't be described; the description nonetheless sets expectations ('counts only'). For a zero-parameter diagnostic tool with full annotation coverage, nothing an agent needs is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes no parameters, so the baseline is 4. There is nothing for the description to disambiguate, and it correctly spends no words on parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Name and description pair a specific verb ('report how full') with the exact resources measured: working-memory slots, interactions, episodes, due-for-review facts, procedures, core blocks. It also distinguishes itself from the sibling zenbrain_recall by stating it returns counts, not contents.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly frames the tool as a cheap, read-only, always-safe call and names the alternative (zenbrain_recall) for the adjacent need of reading actual memories. Both the when-to-use and the when-to-use-something-else conditions are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

zenbrain_recallRecall memoriesA
Read-only

Search long-term memory for anything relevant to a query. By default it searches the episodic, semantic, procedural and core layers; working memory (a handful of recently stored items, held only while the server runs) is searched only when named in layers. Returns results ranked by relevance, each tagged with the layer it came from. Use this before answering when the user refers to something from an earlier session. To see how much is stored rather than what, use zenbrain_health; to save something, use zenbrain_store.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoHow many results to return at most, after ranking (default 10).
queryYesWhat to look for, in plain language.
layersNoRestrict the search to these layers. Defaults to episodic, semantic, procedural and core; add 'working' to include recently stored items held in memory.
minConfidenceNoLeave out results whose confidence is below this value. Applied before the final ranking, so fewer than `limit` results may come back. Facts use their stored confidence, procedures their success rate; core blocks, episodes and working-memory items count as 1 and are kept.

Output Schema

ParametersJSON Schema
NameRequiredDescription
countYesHow many memories were returned.
resultsYesMatching memories, most relevant first.
skippedYesRows that carried no readable content and were left out of `results`.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true and openWorldHint=false, so safety is covered. The description adds real context beyond that: working memory is volatile and only lives while the server runs, default layers searched, results are ranked and layer-tagged. It stops short of describing pagination or result shape, but the output schema handles that.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, each earning its place: capability, default layer behavior, return characteristics, then usage and alternatives. Front-loaded with the core verb and scope with no throat-clearing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only search tool with an output schema and full parameter coverage, the description covers purpose, default behavior, volatility caveat, and sibling routing. Nothing needed to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so all four parameters are documented in the schema itself, and the description largely echoes the layers default. Baseline 3 is appropriate; the description adds minimal meaning beyond what the schema already states about layers and limit.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (search) and resource (long-term memory) with clear scope: relevance-ranked recall across named layers. An agent can tell it apart from zenbrain_store and zenbrain_health without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to use it ('before answering when the user refers to something from an earlier session') and names both alternatives with their own purposes ('zenbrain_health' for how much is stored, 'zenbrain_store' to save). Routing is fully determined.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

zenbrain_storeStore a memoryA

Write something into long-term memory so it survives this conversation. Routing is automatic by default: steps or instructions become a procedure; content with an emotional weight above 0.5 (detected, or set via emotionalWeight) becomes an episode; a confidence above 0.9 makes it a pinned core memory; anything else becomes a semantic fact. Set type only when you want to override that. Every call adds a new memory, except that storing the same core memory again updates it; the content is also kept in working memory while the server runs. Returns the id of the stored memory. If it may already be stored, check with zenbrain_recall first: storing it again adds a duplicate.

ParametersJSON Schema
NameRequiredDescriptionDefault
typeNoRouting hint. 'auto' (default) decides from the content.
stepsNoOrdered steps of a procedure. Optional: without them the steps are taken from the content (numbered or bulleted lines, otherwise every line).
toolsNoTools a procedure uses.
sourceNoWhere this came from, e.g. 'user', 'ai', 'import'.
contentYesThe memory to store, in plain language.
contextNoContext domain, e.g. 'work', 'personal', 'learning'.
outcomeNoWhat the procedure achieves.
confidenceNoHow certain this is (0–1). Above 0.9 routes to core memory.
emotionalWeightNoEmotional significance (0–1). Detected from the content when omitted. Above 0.5 routes to episodic memory.

Output Schema

ParametersJSON Schema
NameRequiredDescription
idYesIdentifier of the stored memory.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare a non-read-only, non-idempotent write, and the description goes well beyond that: it discloses that every call creates a new memory except duplicate core memories (which update), that content is mirrored into working memory while the server runs, that duplicates are possible, and what the call returns. These are exactly the mutation/dedup traits annotations cannot express.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose, then routing, then the dedup caveat and return value. Dense and mostly waste-free, though the routing sentence is long and re-states thresholds the schema already carries.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter write tool with an output schema, the description covers the decision logic, mutation semantics, dedup risk, and return value. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds value by tying parameters into composite routing logic (the interaction of `type`, `confidence`, and `emotionalWeight`) and by clarifying that `type` is an override rather than a requirement — the schema documents the thresholds individually but not the routing precedence.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Write something into long-term memory') and immediately distinguishes itself from the read-side sibling by naming zenbrain_recall as the dedup check. An agent can tell what this tool does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit routing rules ('steps or instructions become a procedure; emotional weight above 0.5 becomes an episode; confidence above 0.9 becomes core'), an explicit override condition for `type`, and a when-to-use-otherwise instruction ('If it may already be stored, check with zenbrain_recall first'). This is a near-complete decision procedure.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.4.10
    • Changedzenbrain_recall1 field changed
      • changedInput schema / properties / minConfidence / description
        Previous value: -"Leave out results whose stored confidence is below this value; results stored without a confidence count as 1 and are kept."New value: +"Leave out results whose confidence is below this value. Applied before the final ranking, so fewer than `limit` results may come back. Facts use their stored confidence, procedures their success rate; core blocks, episodes and working-memory items count as 1 and are kept."
  2. 2 tool updatesv0.4.9
    • Changedzenbrain_recall4 fields changed
      • removedInput schema / properties / includeContext
        Removed value: -{
        -  "description": "Boost results whose stored context (time of day, weekday, task type) matches the current one.",
        -  "type": "boolean"
        -}
      • changedInput schema / properties / limit / description
        Previous value: -"Maximum results (default 10)."New value: +"How many results to return at most, after ranking (default 10)."
      • changedInput schema / properties / minConfidence / description
        Previous value: -"Drop results below this confidence."New value: +"Leave out results whose stored confidence is below this value; results stored without a confidence count as 1 and are kept."
      • removedInput schema / properties / taskType
        Removed value: -{
        -  "description": "Current task, e.g. 'coding', 'writing'. Only has an effect together with `includeContext: true`.",
        -  "type": "string"
        -}
    • Changedzenbrain_store1 field changed
      • changedInput schema / properties / steps / description
        Previous value: -"Ordered steps. Required when type is 'procedure'."New value: +"Ordered steps of a procedure. Optional: without them the steps are taken from the content (numbered or bulleted lines, otherwise every line)."
  3. 4 tool updatesv0.1.7
    • Changedzenbrain_consolidate2 fields changed
      • changedOutput schema / properties / promoted / description
        Previous value: -"Episodes promoted to semantic facts."New value: +"Episodes promoted to semantic facts in this pass."
      • changedOutput schema / properties / pruned / description
        Previous value: -"Items pruned below the retention threshold."New value: +"Always 0: consolidation deletes nothing from long-term memory."
    • Changedzenbrain_health1 field changed
      • changedOutput schema / (root)
        Previous value: -nullNew value: +{
        +  "$schema": "http://json-schema.org/draft-07/schema#",
        +  "additionalProperties": false,
        +  "properties": {
        +    "core": {
        +      "additionalProperties": false,
        +      "description": "Core memory blocks.",
        +      "properties": {
        +        "blocks": {
        +          "type": "number"
        +        }
        +      },
        +      "required": [
        +        "blocks"
        +      ],
        +      "type": "object"
        +    },
        +    "episodic": {
        +      "additionalProperties": false,
        +      "description": "Episodes stored.",
        +      "properties": {
        +        "count": {
        +          "type": "number"
        +        }
        +      },
        +      "required": [
        +        "count"
        +      ],
        +      "type": "object"
        +    },
        +    "procedural": {
        +      "additionalProperties": false,
        +      "description": "Procedures stored.",
        +      "properties": {
        +        "count": {
        +          "type": "number"
        +        }
        +      },
        +      "required": [
        +        "count"
        +      ],
        +      "type": "object"
        +    },
        +    "semantic": {
        +      "additionalProperties": false,
        +      "description": "Facts stored, and how many are due for spaced-repetition review.",
        +      "properties": {
        +        "count": {
        +          "type": "number"
        +        },
        +        "dueForReview": {
        +          "type": "number"
        +        }
        +      },
        +      "required": [
        +        "count",
        +        "dueForReview"
        +      ],
        +      "type": "object"
        +    },
        +    "shortTerm": {
        +      "additionalProperties": false,
        +      "description": "Interactions held in short-term memory.",
        +      "properties": {
        +        "interactions": {
        +          "type": "number"
        +        }
        +      },
        +      "required": [
        +        "interactions"
        +      ],
        +      "type": "object"
        +    },
        +    "working": {
        +      "additionalProperties": false,
        +      "description": "Working-memory slots in use and the slot limit.",
        +      "properties": {
        +        "max": {
        +          "type": "number"
        +        },
        +        "used": {
        +          "type": "number"
        +        }
        +      },
        +      "required": [
        +        "used",
        +        "max"
        +      ],
        +      "type": "object"
        +    }
        +  },
        +  "required": [
        +    "working",
        +    "shortTerm",
        +    "episodic",
        +    "semantic",
        +    "procedural",
        +    "core"
        +  ],
        +  "type": "object"
        +}
    • Changedzenbrain_recall3 fields changed
      • changedInput schema / properties / includeContext / description
        Previous value: -"Boost results matching the current context."New value: +"Boost results whose stored context (time of day, weekday, task type) matches the current one."
      • changedInput schema / properties / layers / description
        Previous value: -"Restrict the search to these layers. Defaults to all but working."New value: +"Restrict the search to these layers. Defaults to episodic, semantic, procedural and core; add 'working' to include recently stored items held in memory."
      • changedInput schema / properties / taskType / description
        Previous value: -"Current task, e.g. 'coding', 'writing' — used for context matching."New value: +"Current task, e.g. 'coding', 'writing'. Only has an effect together with `includeContext: true`."
    • Changedzenbrain_store1 field changed
      • changedInput schema / properties / emotionalWeight / description
        Previous value: -"Emotional significance (0–1). Detected from the content when omitted."New value: +"Emotional significance (0–1). Detected from the content when omitted. Above 0.5 routes to episodic memory."
  4. 4 tool updatesv0.1.5
    • First observedzenbrain_consolidate
    • First observedzenbrain_health
    • First observedzenbrain_recall
    • First observedzenbrain_store

TDQS

A4.6/5.0

Scored across 4 tools

Disambiguation5/5

Each tool has a clearly distinct role: recall reads, store writes, consolidate maintains, and health reports stats. Descriptions explicitly cross-reference to avoid confusion (e.g., 'to see how much is stored rather than what, use zenbrain_health'), and no overlap exists among the four operations.

Naming Consistency4/5

All tools share the zenbrain_ prefix and use snake_case, forming a predictable pattern. However, the action words mix verbs (recall, store, consolidate) with a noun (health), a minor deviation from a pure verb-based convention.

Tool Count5/5

Four tools cover the core memory lifecycle: write, read, maintain, and monitor. The count is well within the 3-15 well-scoped range, and each tool clearly earns its place without redundancy.

Completeness4/5

The surface covers storing, recalling, consolidating, and monitoring memory layers. Minor gaps exist: no explicit update or delete for non-core memories, though storing anew works around this and deletion is intentionally absent.

Maintenance

ActivityActive
ResponsivenessSlow

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    A local memory engine for AI agents. Stores conversation episodes, consolidates knowledge through a neuroscience-inspired lifecycle, and builds a personal knowledge graph — all in a local SQLite database.
    17
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Local-first persistent memory layer for AI agents. Provides hybrid search (FTS5 keyword + vector embeddings) over 223K+ knowledge chunks via MCP. Tools: brain_search, brain_store, brain_entity, brain_subscribe. Features pub/sub with stable agent identity, delivery tracking, and Claude --channels integration. SQLite + BrainBar Swift daemon on Unix socket.
    2,361 PyPI
    9
    Apache 2.0
  • A
    license
    Not graded
    quality
    A
    maintenance
    Provides persistent, cooperative memory for LLMs via MCP, with SQLite storage and tools for capturing, recalling, consolidating, crystallizing, and forgetting memories across sessions.
    3
    MIT