compound-memory
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@compound-memorysave a memory that this project uses pnpm, not npm"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
compound-memory
English | 简体中文
Local-first shared memory for multiple AI agents — plain Markdown files that compound in value as they are used. Memory lives on your disk as frontmatter-annotated Markdown, gets stronger with every confirmed use, decays into a revivable archive when neglected, and auto-commits to a local git history on every write.
Why
Every agent session starts from zero: preferences get re-asked, project conventions get re-discovered, the same pitfall gets hit twice. compound-memory gives all your agents one shared store:
Local-first — nothing leaves your machine; memories are human-readable Markdown files, not rows in an opaque database.
MCP-native — exactly 5 tools (
memory_write/memory_search/memory_get/memory_link/memory_feedback) as the single read-write boundary; works with any MCP host (Claude Code, ZCode, WorkBuddy, …), plus a full CLI for operations.Compounding — confirmed usage raises confidence, related memories are recalled as neighbors, validation from a different host counts as independent evidence, and distillation merges many raw memories into fewer, denser ones.
Multi-agent by design — a
_sharednamespace everyone reads, plusagent-*private namespaces each host owns; cross-host validation is tracked per host.Optional semantic recall — vector search via sqlite-vec + BGE embeddings, with automatic graceful fallback to pure lexical search when unavailable.
Related MCP server: SharedBrain
Quick Start
For AI agents
Paste this one-liner into your coding agent (Claude Code, Cursor, ZCode, …) and let it do the rest:
Set up compound-memory (https://github.com/chinwe/compound-memory) — a local-first multi-agent shared memory (MCP server + CLI) — on this machine: install it (`uv tool install compound-memory`, or clone the repo and `uv sync --extra dev`), initialize the store (`compound-memory init`, defaults to ~/.agents/memory), register its stdio MCP server in this host's MCP config — command `compound-memory-server` (PyPI install) or `uvx --from compound-memory compound-memory-server`, env `COMPOUND_MEMORY_ROOT=~/.agents/memory` and `COMPOUND_MEMORY_AGENT_ID=agent-<your-host-id>` — then verify by calling `memory_search` and expecting a `{"hits": [...]}` response; if the host needs a restart to load MCP servers, tell me. Host-specific configs and the usage protocol: docs/agent-integration.md in the repo.For humans
1. Install
Python ≥ 3.11. Either route works:
# Route A: clone the repo (uv-managed; same path the MCP config uses)
git clone https://github.com/chinwe/compound-memory.git
cd compound-memory && uv sync --extra dev
# Route B: install from PyPI (no clone needed)
uv tool install compound-memory # or: pip install compound-memory2. Initialize your store
Defaults to ~/.agents/memory; override with the COMPOUND_MEMORY_ROOT env var.
uv run compound-memory init3. Wire it into your MCP host (recommended)
This lets your everyday agents read/write the shared store automatically:
{
"mcpServers": {
"compound-memory": {
"type": "stdio",
"command": "uv",
"args": ["run", "--directory", "<repo>", "compound-memory-server"],
"env": {
"COMPOUND_MEMORY_ROOT": "~/.agents/memory",
"COMPOUND_MEMORY_AGENT_ID": "agent-<your-host-id>"
}
}
}
}Installed from PyPI? Swap command/args for uvx + ["--from", "compound-memory", "compound-memory-server"] — no repo clone needed. Setting COMPOUND_MEMORY_AGENT_ID is strongly recommended: the store then resolves caller identity from the process env, so a model misreporting its identity (or forging someone else's source) is rejected loudly.
Verify: ask your agent to call memory_search (any keyword) — a {"hits": [...]} response means you're connected. Or run uv run compound-memory stats from the CLI.
4. Next step
Inject the usage protocol from skills/compound-memory/SKILL.md into your host (the search → feedback → distill loop), per docs/agent-integration.md §6.
Demo

One full loop: write → search → feedback (with cross-host first-validation bonus) → store health. Real output from v0.4.0, long payloads trimmed:
uv run compound-memory init{ "ok": true, "root": "~/.agents/memory" }Two different hosts each write one stable fact (new memories start at confidence 0.5, uses 0):
uv run compound-memory write \
"Deploy serverless functions on this platform times out at 10s — keep handlers under that budget" \
fact agent-claude --key vercel-timeout{
"id": "20261007_86adf1",
"ns": "_shared",
"type": "fact",
"source": "agent-claude",
"content": "Deploy serverless functions on this platform times out at 10s — keep handlers under that budget",
"confidence": 0.5,
"uses": 0,
"key": "vercel-timeout",
"validated_by": []
...
}uv run compound-memory write \
"User prefers concise replies with tables and code examples" \
fact agent-zcode --key user-styleSearch ranks by score (--explain attaches per-hit ranking components for debugging):
uv run compound-memory search "serverless timeout"[
{
"id": "20261007_86adf1", "score": 1.0292, "similarity": 1.0,
"type": "fact", "source": "agent-claude",
"content": "Deploy serverless functions on this platform times out at 10s — keep handlers under that budget",
"neighbors": []
},
{
"id": "20261007_6a0c0c", "score": 0.5211, "similarity": 0.4919,
"type": "fact", "source": "agent-zcode",
"content": "User prefers concise replies with tables and code examples",
"neighbors": []
}
]A different host used this memory and reported it back — uses +1, conf +0.1; and since the reporter agent-workbuddy ≠ source agent-claude, the first cross-host validation adds another +0.15:
uv run compound-memory feedback 20261007_86adf1 agent-workbuddy{
"id": "20261007_86adf1",
"confidence": 0.75,
"uses": 1,
"last_used": "2026-10-07",
"validated_by": ["agent-workbuddy"],
"evidence": {
"success_count": 1, "failure_count": 0, "contradiction_count": 0,
"last_verified": "2026-10-07",
"recent": [{ "date": "2026-10-07", "agent": "agent-workbuddy", "outcome": "success" }]
}
...
}Store health at a glance (fixed-bucket histograms, liveness, distillation yield):
uv run compound-memory stats{
"total": 2, "archived": 0, "active": 2,
"avg_confidence": 0.625,
"by_type": { "fact": 2 },
"by_ns": { "_shared": 2 },
"review_queue_entries": 0,
"uses_histogram": { "0": 1, "1-2": 1, "3-5": 0, "6-9": 0, "10+": 0 },
"confidence_histogram": { "<0.3": 0, "0.3-0.6": 1, "0.6-0.8": 1, "0.8-1.0": 0 },
"recent_feedback_7d": 1, "cross_validated": 0,
"distilled_total": 0, "distilled_recent_7d": 0
}Three things to notice:
New memories start at
confidence0.5 and move on evidence — feedback carries an outcome:successraises it,failurelowers it (floor 0.05),contradictionfreezes it into the review queue,obsoletearchives immediately.First validation from a different host earns an independent bonus (once per host per memory), with
validated_by/evidencetrails — confidence is evidence of correctness, not popularity.Hits embed one-hop neighbors automatically (empty here — no links yet;
memory_linkcreates bidirectional links that get recalled for free).
How compounding works
Interest source | Mechanism |
① Usage reinforcement |
|
② Link value |
|
③ Distillation |
|
④ Cross-agent validation | Feedback from an agent other than the source adds conf +0.15 |
Scoring (weights are the W_* constants in src/compound_memory/scoring.py): 0.70·similarity + 0.15·confidence + 0.10·recency(0.5+0.5·e^(−Δt/τ)) + 0.05·type weight. With the vector channel enabled, ranking switches to RRF fusion with an ε=0.04 prior tie-break (see the spec, "index as cache").
The 5 MCP tools
Tool | Purpose | Key points |
| Write a memory |
|
| Retrieve | Returns |
| Fetch by id | Always contains a |
| Link two memories | Bidirectional; both sides must be in the same ns; private-ns links require owner identity |
| Report "this memory was actually used" | Default |
Tool descriptions embed the protocol rules themselves, so agents keep the loop intact even without host-side rules injected. Full parameter reference: docs/agent-integration.md.
CLI
uv sync --extra dev # first clone: build .venv (later `uv run` reuses it)
uv run compound-memory init # initialize an empty store
uv run compound-memory write "Vercel Serverless has a 10s timeout" episode agent-workbuddy
uv run compound-memory search "Vercel timeout" # hits embed one-hop neighbors (limit 3, --no-neighbors to disable)
uv run compound-memory feedback <id> agent-claude
uv run compound-memory decay # run from cron
uv run compound-memory revive <id> # revive an archived memory
uv run compound-memory distill-plan # distillation candidates: merge_with (same-key strong) + possible_dup_of (BM25 weak) + promotion_candidate (high-activity episodes)
uv run compound-memory distill-apply "the merged insight" insight agent-workbuddy --sources <id1>,<id2> # atomic: product (links, origin=distillation) + source archival, one commit
uv run compound-memory stats # health: uses/confidence buckets + liveness + distillation yield
uv run compound-memory rebuild-index # rebuild the search cache anytime
uv run compound-memory review-queue # conflict queue (CLI-only entry)
uv run compound-memory git-log # audit trailMore operations: explain <id> (confidence composition + evidence detail for one memory), forget <id> --agent <id> (terminal removal, ADR-0009), review-resolve (adjudicate conflicts), extract <transcript|dir> (deterministic session-transcript mining).
Architecture
Agent (MCP client / CLI)
└─ memory_write | memory_search | memory_get | memory_link | memory_feedback
└─ MemoryStore (~/.agents/memory)
├─ namespaces/_shared/{episode,fact,insight,skill}/*.md shared area
├─ namespaces/agent-*/... private areas
├─ archive/... decayed archive (revivable)
├─ index/tokens.json rebuildable search cache
├─ review-queue.md fact/insight conflict queue
└─ .git/ auto-commit on every writeScheduled distillation prep (launchd / cron / systemd)
Per ADR 0001, the deterministic prep runs on a schedule while judgment (summarizing / merging) stays with the calling agent. Every day at 09:00 the candidate list lands in <root>/distill/last-plan.json. Pick one scheduler — launchd (macOS standard, catches up after sleep), systemd user timer (Persistent=true, same catch-up), or cron (most portable, no catch-up) — all three drive the same platform-neutral scripts/distill-prepare.sh. Ready-made templates with copy-paste instructions: scripts/com.compound-memory.distill-prepare.plist.tmpl (launchd), scripts/compound-memory-distill-prepare.{service,timer}.example (systemd), and the Chinese README for cron. The script runs set -eu: any failure exits non-zero (visible via launchctl list / systemctl --user list-timers / cron mail, log at distill/prepare.log). distill/ is a runtime artifact directory (auto-gitignored) — no commit noise; only distill-apply after agent judgment lands one atomic commit.
Documentation
docs/specs/0001-compound-memory-spec.md— design specdocs/agent-integration.md— per-host MCP configs + the unified usage protocol (Chinese)docs/adr/— architecture decision recordsCONTEXT.md— glossary (Chinese)skills/compound-memory/SKILL.md— usage rules for hosts (Chinese)
Development
uv run pytest tests/ -q # full suite (MCP tool boundary + distillation + lifecycle/index/CLI + input defense)
uv run mypy src/compound_memory/Test seams: the MCP tool boundary via in-process mcp.Client(server) (no subprocess) plus unit tests for core modules (scoring / index / store ops). CI runs tests, type checks, and a pure-wheel install smoke across Python 3.11/3.12/3.13.
Release
PyPI versions are immutable and the tag must match pyproject.toml's version (the release workflow verifies this and fails loudly). Releases go through GitHub Actions + PyPI Trusted Publisher (OIDC, no token): push a tag like v0.1.0 and release.yml builds and publishes automatically.
MCP Registry name: mcp-name: io.github.chinwe/compound-memory
Available Tools
5 toolsmemory_feedbackA
Report feedback on a memory with an outcome (closes the compounding loop — call after actually adopting a memory). outcome: 'success' (default, the memory worked), 'failure' (it misled you — confidence drops 0.2, floor 0.05), 'contradiction' (you dispute it — confidence frozen and a review entry is queued pending adjudication), 'obsolete' (it is superseded — archived immediately), 'unknown' (records the event only). Anything else is rejected. Success raises confidence (+0.1; extra +0.15 when a different agent validates for the first time). agent must be your own source agent id. Memories in a private 'agent-' namespace accept feedback only from the owner (agent = 'agent-' or ''). Archiving is reversed on feedback (except outcome=obsolete, which archives instead).
| Name | Required | Description | Default |
|---|---|---|---|
| agent | Yes | ||
| mem_id | Yes | ||
| outcome | No | success |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the entire behavioral burden and does so richly: exact confidence deltas (+0.1 success, +0.15 for first-time cross-agent validation, -0.2 failure with 0.05 floor), freezing and queued review for 'contradiction', immediate archiving for 'obsolete', rejection of any other value, ownership restriction for private 'agent-<name>' namespaces, and archiving being reversed on feedback. This is far beyond what any structured field provides.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose and the trigger before enumerating outcomes, and every clause carries information not present in the schema. The dense run of parentheticals makes it slightly heavy to scan, but there is little genuine filler to cut.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a three-parameter mutation tool with no annotations, no output schema, and 0% schema coverage, the description covers timing, per-outcome behavior, and authorization thoroughly. The one remaining gap is what the caller receives back (returned confidence, review-entry id), though the state changes themselves are disclosed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does: every meaningful value of 'outcome' is defined with its side effect and the default is noted, and 'agent' is constrained to "your own source agent id" with the private-namespace ownership rule spelled out. Only 'mem_id' is left implicit, which is reasonable for a self-evident identifier.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Report feedback on a memory with an outcome") and immediately differentiates itself from siblings like memory_write/memory_get by framing feedback as the step after a memory is adopted. The compounding-loop phrasing makes its role in the memory lifecycle unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"call after actually adopting a memory" gives a clear trigger condition, and the per-outcome effects tell the agent which outcome applies when. It stops short of explicitly comparing itself to the sibling tools or stating when not to call it, so it is clear context without full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_getA
Fetch a memory by id; one-hop link neighbors are included by default. reader: your own source agent id — required when the memory lives in a private 'agent-' namespace (readable only by its owner host). project: your workspace project slug — get-by-id itself is always readable regardless of project, but embedded neighbors are filtered to global ones plus your project's. After adopting it, call memory_feedback (agent = your source id).
| Name | Required | Description | Default |
|---|---|---|---|
| mem_id | Yes | ||
| reader | No | ||
| project | No | ||
| include_neighbors | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does substantial work: it discloses that neighbors are returned by default, that private-namespace reads require the owner's reader id, and that embedded neighbors are filtered to global plus the caller's project. It does not describe return shape, error modes, or whether missing ids fail loudly, leaving some behavioral gaps for a 4-parameter read tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core action and default behavior are front-loaded in the first clause, and the parameter notes follow in order. The longer em-dash sentences and the trailing memory_feedback instruction add real information but make the block denser than it needs to be.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description supplies the permission model, the default neighbor behavior, and the project-scoping rule — enough for an agent to call it correctly in the common case. It stops short of describing the response payload or failure behavior, which a 0%-coverage schema leaves entirely undocumented.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate and largely does: reader gets meaning plus a required-when condition, project gets meaning plus its filtering effect on neighbors, and include_neighbors is implied by 'included by default'. Only mem_id is left undefined, which is self-evident from the purpose sentence.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Fetch a memory by id', and immediately narrows scope by noting that one-hop link neighbors are included by default. It never explicitly contrasts with memory_search or memory_link, so an agent gets the gist but must infer the boundary between this and the search sibling on its own.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states a real conditional rule ('reader: required when the memory lives in a private agent-<name> namespace') and a follow-up workflow ('After adopting it, call memory_feedback'). What is missing is a negative case — there is no statement of when to prefer memory_search or memory_link instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_linkA
Create a bidirectional link between two memories (compounding source #2: association). Both memories must live in the same namespace; cross-namespace links are rejected. Memories in a private 'agent-' namespace accept links only from the owner (agent = your own source agent id).
| Name | Required | Description | Default |
|---|---|---|---|
| id_a | Yes | ||
| id_b | Yes | ||
| agent | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it does disclose the mutating effect (bidirectional link), the namespace constraint, and the ownership/permission rule. However, it is silent on idempotency, what happens if the link already exists, whether links can be removed, and failure behavior beyond rejection.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then constraints in descending order of importance. The '(compounding source #2: association)' aside is internal jargon that interrupts the flow slightly, but the paragraph is otherwise tight and waste-free.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an unannotated mutation tool with no output schema, the description covers the key operational constraints (namespace, ownership) an agent needs before calling. Missing details on return/error behavior are minor since no output schema exists to lean on.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It clarifies the agent parameter ('your own source agent id') and ties it to the ownership rule, but id_a and id_b get no explanation beyond their names, leaving two required params essentially undocumented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Create a bidirectional link between two memories'. The bidirectionality and the association framing are precise. It does not name a sibling such as memory_feedback to disambiguate, but the purpose is otherwise unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear conditions and exclusions: both memories must share a namespace, cross-namespace links are rejected, and private 'agent-<name>' namespaces accept links only from the owner. It stops short of comparing against siblings like memory_feedback or memory_write, so it is strong context without explicit alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_searchA
Search memories. Fuses lexical (BM25) and, when the vec extra + model are installed, vector (BGE) recall via RRF; otherwise falls back to lexical only. Each hit's 'similarity' field carries the normalized RRF fusion score (fused/rrf_max) in dual-channel mode, or the normalized BM25 score in lexical-only (linear) mode; per-hit rank components and the evidence summary come via explain=true. Confidence/recency/type act only as a small tie-break. Default scope is _shared PLUS your own private 'agent-' namespace (when your identity is known via attested process id or explicit reader) — private hits surface automatically, no extra query needed. Pass ns explicitly ('_shared' or 'agent-') to search a single namespace. project: your workspace project slug — results span global memories plus that project's; omit it and ONLY global (untagged) memories are returned (fail-closed: project memories never leak into general sessions). Each hit embeds up to 3 trimmed one-hop neighbors (active only, project-filtered) unless include_neighbors=False. explain: optional debugging carrier — pass true to attach a per-hit 'explain' object (lexical/vector rank, RRF score, prior term breakdown, retrieval channel) plus an 'evidence' summary line (outcome counts, last_verified, origin); omitted by default so the default hit shape stays unchanged. reader: your own source agent id — REQUIRED when ns is 'agent-' (private namespace, readable only by its owner host). Returns {'hits': [...]} sorted by score. Compounding rule: after actually adopting a hit, call memory_feedback (agent = your source id) — skipped feedbacks leave the store static.
| Name | Required | Description | Default |
|---|---|---|---|
| ns | No | ||
| query | Yes | ||
| top_k | No | ||
| reader | No | ||
| explain | No | ||
| project | No | ||
| include_neighbors | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden and does so richly: dual-channel vs lexical-only fallback, the meaning of the 'similarity' field per mode, that confidence/recency/type are only tie-breaks, default-scope semantics, the private-namespace reader requirement, the fail-closed project isolation, neighbor embedding behavior, and the returned shape {'hits': [...]} sorted by score.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with 'Search memories' before the mechanism detail, which is good, but the body is a dense run-on of em-dash clauses mixing score semantics, scope rules, and debugging options without paragraphing or bullets. Nearly every clause carries information, but the packing makes it hard to scan and pushes the actionable usage rules behind implementation trivia.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-parameter search tool with no output schema and no annotations, the description is nearly complete: it documents return shape, ordering, scope defaults, and the downstream memory_feedback step. Only top_k's role (result count) is left implicit, which is minor given its descriptive name and default.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate and largely does: ns, reader, explain, project, and include_neighbors all get concrete semantics beyond their terse titles, including a hard requirement on reader. The main gap is top_k, which is never mentioned in the description, though its name plus default is fairly self-evident.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Search memories') and immediately characterizes the mechanism (BM25 + optional BGE vector fused via RRF). It also differentiates from siblings by naming memory_feedback and the adoption/feedback loop, so an agent can place it in the memory toolset without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use guidance for every optional parameter: default scope is '_shared' plus private namespace, pass ns explicitly to narrow to one namespace, omit project to get only global memories (fail-closed), reader REQUIRED when ns is 'agent-<name>', explain for debugging, include_neighbors=False to suppress neighbors. When-not conditions (project omission consequences, reader requirement) are spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_writeA
Write a memory. type: episode|fact|insight|skill|decision; source: writing agent id; ns: '_shared' or 'agent-'. key: stable id for fact/insight/decision (enables conflict review). Write only stable facts (preferences, conventions, environment constraints, pitfalls), not session-temporary details; volatile status notes (in-progress work, remaining todos) either carry valid_until or stay out — a stale status memory is worse than none; prefer reusing an existing key over a new entry. valid_from/valid_until: optional ISO dates (YYYY-MM-DD) marking the fact's validity window — once valid_until has passed, the memory is excluded from search results but still readable via memory_get. project: optional lowercase-slug scope tag (lowercase alphanumeric segments joined by dashes) for workspace-specific memories — tagged memories are hidden from searches that do not pass the same project, while untagged (global) memories stay visible everywhere; declare your workspace project per the host usage rules and pass it on every write/search in that workspace. Returns the stored memory; conflict: true means a different version with the same key exists and a review entry was queued.
| Name | Required | Description | Default |
|---|---|---|---|
| ns | No | _shared | |
| key | No | ||
| type | Yes | ||
| links | No | ||
| source | Yes | ||
| content | Yes | ||
| project | No | ||
| valid_from | No | ||
| valid_until | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses conflict handling ("conflict: true means a different version with the same key exists and a review entry was queued"), the post-expiry lifecycle (excluded from search but readable via memory_get), and project-based visibility scoping. It stops short of stating auth requirements, rate limits, or upsert/idempotency semantics for repeat writes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then grouped by parameter and lifecycle concern; nearly every sentence adds usable information. It is dense and semicolon-packed, which costs a little readability, but there is no obvious filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter mutation tool with no annotations and no output schema, the description covers scope, conflict behavior, temporal validity, and project isolation, plus the return value. The remaining gaps are the undocumented links parameter and content semantics, which are minor but real.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate and largely does: it defines type values (episode|fact|insight|skill|decision), ns values ('_shared' or 'agent-<name>'), key semantics (stable id enabling conflict review), valid_from/valid_until format (ISO YYYY-MM-DD) and behavior, and project format (lowercase-slug) and visibility effect. Only links and content are left undocumented beyond their names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource ("Write a memory") and immediately enumerates the memory types, which tells an agent exactly what this tool stores. It does not name a sibling explicitly, so an agent must infer that this is the write path versus memory_search/memory_get, keeping it just below a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-and-when-not guidance: write stable facts (preferences, conventions, constraints, pitfalls) and not session-temporary details, with a stated rationale ("a stale status memory is worse than none"). It also directs the agent to reuse an existing key over creating a new entry and explains the valid_until escape hatch, leaving little to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
5 tool updates
v0.4.3- First observed
memory_feedback - First observed
memory_get - First observed
memory_link - First observed
memory_search - First observed
memory_write
TDQS
Scored across 5 tools
Each tool maps to a distinct action: search, get, write, link (association), and feedback (outcome reporting). The boundaries are clear with no functional overlap between any pair.
All five tools follow an identical memory_<verb> pattern (write, search, get, link, feedback). Fully predictable and consistent.
Five tools cover the core memory operations (persist, retrieve, associate, evaluate) without redundancy. The set is well-scoped for a memory substrate.
Write/search/get/link/feedback cover the lifecycle, and update is handled implicitly via same-key writes while obsolete-archiving substitutes for delete. Explicit update/edit or a bulk-list operation are minor gaps an agent can work around.
Maintenance
Related MCP Connectors
Persistent memory for AI agents. Search and store durable facts, preferences and decisions.
Persistent memory for AI agents. Semantic search, memory graph, W3C DID identity.
One memory, every AI. A shared, user-owned markdown memory your AI clients read and write over MCP.
Shared memory for coding agents. Stop re-explaining your codebase every session.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceLocal-first, file-based memory layer for AI agents — one shared Markdown vault across Claude, Codex, Gemini, Cursor and any MCP client. Provides read/write memory tools with an audit trail, per-agent trust levels, and Git sync; no cloud and no lock-in.2MIT
- AlicenseNot gradedqualityBmaintenanceLocal-first, multi-user shared memory for AI agents with semantic search, offline support, and team synchronization.MIT
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to capture, structure, remember, and retrieve source-backed memory as local Markdown files, with reviewable writes and no cloud dependency.199MIT
- FlicenseNot gradedqualityCmaintenanceProvides AI agents with persistent, local cross-session shared memory by combining vector semantic retrieval with knowledge graph relationships, and supports short/long-term memory management and local backups.-