mneme
The mneme server provides vault-native memory capabilities for AI assistants, storing and retrieving knowledge as plain Markdown files using FTS5 full-text search and temporal reasoning — no LLM call required on the critical path.
Search the vault (
mneme_search): Free-text queries using FTS5 BM25 with optional filters by document type (session,topic,reference) and date range. Returns ranked hits with snippets, trust scores, confidence labels, and backend provenance.Recall documents (
mneme_recall): Retrieve vault documents bysession_idor date range, returning file paths, titles, modification times, and optionally the full Markdown body.Write or update content (
mneme_write): Atomically append or replace a Markdown section (H2 heading) in a vault file, with frontmatter support for new files and path containment enforcement for security.Summarize a topic (
mneme_summarize): Generate a grouped summary of documents related to a topic using FTS5 search, with optional date filtering and knowledge-graph augmentation.Build a timeline (
mneme_timeline): Retrieve chronologically ordered references for a subject, sorted by modification time, with bi-temporal filtering (valid_from,valid_to,as_of) support.Prime a new session (
mneme_prime): Preflight a context bundle combining recent sessions and topic-relevant matches within a configurable token budget, with deduplication and progressive format selection to maximize token efficiency.
Provides an optional knowledge graph enrichment backend using Neo4j for full-profile knowledge graph capabilities, enabling graph-based retrieval and temporal reasoning via Graphiti.
mneme Record
Vault-native memory for Claude Code. Markdown is ground truth.
Every session starts from zero. You re-explain the same architecture, the same constraints, the same decision you already settled yesterday, and the tools that promise to fix it mostly store your conversation history in opaque SQLite blobs and call an LLM every time you finish a session. The record of your own work ends up somewhere you cannot read, cannot grep, and cannot take with you.
mneme records what happened in each Claude Code session as plain markdown files in a directory you own (the vault) and indexes them with SQLite FTS5. The next session opens with a preflight block — today's headings, a git status summary, and the five most recently modified session documents — and the agent queries the index on demand through mneme_search.
Here is the file the Stop hook writes. The frontmatter, heading, and summary shape come from packages/mneme-cc-plugin/src/mneme_cc_plugin/hooks/stop.py and packages/mneme-core/src/mneme_core/distill/templates/summary-en.md; the contents below are illustrative.
---
id: session-2026-07-24
type: session
created: 2026-07-24T09:12:41.087213+00:00
schema_version: 1
---
# Sessions 2026-07-24
## 09:12 session a3f19c2e
transcript: `~/.claude/projects/mneme/a3f19c2e.jsonl`
**Session intent**: make the retrieval guard fail on the negative probe too
**Files touched**
- `benchmarks/retrieval/regression_guard.py`
- `benchmarks/retrieval/baseline.json`
**Tool activity** (34 events, 08:41:02 → 09:12:38)
- Edit: 11
- Bash: 9
- Read: 8
*Deterministic extractive summary (zero-LLM). Edit freely — this file is yours.*pipx install mneme-cc-plugin && mneme installThat installs the plugin and registers the lifecycle hooks; mneme doctor verifies the result, and profiles and per-client installs are under Three-Tier Install. Claude Code registers six hook events; Codex and Antigravity map four. Any other MCP client (Kimi, Qwen, Cline, Cursor) gets the ten MCP tools through the open adapter, with no lifecycle hooks and no automatic capture.
What it stores is a file you can open. When a session has actually changed something in the vault, the Stop hook appends a timestamped
## HH:MM sessionblock tovault/sessions/YYYY-MM-DD.mdwithtype: sessionfrontmatter, written atomically under a cross-process lock;mneme index rebuildreconstructs the FTS5 index over every markdown file in the vault, so the database is derived state you can delete.Closing a session costs nothing and takes 2 ms. "No LLM call on the critical path" is enforced in CI rather than promised in prose.
tools/spec_verify.pyparses all six lifecycle hook modules and fails the build on any import of seven network-capable roots (anthropic,openai,requests,httpx,urllib.request,urllib3,aiohttp), andpackages/mneme-cc-plugin/tests/integration/test_c3_no_network.pyre-checks the full transitive import closure of the three hot-path hooks at runtime, where a static scan cannot see. The Stop-hook proxy benchmark measures 2 ms at p95 over 100 sessions against a 1000 ms ceiling:benchmarks/latency/p95_guard.pyenforces that ceiling inside the same path-scoped benchmark workflow described below, andpackages/mneme-cc-plugin/tests/unit/test_stop_performance.pyre-checks it over 100 real Stop calls on every CI run, which carries no path filter.Search quality cannot silently degrade. A pull request that drops production-FTS5 nDCG@5 more than 0.02 below the locked baseline (0.8006, reported as 0.801), drops Recall@10 more than 0.05 below 1.00, or fails the out-of-vocabulary negative probe fails the build —
benchmarks/retrieval/regression_guard.py, run by.github/workflows/bench.ymlon every pull request and every push tomainthat touchespackages/mneme-core,packages/mneme-mcp,benchmarks/, or the workflow file itself.
Scope and limits
Those numbers come from the in-repo benchmark suite, seeded with MNEME_BENCH_SEED=42. Benchmark A uses a 500-document corpus. Benchmark E uses its default 300-document, 30-query fixture. Reproduce with make bench-all. Both figures above — the 0.801 and the 2 ms — carry the note that governs every figure in this README:
Note: All figures below are deterministic regression anchors computed on a seeded synthetic corpus; they are not real-world quality measurements (see ADR-012).
Retrieval claims follow the reachable path. The production mneme_search path is FTS5 BM25. The Python core contains an experimental feature-hashed lexical-vector backend and an RRF fusion protocol used by unit tests and synthetic benchmarks, but that backend is not wired into the MCP server or installer. Full-profile summarize and timeline can add gated local Graphiti and Neo4j fields. A true semantic embedding backend remains roadmap.
Privacy and network. Inline <private> tag redaction at staging write with SHA256 audit log. Zero outbound network calls except opted-in compression LLM and optional local Neo4j. Compression happens in the background, opt-in, with a cost cap.
Temporal reasoning. The deterministic claim lifecycle (valid-from/to, supersedes, as-of queries, contradiction detection, temporal blame provenance time-travel) is built in on every profile — pure SQLite, no extra dependency. Graphiti export and LLM claim extraction remain optional and never run on the Stop or critical path.
Context Continuity Engine (opt-in). Checkpoints are plain markdown in the vault, zero-LLM, default off.
Obsidian is fully optional. A vault is simply a plain directory of markdown files. mneme requires no specific editor, no external application, and no Obsidian installation. You can work with your vault using grep, git, VS Code, or any text editor. The term "vault" is borrowed convention for a self-contained markdown directory, not a dependency on any particular tool. Because the vault is plain markdown, a user who already uses Obsidian can point it at the same directory and get rendered notes, backlinks, and graph-view navigation over the wikilinks mneme writes. The two tools coexist cleanly: mneme stores all derived state (indexes, staging, audit logs) inside a .mneme directory that Obsidian ignores as a dot folder, and mneme's indexer excludes the .obsidian settings folder from indexing, so neither tool disturbs the other. Obsidian is a convenient viewer and navigator for vault content. It is not part of mneme's capture, indexing, or retrieval path, and it must not be treated as an installation prerequisite.
The full shipped / gated / roadmap ledger is in Implementation Status; the capabilities mneme does not ship at all are listed under What 2.0 Does Not Ship Yet.
Status: 4.1.0 public release. Package, plugin, runtime, citation, and documentation version sources are kept in lockstep by tools/version_bump.py (18 sources including this line, verified in CI), so no single declared version can drift. Upgrading from an earlier line: docs/UPGRADING.md.
Related MCP server: auxly-memory-cli
Tools
The MCP server registers ten tools. Every client that speaks MCP gets all ten; lifecycle hooks
and automatic capture are a separate layer that only Claude Code, Codex, and Antigravity provide.
The authoritative list is packages/mneme-mcp/src/tool_registry.ts.
Tool | What it does |
| FTS5 BM25 retrieval over the vault with Turkish casefold normalization, plus date, memory-type, and scope filters. Returns ranked hits and EvidenceCards carrying content hashes, trust, confidence, and the backend that actually ran. |
| Retrieves indexed documents by session identifier, date range, and scope. Returns paths, titles, modification times, memory types, and optionally the full markdown body. |
| Atomically appends or replaces a markdown section in a vault file. Enforces vault path containment and redacts private spans before storage. |
| Groups FTS5 matches for a topic by vault directory within optional date and scope filters. Graphiti enrichment appears only when that optional local graph integration is configured. |
| Returns scope-restricted references for a subject in chronological order. Graphiti facts and bi-temporal filtering appear only when the optional local graph integration is available. |
| Builds a token-budgeted preflight context bundle from recent sessions and topic-relevant matches. A caller session identifier enables per-session injection deduplication and full, keypoints, or reference formatting. |
| Queues a redacted memory-edit proposal for the policy drain. The server does not apply the edit directly; durable categories always require human approval. |
| Lists recent Context Continuity Engine checkpoints from the active scope, newest first. Missing checkpoint state returns an empty list. |
| Loads salience-ranked working-set items from a Context Continuity Engine checkpoint. Unknown and out-of-scope anchors return the same neutral not-found result. |
How mneme compares
Memory tools in the Claude Code and agent ecosystem make different trade-offs. The table below compares architectural capabilities across the dimensions mneme commits to, and it deliberately includes the rows where another tool leads. These cells describe design properties that are publicly verifiable from each tool's documentation. They are not a benchmarked ranking. For mneme's own reproducible numbers see Reproducible Numbers; for per-tool detail and an honest "where mneme is not the best fit" list see docs/COMPETITIVE.md.
Legend: ✓ built in · gated shipped, needs an opt-in dependency or flag · ~ partial · — not available · n/a the dimension does not apply.
Dimension | mneme | claude-mem | mem0 | Letta | Zep | Supermemory |
Plain-markdown store you can | ✓ | — | — | ~ | — | — |
Built-in | ✓ | — | — | — | — | — |
Deterministic Stop capture, no LLM call | ✓ | — | n/a | n/a | n/a | n/a |
Hybrid retrieval in the normal user path | ~ | ~ | ~ | ~ | ✓ | ✓ |
Temporal claim lifecycle (valid-from/to, supersedes, blame) | ✓ | — | ~ | ~ | ✓ | ~ |
Project and code graph (tree-sitter, PR-impact) | gated | ~ | — | — | — | — |
Adaptive token and context budget | ✓ | — | — | — | — | — |
Agent security: capability firewall, taint, approval gate | ✓ | — | — | — | — | — |
One-command lossless migration from claude-mem | ✓ | n/a | — | — | — | — |
Local-first, no cloud account required | ✓ | ✓ | ~ | ✓ | — | — |
Runs in Claude Code, Codex, Antigravity, any MCP client | ✓ | ~ | ~ | ~ | ~ | ~ |
License | Apache-2.0 | Apache-2.0 | Apache-2.0 | Apache-2.0 | cloud | open source |
Team memory with a web graph UI (mneme: self-hosted git sync + local console) | ✓ | — | ~ | — | ✓ | ✓ |
Agent autonomously rewrites its own memory (mneme: policy-graduated, rollback, audit chain) | ✓ | — | ~ | ✓ | — | — |
Auto-summarization at session end, on by default (mneme: deterministic zero-LLM) | ✓ | ✓ | — | — | ~ | ~ |
Localized observation-prompt presets (mneme: en + tr) | ~ | ✓ | — | — | — | — |
The 3.0 line closed the former gap rows on mneme's own terms. Team memory is self-hosted (any git remote, redaction-before-share, optional age end-to-end encryption) with a loopback-only web console rather than a vendor cloud. Autonomy is policy-graduated: the agent applies operator-allowed low-risk edit classes on its own, every change is journalled for one-command rollback and chained into a tamper-evident HMAC audit log, and durable categories always keep a human in the loop. The default-on session summary is deterministic and zero-LLM — no key, no cost, no latency — with LLM compression as the opt-in richer layer. Localized presets ship for English and Turkish today (claude-mem still leads on raw language count, hence the honest ~). Where a hosted product is genuinely the better fit, docs/COMPETITIVE.md says so.
Implementation Status
An honest, at-a-glance map of what is shipped today versus what is gated behind optional infrastructure or still on the roadmap. Shipped means present in the default install path and covered by CI. Gated means implemented but inactive until you provide the optional dependency or flag. Roadmap means designed (often with a seam or protocol already in place) but not yet packaged.
Capability | Status | Detail |
FTS5 BM25 retrieval ( | Shipped | default MCP search path |
RRF fusion protocol | Experimental | Python API plus synthetic benchmark harness; not wired into |
| Shipped | Python + TypeScript mirror; staging write |
Zero-LLM deterministic Stop capture | Shipped |
|
Adaptive context layer (shell compress, injection dedup, adaptive top-k) | Shipped |
|
Pattern + trajectory memory | Shipped | vault-markdown primitives |
Claude Code / Codex / Antigravity native plugins | Shipped (native) | Claude Code registers 6 hook events; Codex and Antigravity map 4; 2 skills + MCP |
Open MCP adapter (Kimi, Qwen, any MCP client) | Shipped (non-native) | MCP tools only, no auto-capture |
Background AI compression | Shipped (opt-in, default off) | monthly cost-cap ledger |
Feature-hashed lexical-vector retrieval | Experimental, disconnected | Implemented in Python core and benchmarks; no documented installer or MCP user path |
Temporal claim lifecycle + rule-based claim extraction + | Shipped | Graphiti export gated; LLM extraction optional, never on the Stop/critical path |
Project + code graph (mneme-graph) | Shipped (separate package) | tree-sitter Python/JavaScript/TypeScript extraction, community detection, PR-impact, entity canonicalization |
Code memory (mneme-code) | Shipped (separate package) | AGENTS.md procedural parsing, test-output to failure memory, fix-trajectory |
Domain modes | Shipped | vault-config user modes + CLI; clinical and security-review modes block external extraction and artifact upload; user config can never weaken a built-in privacy mode or disable redaction |
Agent security | Shipped | capability firewall, data-flow taint tracking, human-approval gate for durable edits, poisoned-vault benchmark |
Read-only console | Shipped | self-contained, offline, injection-safe HTML audit report |
Connectors (Obsidian local + GitHub injected-transport) | Shipped (opt-in, default off) | redaction-before-ingest; revoke by disabling |
KG temporal enrichment via live Neo4j/Graphiti writes (summarize/timeline) | Gated | full profile: Docker + Neo4j |
Packaged semantic embedding adapter | Roadmap | adapter protocol exists; no packaged semantic model or production MCP wiring |
Web-based knowledge-graph visual explorer | Roadmap | planned |
Multi-user team features (merge-conflict resolution, per-user ACL, dashboards) | Roadmap (Team) | read-only shared vaults via git remote work today |
Reproducible Numbers
These come from the in-repo benchmark suite, seeded with MNEME_BENCH_SEED=42. Benchmark A uses a 500-document corpus. Benchmark E uses its default 300-document, 30-query fixture. Reproduce with make bench-all.
Note: All figures below are deterministic regression anchors computed on a seeded synthetic corpus; they are not real-world quality measurements (see ADR-012).
Benchmark | Metric | Result |
A. Retrieval quality | nDCG@5, production FTS5 | 0.801 (Recall@10 1.00, MRR 0.734) |
B. Stop hook latency | p95 | 2 ms (constraint budget 1000 ms) |
B. Retrieve latency | p95 | 3 ms on indexed 500-doc corpus |
C. Shell output compression | reduction | 88 percent on redundant Bash logs |
C. Injection deduplication | skip rate | 95 percent in tight 20-turn sessions |
C. Compressed format | savings | keypoints 46 percent, ref 88 percent vs full |
D. Migration tool | assertions | 4 of 4 pass (migrated, idempotent, dedup, redaction) |
E. Head-to-head adapter | mneme leg | nDCG@5 0.831, MRR 0.772 on 300-doc fixture |
CI regression guards lock the path-scoped benchmark surface. Pull requests touching benchmarked code run the benchmark workflow. Any run that drops production FTS5 Benchmark A nDCG@5 by more than 0.02 or breaches the 1000 ms Stop hook p95 fails the build. The BoW RRF condition is reported only as a lexical-surrogate ablation.
Three-Tier Install
# Lite: FTS5 + Stop hook + privacy redaction + 10 MCP tools (Python + Node only)
pipx install mneme-cc-plugin
mneme install --profile=lite
# Standard: lite plus the standard optional dependency profile.
# The normal MCP search path remains FTS5. No --enable-dense installer flag ships.
mneme install --profile=standard
# Full: standard + gated Graphiti temporal knowledge graph enrichment (Docker + Neo4j)
mneme install --profile=fullUpgrade in place without losing data.
mneme upgrade --profile=standardVerify a healthy install.
mneme doctorUsing mneme with Codex
mneme is Claude-Code-native by origin. Because its retrieval core (mneme-core), its MCP server (mneme-mcp), and its vault contract are client-neutral, mneme also runs inside the OpenAI Codex CLI as an additive layer, with no loss of fidelity.
# Plugin: skills, MCP server, and lifecycle hooks together
codex plugin marketplace add OnourImpram/mneme
# Or wire just the MCP server into ~/.codex/config.toml
mneme install --client=codexCodex gets the same ten MCP tools, the same two skills, and the same vault. Four of mneme's six registered Claude Code hook events map to native Codex lifecycle events (SessionStart, PostToolUse, Stop, PreCompact), and SessionEnd folds into Stop. UserPromptSubmit has no Codex mapping. See docs/CODEX.md for the full coverage table and ADR-014 in docs/ARCHITECTURE.md for the multi-client design.
Using mneme with Antigravity
Antigravity (Google's agentic IDE) uses the Gemini-CLI extension model, and mneme ships a native extension for it.
mneme install --client=antigravityThis installs the mneme extension into ~/.gemini/extensions/, wiring the same ten MCP tools, the same two skills, a GEMINI.md rules file, and lifecycle hooks (SessionStart, PostToolUse, Stop, PreCompact) that map to the same mneme hook <event> core path Claude Code and Codex use. Because Antigravity exposes a Stop hook, session capture has full native parity.
Other MCP clients (open adapter)
Any MCP-capable client (Kimi, Qwen, Cline, Cursor, and others) can use mneme through the open adapter. This is the non-native tier: the ten MCP tools are available for the model to call, but there are no lifecycle hooks and no automatic capture.
mneme install --client=mcp --config <path-to-your-clients-mcp-config.json>mneme merges only its own server entry and leaves every other server in the config untouched. See docs/INTEGRATIONS.md for the client-tiering details and examples/ for a config snippet and a portable AGENTS.md template.
What 2.0 Ships
10 MCP tools:
mneme_search,mneme_recall,mneme_write,mneme_prime,mneme_summarize,mneme_timeline,mneme_propose,mneme_health,mneme_checkpoint_list,mneme_working_set_load. Default search is FTS5.mneme_healthreports the installation's own state — index schema and age, the locale profile the index carries, document count, staging depth — with a remedy attached to every warning. Full-profile summarize and timeline can add KG fields when the local graph is active.mneme_checkpoint_listandmneme_working_set_loadsupport the Context Continuity Engine (CCE): list available working-set checkpoints and load a checkpoint's salience-ranked items for JIT context re-injection.5 Claude Code hooks:
PostToolUse,SessionStart,Stop,PreCompact,SessionEnd.3 slash commands:
/mneme:prime,/mneme:recall,/mneme:migrate.2 skills:
mneme-prime,mneme-search.7-benchmark suite (
make bench-all): retrieval quality (A), Stop/retrieve latency (B), adaptive-context cost (C), claude-mem migration (D), head-to-head adapter (E), LongMemEval (F), CCE compaction-recall (G).One-command migration:
mneme-migrate migrate-from-claude-memwith tri-state archive flag, idempotent re-run, and a per-run hash-checked rollback manifest.Adaptive Context Layer:
distill.shell_compress,distill.injection_dedup,distill.adaptive_topk,distill.compressed_format, plusmneme auditfor token reports andmneme audit-logfor redaction audit entries.Pattern memory:
mneme patterns {store, search, list, show, delete}writing vault-markdown Signal/Action/Outcome documents.Trajectory recorder:
mneme trajectory {start, step, end, show, list}capturing per-session decision trails undervault/trajectories/.Background AI compression (opt-in, default off):
mneme compress {enable, disable, status, dry-run, run}with monthly cost cap ledger.
The 2.0 advanced line
Eight modules extend mneme's core for specialized workloads. All are gated or shipped as separate packages. All ship with redaction-before-store, provenance on every record, and confidence labels on every extracted claim. None runs on the Stop or critical path.
Project graph (mneme-graph): tree-sitter extraction for Python, JavaScript, and TypeScript; community detection; PR-impact analysis; entity canonicalization.
Code memory (mneme-code): AGENTS.md procedural parsing, test-output to failure memory, fix-trajectory capture.
Domain modes: clinical and security-review modes block external extraction and artifact upload at config layer. A user config can never weaken a built-in privacy mode or disable redaction.
Agent security: capability firewall, data-flow taint tracking, human-approval gate for durable edits, poisoned-vault benchmark.
Read-only console: self-contained, offline, injection-safe HTML audit report requiring no server.
Experimental feature-hashed lexical-vector retrieval: a Python-core backend and RRF seam used by tests and synthetic benchmarks. It is not connected to the installer or production MCP search path.
Temporal extraction + Graphiti export: rule-based claim extraction, valid-from/to lifecycle, supersedes links, and export to a local Graphiti instance. LLM extraction is optional and never on the critical path. Live Neo4j writes are gated on the full Docker + Neo4j profile.
Connectors (Obsidian local + GitHub injected-transport): opt-in, default off. Redaction runs before every ingest. Revoke by disabling in config.
What 2.0 Does Not Ship Yet
A credible "best in market" claim requires honest scope acknowledgment.
No packaged semantic embedding backend, and no installed dense-retrieval user path. The feature-hashed lexical-vector implementation remains an experimental Python API.
No dense or KG leg inside
mneme_search. MCP search is FTS5. KG enrichment is gated to summarize and timeline when local full-profile graph state is active.No cloud SaaS option. mneme is local-first by architectural conviction.
No web-based knowledge graph visual explorer. Planned.
No multi-user team features with merge-conflict resolution, per-user ACL, or team dashboards. Read-only shared vaults via git remote work today. Full team support is roadmap.
See docs/COMPETITIVE.md for the full landscape and which tools may suit those needs better.
Documentation
docs/ARCHITECTURE.md: design philosophy and the 16 Architecture Decision Records (ADR-001 through ADR-016, with ADR-006 superseded by ADR-015).docs/CONSTRAINTS.md: six sacred constraints and how to verify each.docs/VAULT.md: vault contract, frontmatter specification, atomic write pattern.docs/HOOKS.md: hook integration guide, timing budgets, fail-soft contract.docs/MCP.md: tool API reference with JSON schemas and example calls.docs/RELEASE.md: GitHub tag, release, and metadata checklist.docs/COOKBOOK.md: ten worked recipes with full Claude Code transcripts.docs/MIGRATION-FROM-CLAUDE-MEM.md: one-command migration with tri-state archive and hash-checked rollback walkthrough.docs/BENCHMARKS.md: methodology and the locked baseline numbers.docs/COMPETITIVE.md: living landscape document (monthly refresh).docs/PRIVACY.md: outbound network call audit and telemetry policy (zero by default).docs/GOVERNANCE.md: maintenance model, release authority, succession.
License
Apache License 2.0. See LICENSE and NOTICE. Releases up to and including the 2.x line were published under MIT and remain so.
Acknowledgments
Maintained by Onour Impram (@OnourImpram). The Adaptive Context Layer and the pattern and trajectory primitives draw conceptually from token-compression and agent-DB patterns proven in production internal tooling. The architecture is mneme-native, the lineage is operator experience.
Available Tools
6 toolsmneme_primeA
Preflight context bundle for a new session. Combines recent session-typed docs and topic-relevant matches inside a token budget. v1.0 uses the full injection format; Phase F.5 adds the keypoints/ref Adaptive Context Layer. Pass session_id (from CLAUDE_SESSION_ID) to activate per-session injection deduplication and progressive format selection (full → keypoints → ref as context fills).
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | No | Caller session identifier (e.g. CLAUDE_SESSION_ID). Enables per-session injection deduplication and progressive format selection (full to keypoints to ref). | |
| budget_tokens | No | ||
| topic_doc_count | No | ||
| task_description | Yes | ||
| recent_session_count | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description bears full responsibility. It discloses the injection format ('full'), progressive format selection, per-session deduplication, and token budgeting. However, it does not mention safety or side effects; as a read-only context fetch, this is acceptable but could be more explicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single paragraph of three sentences. It includes forward-looking version details ('v1.0 uses...Phase F.5 adds...') which may be unnecessary for the immediate AI agent use. The core purpose is front-loaded, but the version info adds clutter.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 5 parameters and no output schema, the description provides a reasonable overview but lacks details on return format, pagination, or error cases. It explains the adaptive layer and progressive format, but missing output schema leaves agents needing to infer behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 20% (session_id only). The description adds meaning for session_id (purpose, activation of features) but does not explain task_description, budget_tokens, recent_session_count, or topic_doc_count beyond the schema defaults. It hints at 'token budget' and 'topic-relevant matches', but lacks full compensation for the low coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Preflight context bundle for a new session.' It specifies what it combines (session-typed docs and topic-relevant matches) and mentions a token budget. This distinguishes it from sibling tools like mneme_recall (retrieval) and mneme_write (creation).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool (new session startup) and mentions passing session_id to activate deduplication and format selection. It does not explicitly state when not to use or compare to alternatives, but the context of 'preflight' and sibling names provide implied guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mneme_recallB
Retrieve documents by session_id or date range. Returns paths, titles, mtimes, and (when include_body) the markdown body of each match.
| Name | Required | Description | Default |
|---|---|---|---|
| top_n | No | ||
| date_to | No | ||
| date_from | No | ||
| session_id | No | ||
| include_body | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries full burden. It implies a read-only operation ('Retrieve'), but it does not explicitly state that the tool is non-destructive, idempotent, or requires certain permissions. The description is functional but lacks explicit behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences: first states the core functionality and filter options, second lists returned fields. It is front-loaded, concise, and avoids unnecessary details—every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has 5 parameters with no output schema and no annotations. The description covers the filtering logic and return fields but omits details like top_n default, whether date ranges are inclusive, how session_id and dates interact, and the exact structure of the returned data. This leaves ambiguity for the AI agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, so the description must compensate. It mentions 'by session_id or date range' and the return fields, but does not explain each parameter's purpose, constraints (e.g., top_n defaults to 10, max 50), or whether parameters are optional/required. The description adds some value but is insufficient for a 5-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool retrieves documents filtered by session_id or date range, and specifies the returned fields (paths, titles, mtimes, optional body). This distinguishes it from sibling tools like mneme_search (likely full-text) and mneme_write (obviously writing).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like mneme_search or mneme_timeline. It does not mention prerequisites, limitations, or cases where the tool should not be used. For example, it is unclear whether session_id and date range can be combined.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mneme_searchA
Retrieval over the vault. v1.0 ships FTS5 BM25 with Turkish casefold normalization. Optional date and frontmatter type filters. Returns ranked hits with snippets. The cards array (EvidenceCard) is the preferred output and carries content_hash, trust, and confidence_label. Each card carries a backend field identifying which retrieval leg produced it (fts5, dense, kg). The backends_used array lists every backend that returned at least one hit. hits is kept for backward compatibility.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Free-text query. | |
| top_k | No | ||
| filters | No | ||
| min_query_length | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full burden. It discloses the retrieval algorithm (FTS5 BM25, Turkish normalization), optional filters, and important output details (cards array with content_hash, trust, confidence, backend field, backends_used array). It also notes backward compatibility for 'hits', adding transparency without contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficiently structured: main purpose, algorithm, filters, output details. It is concise with no fluff, though it could be organized into bullet points for easier parsing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema, the description thoroughly explains the output structure (cards with fields, backends_used). It covers filters and backward compatibility. Missing details like snippet format or pagination, but top_k covers result count.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is low (25%), so description must compensate. It explains 'query' as free-text, 'top_k' default 10 max 50, and implies filters for date and type. However, it does not detail 'min_query_length' and misses that the schema already defines filter enum values. Adds some value but not enough for full clarity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it performs retrieval over the vault with FTS5 BM25 and Turkish normalization. It specifies optional filters and returns ranked hits with snippets. However, it does not explicitly differentiate from sibling tools like mneme_recall or mneme_timeline, relying on the name to imply search.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions optional date and frontmatter type filters but provides no guidance on when to use this tool versus its siblings (mneme_prime, mneme_recall, etc.). There is no explicit alternative naming or when-not-to-use advice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mneme_summarizeB
Topic summary grouped by directory. v1.0 uses FTS5 and can add Graphiti fields when full-profile KG state is active.
| Name | Required | Description | Default |
|---|---|---|---|
| top_k | No | ||
| topic | Yes | ||
| date_range | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must fully disclose behavioral traits. It mentions that v1.0 uses FTS5 and can add Graphiti fields under condition, but it does not explain what happens when KG state is inactive, whether the tool is read-only, any rate limits, or side effects. This leaves significant behavioral ambiguity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise at two sentences, with no wasted words. The first sentence states the core function, and the second adds technical context. However, the second sentence's jargon (FTS5, Graphiti) might be slightly opaque to an AI agent without further context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has 3 parameters, no output schema, and no annotations, the description is incomplete. It lacks parameter guidance, return value description, behavioral details, and error conditions. The conditional Graphiti mention is helpful but insufficient for full contextual completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, and the tool description does not add any meaning to the parameters (topic, date_range, top_k) beyond their schema definitions. For a tool with 3 parameters and no schema documentation, the description should compensate but fails to do so.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool produces a 'Topic summary grouped by directory,' which is a specific verb-object pair. The mention of FTS5 and Graphiti fields differentiates it from sibling tools like mneme_search (search) and mneme_recall (recall), making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for obtaining a structured summary with directory grouping and optional Graphiti fields, but it does not explicitly state when to use this tool over alternatives (e.g., mneme_search for raw results) or when not to use it. No exclusions or prerequisites are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mneme_timelineC
Temporal-ordered references for a subject. v1.0 returns FTS5 hits sorted by mtime ascending and can add bi-temporal Graphiti facts when full-profile KG state is active.
| Name | Required | Description | Default |
|---|---|---|---|
| as_of | No | ||
| top_k | No | ||
| subject | Yes | ||
| valid_to | No | ||
| valid_from | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Without annotations, the description must carry full burden. It discloses sorting order (mtime ascending) and a conditional side effect (adds bi-temporal facts). However, it omits key traits like default limit (top_k), pagination, or whether it only reads or can mutate data permanently. The version note adds transparency but is not comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very short (two sentences) with no redundancy. However, the first sentence is a noun phrase rather than a complete sentence, slightly harming readability. It earns points for efficiency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 5 parameters, no output schema, and no annotations, the description is insufficient. It fails to explain the purpose of date filters, the default limit, or the output format. The mention of FTS5 and Graphiti does not compensate for the missing parameter and return value documentation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%. The description does not explain any parameter beyond implying 'subject' is the subject. The date parameters (valid_from, valid_to, as_of) and top_k are entirely undocumented, leaving the agent without meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool provides 'temporal-ordered references for a subject,' which gives a general purpose. However, it lacks a clear verb (e.g., 'retrieve' or 'list') and relies on jargon like 'FTS5 hits' and 'Graphiti facts.' It vaguely distinguishes from siblings like 'mneme_search' but doesn't explicitly differentiate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions v1.0 behavior and a condition ('when full-profile KG state is active'), giving some context. But it does not specify when to use this tool versus alternatives like mneme_search or mneme_recall, nor does it provide any when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mneme_writeA
Atomically append or replace a markdown section in a vault file. Enforces assertWithinVault path containment. Optional frontmatter applies to newly created files only.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Path relative to vault root. | |
| content | Yes | Body content for the section. | |
| replace | No | ||
| section | Yes | H2 heading text without `## `. | |
| frontmatter | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses atomicity, path containment, and frontmatter applicability. However, it does not explain the behavior when replacing a section (e.g., whether it overwrites or merges), nor what happens if the file does not exist (implicitly creates? crashes?). Some key behaviors are omitted.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two efficient sentences front-load the purpose and key behaviors. Every phrase adds value: atomicity, operation type, path safety, frontmatter condition. No wasted words or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 5 parameters, 3 required, no output schema, and moderate complexity, the description is functional but incomplete. It fails to explain the return value (if any), error conditions, or the exact semantics of 'replace'. A more complete description would address these gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 60% coverage with descriptions for 3 of 5 parameters. The description adds context for 'path' (enforces assertWithinVault) and 'frontmatter' (applies only to new files). However, the 'replace' parameter lacks explanation of its effect, and 'content' and 'section' rely solely on schema definitions. Not significantly more value than schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool appends or replaces a markdown section in a vault file atomically, with precise scope and resource. The action and object are specific, and the name 'mneme_write' distinguishes it from sibling tools like mneme_recall (read) or mneme_search (search).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus siblings such as mneme_recall or mneme_search. There is no mention of prerequisites, alternatives, or scenarios where this tool should be avoided. The context signals show sibling tools exist, but the description offers no differentiation in usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v2.0.2- First observed
mneme_prime - First observed
mneme_recall - First observed
mneme_search - First observed
mneme_summarize - First observed
mneme_timeline - First observed
mneme_write
TDQS
Scored across 6 tools
Each tool serves a unique purpose: prime for session context, recall for retrieval by id/date, search for full-text search, summarize for topic summaries, timeline for temporal references, and write for file operations. No overlapping functionality.
All tools follow the consistent pattern 'mneme_<verb>', with verbs that clearly indicate the action (prime, recall, search, summarize, timeline, write). No mixing of conventions.
With 6 tools, the server is well-scoped for a knowledge vault system. Each tool addresses a core operation without redundancy or excess, fitting the typical 3-15 range perfectly.
The tool set covers essential operations: session priming, retrieval by criteria, full-text search, summarization, timeline queries, and writing. A minor gap might be the absence of a dedicated delete tool, but write's replace functionality and overall coverage make this a minor shortfall.
Maintenance
Related MCP Connectors
Persistent, governed institutional memory for Claude Code — specs, decisions, learnings.
- TaprootOAuthcom.taproothq
Persistent memory layer for AI tools. Save and recall notes across Claude and other MCP clients.
- mcpOAuthai.butlerbrain
Persistent memory for AI assistants. Save once; recall from Claude, ChatGPT, or any MCP client.
One memory, every AI. A shared, user-owned markdown memory your AI clients read and write over MCP.
Related MCP Servers
- AlicenseAqualityBmaintenanceLocal, searchable project memory for AI coding agents. Markdown source of truth, MCP interface, safe structured updates310Apache 2.0
- AlicenseNot gradedqualityAmaintenanceLocal-first, file-based memory layer for AI agents — one shared Markdown vault across Claude, Codex, Gemini, Cursor and any MCP client. Provides read/write memory tools with an audit trail, per-agent trust levels, and Git sync; no cloud and no lock-in.2MIT
- AlicenseNot gradedqualityBmaintenancePersonal, remotely hosted memory service for Claude, Codex, and other MCP clients that preserves research, project state, and decisions in an auditable revision store with a web UI.10 npm1MIT
- AlicenseAqualityAmaintenancePrivacy-first local memory vault every AI shares over MCP. Markdown + SQLite on your machine; Claude, ChatGPT, Cursor, and any MCP client read and write it live. No cloud, no account, no telemetry124MIT