mcp-light-memory
What is this?
MCP Light Memory is a lightweight, local-first, persistent memory system for coding agents and MCP clients (Warp, OpenCode, JetBrains AI Assistant / PyCharm, Claude Code, Cursor). It acts as a checkpoint + retrieval layer — it stores the minimum durable state needed to resume complex work across sessions, without keeping the full conversation in the model's context window.
When your agent starts a task, it calls context and gets back relevant past decisions, gotchas, constraints, and hypotheses — ranked, deduplicated, and trust-bounded. When it finishes, it checkpoints the working state. Next session, even after a restart, the memory is there.
Related MCP server: M3 Memory
Why use it?
Problem | How MCP Light Memory solves it |
Agents forget everything between sessions | Markdown files persist on disk; the agent retrieves them via BM25 + optional embeddings |
Full session history is too large for context | Only relevant memories are retrieved (token-budgeted, MMR-diversified) |
Cloud dependency / privacy concerns | 100% local, offline, zero network calls, no daemon |
Heavy setup / dependencies | Zero required runtime deps (pure Python 3.8+ stdlib); optional |
Prompt injection via stored memory | Every retrieved memory is explicitly |
Multi-project isolation | Router with registry allowlist, |
MCP protocol drift | Dual-era support: modern |
How it works (mechanisms)
Markdown is the source of truth. Every memory is a
.mdfile with YAML frontmatter (id,type,status,tags,sources,links,valid_from,valid_to,supersedes). Human-readable, diffable, durable.SQLite is a rebuildable cache. BM25/FTS5 index + optional embedding vectors + usage tracking. Delete it and everything rebuilds from Markdown.
Retrieval: pure-Python BM25 + optional dense embeddings → RRF fusion → MMR diversification → policy boosts (type/status/temporal) → token-budget cut. Adaptive mode: sparse first, dense only if weak.
Lifecycle:
remember→update→supersede(links both directions, never deletes history) →forget(archives, never deletes) →timeline(temporal view).search --at YYYY-MM-DDfor historical queries.Trust boundary: retrieved content is wrapped in
=== BEGIN/END INTERNAL_RAG MEMORY ===with aSECURITY NOTICEheader. Structured JSON/MCP carriestrust: untrusted+ optionalsecurity_flags: ["instruction_like_content"].Evidence freshness: each result includes
evidence_state(present/missing/unverifiable) for local path-like evidence — derived at retrieval time, never persisted.Multi-project router: one MCP stdio server in front of many projects via a JSON registry.
write:falseblocks mutating tools before spawning a child. Per-call subprocess isolation (no shared state).
Setup
Prerequisites
Python 3.8+ (uses
pylauncher,python, orpython3— the installer auto-detects the real interpreter and rejects the WindowsApps stub)Git (the target project must be a git repo)
Optional:
pip install sentence-transformers numpyfor better semantic retrieval
The current version is defined by the VERSION file — check it (or run mlm.py --version) instead of hard-coding an expected number.
Quick start
Clone this repo once, then install into any project:
# Windows (PowerShell)
git clone https://github.com/PeterPirog/mcp-light-memory.git ~/mcp-light-memory
python ~/mcp-light-memory/install.py . --client warp# Linux/macOS
git clone https://github.com/PeterPirog/mcp-light-memory.git ~/mcp-light-memory
python3 ~/mcp-light-memory/install.py . --client warpThe installer:
copies skill files + creates
INTERNAL_RAG/+AGENTS.mdruns
init+checkpoint+validate(soguardisOKimmediately)auto-registers the MCP server in the client config when it can do so safely (or reports
MANUAL_REQUIRED/ prints JetBrains instructions)writes the absolute path to the verified Python interpreter (survives Windows PATH issues)
python .agents\skills\internal-rag\mlm.py --version # reports the installed version
python .agents\skills\internal-rag\mlm.py status # expect: INTERNAL_RAG ready
python .agents\skills\internal-rag\mlm.py guard # expect: GUARD OKInstallation matrix
One installer, four clients, two config scopes. Full guide: docs/INSTALLATION.md.
Client | Project scope | Global scope |
Warp (config write automatic; project activation may require approval) |
|
|
OpenCode stable (V1) (automatic for safe JSON config writes) |
|
|
OpenCode 2 (V2, beta) (automatic for safe JSON config writes) |
|
|
JetBrains AI / PyCharm (manual in IDE UI) |
|
|
--globalchanges the scope of the CLIENT CONFIG (~/.warp/.mcp.jsonvs{repo}/.warp/.mcp.json,~/.config/opencode/opencode.jsonvs projectopencode.json). The server still points at the target project you installed into.Need one global MCP endpoint for many repositories? Use the multi-project router — docs/MCP-MULTI-PROJECT.md.
JetBrains/PyCharm is assisted, not fully automatic: the installer prepares the JSON + Working Directory; you add the server in Settings → Tools → AI Assistant → MCP and choose Server level = Project or Global.
Manual setup (no installer) per client: docs/INSTALLATION.md + client pages (Warp · OpenCode).
Zero-shot: copy-paste prompts for Warp and OpenCode
You can paste one of these directly into the client agent. Replace C:\Projects\App with the real target repository path.
Warp — install for one project:
Install and configure MCP Light Memory (mcp-light-memory) as an MCP server for project C:\Projects\App in Warp, using project scope. Use the repository https://github.com/PeterPirog/mcp-light-memory. If the tool is not cloned yet, clone it to a stable location outside the project; if it already exists, update it with git pull --ff-only. Apply the canonical installation contract from the repository and run install.py with TARGET_PROJECT=C:\Projects\App and --client warp without --global. Do not force-overwrite an existing configuration. After installation, verify from cwd=C:\Projects\App: mlm.py --version, mlm.py status, and mlm.py guard, and confirm that the Warp configuration contains mcp-light-memory and the C:\Projects\App path. Report success only after MCP REGISTRATION: REGISTERED and successful verification. If Warp requires an additional project activation/toggle/approval, state the exact client-side step and do not claim the server is active before it is completed.Warp — global client config for one project:
Install and configure MCP Light Memory (mcp-light-memory) in Warp globally for project C:\Projects\App. Use the repository https://github.com/PeterPirog/mcp-light-memory. If the tool is not cloned yet, clone it to a stable location outside the project; if it already exists, run git pull --ff-only. Apply the canonical installation contract and run install.py with TARGET_PROJECT=C:\Projects\App, --client warp, and --global. Remember: --global means the global Warp client configuration, while the server must still be bound to C:\Projects\App; do not use the multi-project router. After installation, verify from cwd=C:\Projects\App: mlm.py --version, mlm.py status, and mlm.py guard, and confirm that the global Warp configuration contains mcp-light-memory and the C:\Projects\App path. Report success only after MCP REGISTRATION: REGISTERED and successful verification.OpenCode — install for one project (stable/V1):
Install and configure MCP Light Memory (mcp-light-memory) as an MCP server for project C:\Projects\App in OpenCode. By "OpenCode" I mean stable/V1, so use --client opencode, not opencode2. Use the repository https://github.com/PeterPirog/mcp-light-memory. If the tool is not cloned yet, clone it to a stable location outside the project; if it already exists, run git pull --ff-only. Run install.py with TARGET_PROJECT=C:\Projects\App and --client opencode without --global. Do not force-overwrite an existing configuration. If the installer returns MCP REGISTRATION: MANUAL_REQUIRED (for example because opencode.jsonc exists), do not report success: safely edit the JSONC while preserving comments and unrelated settings if you have appropriate file-editing tools; otherwise report the exact manual action required. After real registration, verify from cwd=C:\Projects\App: mlm.py --version, mlm.py status, and mlm.py guard, and confirm that the OpenCode configuration contains mcp-light-memory and C:\Projects\App.OpenCode — global client config for one project (stable/V1):
Install and configure MCP Light Memory (mcp-light-memory) globally in OpenCode for project C:\Projects\App. By "OpenCode" I mean stable/V1, so use --client opencode. Use the repository https://github.com/PeterPirog/mcp-light-memory. If the tool is not cloned yet, clone it to a stable location outside the project; if it already exists, run git pull --ff-only. Run install.py with TARGET_PROJECT=C:\Projects\App, --client opencode, and --global. --global means the global OpenCode client configuration, while the server must still be bound only to C:\Projects\App; do not use the multi-project router. If the installer returns MCP REGISTRATION: MANUAL_REQUIRED, do not report success and follow the safe JSONC instructions. After real registration, verify from cwd=C:\Projects\App: mlm.py --version, mlm.py status, and mlm.py guard, and confirm that the global OpenCode configuration contains mcp-light-memory and the C:\Projects\App path.For OpenCode 2 / V2, use the same prompts but explicitly say OpenCode 2 / V2 and require --client opencode2. More variants: docs/ZERO-SHOT-SETUP-PROMPTS.md.
Configuration details
Warp
Warp reads MCP server configs from ~/.warp/.mcp.json (global, auto-spawns) or
{repo}/.warp/.mcp.json (project, requires a manual toggle per Warp docs).
Shape: mcpServers.<name> with command, args, working_directory (always set it — the memory store is resolved from it). See examples/warp.example.json and docs/WARP-SETUP.md.
OpenCode stable (V1)
OpenCode reads opencode.json/.jsonc in the project root, or
~/.config/opencode/opencode.json globally. V1 servers are flat under
mcp.<name> (no servers sub-key) with enabled: true and command as an
array — see examples/opencode-legacy.example.json and docs/OPENCODE.md.
OpenCode 2 (V2, beta)
Same config files, different shape: mcp.servers.<name>, command as an
array, and no enabled field (V2 disables via disabled: true) — see
examples/opencode-v2.example.jsonc and docs/OPENCODE.md.
JetBrains AI Assistant / PyCharm
PyCharm does NOT auto-read any MCP config file. The installer prints
ready-to-paste JSON + Working Directory; you add the server in
Settings → Tools → AI Assistant → MCP (STDIO) and choose Server level =
Project or Global. See examples/jetbrains.example.json.
Multi-project router
One MCP connection in front of many projects — registry allowlist, write:false hard boundary, per-call subprocess isolation.
Registry file (projects.json)
{
"projects": {
"backend": { "root": "/abs/path/backend", "write": true },
"shared-lib": { "root": "/abs/path/shared-lib", "write": false }
}
}Warp config for the router
{
"mcpServers": {
"mcp-light-memory-router": {
"command": "python3",
"args": ["/abs/path/mcp-light-memory/.agents/skills/internal-rag/irag_mcp_router.py", "--registry", "/abs/path/projects.json"],
"working_directory": "/abs/path/mcp-light-memory"
}
}
}See docs/MCP-MULTI-PROJECT.md for details.
Workflow
context --task "current task"
↓
recovery, if required (RECOVERY REQUIRED)
↓
checkpoint before first change
↓
implementation
↓
checkpoint after each milestone
↓
guard before finishingCore commands (CLI alias: mlm.py or legacy irag.py):
mlm.py context --task "..."
mlm.py checkpoint --reason "..."
mlm.py search --query "..." --limit 8
mlm.py remember --type decision --title "..." --body "..."
mlm.py show <ref>
mlm.py update <ref> --status superseded
mlm.py status
mlm.py guard
mlm.py validate
mlm.py doctorPath mapping (rebrand: internal-rag → MCP Light Memory)
New name | Legacy path (kept for compatibility) |
|
|
|
|
|
|
|
|
| — |
| — |
The on-disk folder INTERNAL_RAG/ and the skill directory .agents/skills/internal-rag/ are intentionally kept under their legacy names for zero-migration backward compatibility. See docs/MIGRATION-TO-MCP-LIGHT-MEMORY.md.
Durable memory (CRUD)
remember --type decision --title "..." --body "..." --tags "a,b" --evidence "src/x.py:42" --links "decisions/other.md"
show <path-or-id>
show <ref> --section Knowledge
update <ref> --add-tags "new" --append "New evidence: ..."
supersede <ref> --by <new> --reason "..."
forget <ref> # archives, does not delete
link --from <ref> --to <ref>
timeline --limit 20
status
historyTypes: decision, knowledge, constraint, gotcha, failure, hypothesis, session.
Task stack (interrupts)
mlm.py push --task "interrupted work" --reason "user-priority"
mlm.py tasks
mlm.py resume
mlm.py forget-task <id> # drop a specific task
mlm.py forget-task # clear the whole stackConfiguration (.irag.yml, optional)
retrieval:
limit: 10
mmr_lambda: 0.4
min_score: 0.3
embeddings: auto # auto | on | off
profile: english-fast # english-fast (default) | multilingual (PL/EN projects)
embeddings_model: null # explicit model overrides the profile
tokens:
context_budget: 5000
checkpoints:
auto_archive_sessions: true
max_task_stack: 24mlm.py config shows the effective configuration. mlm.py config --init writes a template.
Optional embeddings (better retrieval)
pip install -r requirements-optional.txtWhen the package is available and .irag.yml has embeddings: auto (default), retrieval uses embeddings with fallback to BM25. Override at runtime with --embeddings on|off|auto.
Two retrieval profiles (see docs/EMBEDDINGS.md):
english-fast(default,all-MiniLM-L6-v2)multilingual(intfloat/multilingual-e5-small) — for Polish-English projects
Offline / air-gapped
python pack.py --with-embeddings --profile english-fast
# -> internal-rag-offline-1.8.1.zip (name from pack.py; 1.8.1 = VERSION file)
# On the air-gapped machine:
unzip internal-rag-offline-*.zip -d internal-rag-offline
pip install --no-index --find-links wheels/ -r requirements-optional.txt
python install.py "/path/to/project" --client <warp|opencode|opencode2|jetbrains>See docs/OFFLINE.md for details.
Privacy & Git
The default install mode is local-only. The installer uses .git/info/exclude, not the project's .gitignore, so local memory and integration files are not accidentally committed.
Before publishing a project:
python .\privacy_check.py "D:\path\to\project"Expected: RESULT: PASS
Full removal from a project
python .\uninstall.py "D:\path\to\project"The uninstaller creates a backup outside the repository, then removes INTERNAL_RAG and its integrations. Use --keep-memory to preserve the memory data.
Documentation
Structure in a target project
project/
├── AGENTS.md
├── .irag.yml # optional config
├── INTERNAL_RAG/
│ ├── WORKING_STATE.md
│ ├── INDEX.md
│ ├── .checkpoint.json
│ ├── decisions/ knowledge/ gotchas/ failures/ hypotheses/ sessions/ archive/
│ └── exports/
├── .agents/skills/internal-rag/
│ ├── SKILL.md
│ ├── mlm.py # primary CLI (forwards to irag.py)
│ ├── irag.py # core (legacy alias, still the canonical module)
│ ├── irag_embeddings.py # optional plugin
│ └── irag_hooks.py # optional git hooks
└── .opencode/ # OpenCode integration (optional)Source of truth
current user instructions, 2. current code/tests/configuration, 3. specifications/ADRs, 4. verified memory, 5. session notes, 6. hypotheses.
Memory can be stale. Code takes precedence.
License
MIT.
Changelog
1.8.0 — JetBrains manual setup
--client jetbrainsno longer writes a fake config file (PyCharm ignores MCP config files). Prints ready-to-paste JSON + IDE menu instructions instead.--unregister --client jetbrainsprints a reminder to remove in the IDE UI.
1.7.2 — JetBrains cwd + client-specific messages
JetBrains: writes
working_directoryas a hint + printsWARNINGwith exact path to set inSettings → Tools → AI Assistant → MCP.Client-specific restart messages (Restart PyCharm / Restart Warp / Restart OpenCode).
Memory store: <path>printed in install output for immediate verification.
1.7.1 — Windows Python stub fix
detect_python()rejects the WindowsApps 0-byte stub; preferspy -0p; verifies each candidate with--version.Post-register verification: runs
--versionimmediately after writing the config and reportsPASS/FAIL.--unregisterdeletes empty config files + parent dirs (fixes dead.warp/.mcp.jsonskeleton →GUARD STALE).
1.7.0 — Rebrand to MCP Light Memory
Total rebrand from
internal-ragto MCP Light Memory (mcp-light-memory). New CLI aliasmlm(mlm.py). Logo/icon assets. Migration doc. GitHub rebrand checklist.Backward-compatible:
irag.py,INTERNAL_RAG/, old MCP server names preserved as deprecated aliases.18 rebrand consistency tests.
1.6.1 — Post-v1.6 hardening
Mutation/lifecycle benchmark (11 scenarios). Trust boundary (ADR-015):
trust: untrusted+security_flags. Evidence freshness (ADR-016):evidence_state. Scale benchmark (100/1k/10k). Router security regressions (+12 tests). Docs consistency test. 249 tests pass.
1.6.0 — Retrieval quality + MCP 2026-07-28
Memory-quality benchmark (37 cases). MCP
2026-07-28dual-era (server/discover,_meta,structuredContent,outputSchema). Registry strictwrite. Sources in chunk prefix. Adaptive retrieval. Link-aware context.consolidate --prepare. Router latency benchmark. ADR-010…016.
1.5.0 — Abstention gate + multi-project router
Relevance/abstention gate (
--meta). FTS5 candidate prefilter. Multi-project MCP router. MCP protocol hardening (pure stdout, SDK-verified). 168 tests.
1.4.0 — Chunking + dedup + temporal lifecycle
Section-aware chunking (schema v3). SimHash dedup. Multilingual PL/EN profile. Temporal lifecycle (
valid_from/valid_to/supersedes/--at).consolidate --dry-run.
1.3.0 — Persistent embedding cache
Chunk-level float32 BLOBs in SQLite. Multiple models coexist.
index --vacuum/--embed-missing.
1.0.2 — Token budget + privacy
Token budget enforcement. Stale memory detection. Duplicate detection. Privacy scan at write-time. Auto-checkpoint timer. Offline/air-gapped pack.
1.0.0 — Initial release
BM25 + MMR retrieval. Full memory CRUD. Task stack. MCP server (JSON-RPC stdio). Git hooks. Diagnostics. Export/import. Token budget.
Available Tools
8 toolscheckpointSave Task CheckpointA
Persist current operational task state so work can be resumed after interruption or context loss. Use at meaningful milestones and before ending a work session; use remember for reusable knowledge. This writes local checkpoint/task state, does not delete durable memories, and performs no network access.
| Name | Required | Description | Default |
|---|---|---|---|
| next | No | Recommended next action when work resumes. | |
| phase | No | Current task phase or milestone name. | |
| reason | Yes | Why this checkpoint is being created. | |
| blockers | No | Known blockers or unresolved issues. | |
| completed | No | Work completed since the previous checkpoint. | |
| in_progress | No | Work currently underway. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the safety profile (readOnly=false, destructive=false, openWorld=false), and the description adds the local-write scope, confirms no durable-memory deletion, and states no network access. It does not clarify whether checkpoints overwrite prior state or how the non-idempotent behavior manifests, so it is strong but not fully exhaustive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the core action, then usage and behavioral caveats. Every clause earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description adequately covers when to call it, what state it writes, and what it does not touch, which is sufficient for a checkpoint write. It could be slightly more complete on resume interaction with the sibling 'resume' tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all six parameters (reason, next, phase, blockers, completed, in_progress) are already documented in the schema. The description adds no parameter-level detail beyond what the structured fields provide, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Persist current operational task state') with the added purpose of resuming after interruption or context loss. It explicitly distinguishes itself from the sibling 'remember' by noting that tool is for reusable knowledge, so an agent can tell them apart without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use guidance ('at meaningful milestones and before ending a work session') and names the alternative ('use remember for reusable knowledge') with the condition that selects it. Nothing is left to inference about timing or routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contextLoad Task ContextA
Start or resume a task by retrieving relevant durable project memories and current working state. Use before project changes or after a context reset; use search for a focused lookup. This call may update task or usage state, is non-idempotent, does not delete durable memories, and performs no network access.
| Name | Required | Description | Default |
|---|---|---|---|
| task | Yes | Short description of the task being started or resumed; used to retrieve relevant project context. | |
| limit | No | Maximum number of relevant memories to include in the context packet. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=false, openWorldHint=false, and destructiveHint=false, so part of the description ('non-idempotent', 'no network access') restates structured data. It does add value beyond them by disclosing what state is touched ('may update task or usage state') and a non-obvious negative guarantee ('does not delete durable memories').
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no waste, with the core action front-loaded and usage/behavioral caveats packed into the second sentence. Every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter, non-destructive mutation with full annotation coverage and no output schema, the description covers trigger, alternative, state mutation, idempotency, and network behavior. It stops short of describing what the returned 'context packet' contains, a minor gap given no output schema is declared.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both 'task' and 'limit' are already documented in the schema. The description adds no syntax, format, or defaulting detail beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Start or resume a task') and resource ('durable project memories and current working state'), and explicitly routes the focused-lookup case to the sibling 'search'. An agent can distinguish it from search, resume, and checkpoint without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete triggers ('before project changes or after a context reset') and names the alternative tool and the condition that selects it ('use search for a focused lookup'). Nothing about when-to-use is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
guardCheck Checkpoint FreshnessARead-onlyIdempotent
Check whether project state changed since the last checkpoint. Use before finishing a task to detect uncheckpointed work, then checkpoint if stale. This operation is read-only and idempotent, does not modify memory or task state, and performs no network access.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| ok | Yes | True when project state is consistent with the last checkpoint. |
| fingerprint | No | Fingerprint representing the current project state used by the guard check. |
| changed_files | No | Project files detected as changed since the checkpoint baseline. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, and openWorldHint=false, and the description largely restates these facts ('read-only and idempotent', 'no network access'). It adds only marginal new context via 'does not modify memory or task state', which is implied by readOnlyHint. With annotations carrying the safety profile, this is adequate but adds little.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, with the core purpose and the usage trigger front-loaded before the safety caveat. The final sentence is largely redundant with the annotations, which slightly dilutes the otherwise tight structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only check with an output schema present, the description covers purpose, when to call it, what to do with the result, and the safety profile. Return-value semantics are left to the output schema, which is appropriate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline of 4 applies. No parameter-related claims are made or needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (check) and resource (whether project state changed since the last checkpoint), which is a distinct operation from the sibling 'checkpoint'. The name 'guard' is opaque on its own, but the description resolves it immediately.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the trigger ('Use before finishing a task to detect uncheckpointed work') and the conditional follow-up action ('then checkpoint if stale'), routing the agent to the sibling tool under a stated condition. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rememberStore Durable MemoryA
Store durable project knowledge that should survive across sessions. Use for stable decisions, constraints, gotchas, failures, hypotheses, or reusable knowledge; use checkpoint for temporary task progress. This writes local durable memory, does not delete existing memories, and performs no network access.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | Durable content to preserve: a decision, fact, constraint, gotcha, failure, hypothesis, or reusable session knowledge. | |
| tags | No | Optional comma-separated tags used to organize or retrieve the memory. | |
| type | Yes | Memory category describing the kind of durable knowledge being stored. | |
| scope | No | Optional scope describing where this memory applies. | |
| title | Yes | Short, specific title for the memory. | |
| status | No | Memory confidence state: active for established knowledge or tentative for information that still needs confirmation. | |
| evidence | No | Optional source or project-relative evidence reference supporting the memory. | |
| consequence | No | Optional impact or consequence of this memory for future work. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description's behavioral claims ('does not delete existing memories', 'performs no network access') largely restate destructiveHint=false and openWorldHint=false from the annotations, and it adds the useful locality framing of 'local durable memory'. It does not disclose what the annotations cannot: whether repeated calls with identical content create duplicates (idempotentHint=false implies they do) or what the write returns. With annotations already carrying the safety profile, this is adequate but adds limited new context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with purpose, then usage routing, then behavioral constraints. No filler or repetition; every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter write tool with full schema coverage, no output schema, and annotations present, the description covers purpose, usage routing, and side-effect profile. The only gap is duplicate/return behavior on a non-idempotent write, which is minor.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% with per-field descriptions and enum values, so the schema does the heavy lifting. The description's list of knowledge kinds loosely maps to the type enum but adds no format, syntax, or selection guidance beyond it. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource (store durable project knowledge that survives across sessions) with the durable/temporary distinction made explicit against the checkpoint sibling. An agent can tell immediately what this tool produces and how it differs from checkpoint without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States exactly when to use it (stable decisions, constraints, gotchas, failures, hypotheses, reusable knowledge) and names the alternative with its own condition (checkpoint for temporary task progress). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resumeResume Saved TaskA
Resume and remove the top saved task from the task stack. Use after tasks shows resumable work; use tasks when you only need to inspect the stack. This changes local task state, is non-idempotent, does not delete durable memories, and performs no network access.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false and idempotentHint=false, but the description adds genuinely new context: it removes an entry from a stack (pop semantics), leaves durable memories intact, and performs no network access. The non-destructive clarification is valuable since 'remove' could otherwise read as data loss, though it restates non-idempotency already in annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, both front-loaded: the operation first, then routing guidance and behavioral guarantees. No filler, and every clause carries information an agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-param mutation with no output schema, the description covers the operation, the routing decision, state side effects, and network behavior. The only gap is what 'resume' returns (the resumed task content), which an agent might reasonably want to know, but annotations and description together are otherwise sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; baseline 4 applies. No parameter-related claims are needed or missing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb pair and resource: 'Resume and remove the top saved task from the task stack.' This distinguishes it from sibling 'tasks' (inspect) and from 'checkpoint'/'remember' by naming the exact stack operation performed.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the precondition ('Use after tasks shows resumable work') and the alternative ('use tasks when you only need to inspect the stack'). Both when-to-use and when-not-to-use are covered with the sibling named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
searchSearch Project MemoryARead-only
Search durable project memories and return ranked results with confidence and abstention metadata. Use for focused fact retrieval without starting a task; use context when beginning or resuming work. It does not alter durable memory content or the task stack; repeated calls may update retrieval or usage metadata, so it is not marked idempotent. It performs no network access.
| Name | Required | Description | Default |
|---|---|---|---|
| at | No | Optional YYYY-MM-DD date used to filter memories by temporal validity. | |
| limit | No | Maximum number of ranked memories to return. | |
| query | Yes | Natural-language query describing the fact, decision, constraint, or prior work to retrieve. | |
| types | No | Optional memory types to include; omit to search all supported types. | |
| explain | No | Include per-channel retrieval scoring details for diagnostics. | |
| statuses | No | Optional memory statuses to include; omit to use the server default. |
Output Schema
| Name | Required | Description |
|---|---|---|
| reason | No | Human-readable explanation for the retrieval or abstention decision. |
| results | Yes | Ranked durable memory records returned by the search. |
| admitted | No | Number of candidate memories admitted to the result set. |
| rejected | No | Number of candidate memories rejected by retrieval admission rules. |
| abstained | Yes | True when retrieval intentionally returns no admitted result set. |
| confidence_kind | No | Whether retrieval confidence is heuristic or calibrated. |
| retrieval_confidence | No | Server confidence score for the retrieval result set. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/destructive/openWorld/idempotent hints, but the description goes further by explaining the reason for the non-idempotent hint (repeated calls may update retrieval or usage metadata) and clarifying that durable memory content and the task stack are not altered. The no-network-access statement also adds context beyond the openWorldHint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each load-bearing: purpose/output first, routing guidance second, behavioral caveats third. No repetition of schema content and no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description need not explain return values, and it still flags that results carry confidence and abstention metadata. Combined with full schema coverage and annotations, an agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all six parameters are already documented in the schema. The description adds no syntax, format, or interaction details for query, types, statuses, at, limit, or explain, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('search durable project memories') plus the output character ('ranked results with confidence and abstention metadata'). It explicitly contrasts itself with the sibling 'context' tool, so an agent can route between them without opening either schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when ('focused fact retrieval without starting a task') and a when-not with the named alternative ('use context when beginning or resuming work'). Nothing about tool selection is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
statusInspect Memory StatusARead-onlyIdempotent
Return memory, checkpoint, index, and recovery status for the current project. Use for diagnostics and health checks. This operation is read-only and idempotent, does not modify memory or task state, and performs no network access.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| memories | No | Total number of durable memories in the current project. |
| checkpoints | No | Total number of checkpoints in the current project. |
| index_status | No | Current retrieval index health or synchronization status. |
| last_checkpoint | No | Most recent checkpoint recorded for the current project. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint=false, and the description largely restates these (read-only, idempotent, no state modification, no network access). It adds no information beyond the structured hints, but it also contradicts nothing, so a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the core action and scope, followed by usage and safety context. No filler and nothing repeated needlessly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The operation is parameterless, an output schema exists to describe the return shape, and annotations cover the safety profile; nothing an agent needs in order to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate; per the rubric, a 0-parameter tool has a 4 baseline. The description correctly avoids inventing inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Return) and enumerates the exact resources covered (memory, checkpoint, index, recovery status for the current project), which is well beyond a restatement of the name "status". It stops short of explicitly distinguishing itself from the operational siblings like resume or checkpoint, so it lands just under a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use for diagnostics and health checks" gives a clear context for invoking it, implicitly steering the agent away from action-oriented siblings such as remember or resume. There are no explicit when-not conditions or named alternatives, which keeps it from a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
tasksList Pending TasksARead-onlyIdempotent
List the current task stack and resumable task state. Use before resume when you need to inspect pending work without changing the stack. This operation is read-only and idempotent, does not modify task state, and performs no network access.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| tasks | No | Saved or resumable tasks in the current task stack, in server-defined order. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the description's 'read-only and idempotent, does not modify task state' largely restates structured data rather than adding new behavior. 'Performs no network access' is a marginal addition but is already implied by openWorldHint=false.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the purpose in the first sentence and stays to three short sentences, though the safety restatement ('read-only and idempotent, does not modify task state') is somewhat redundant with annotations and could be trimmed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description need not explain return values, and it covers purpose, timing, and the safe-read nature adequately for a zero-parameter read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4; there is no parameter semantics to document, and the description correctly says nothing about inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('List the current task stack and resumable task state') that an agent can distinguish from sibling tools like resume or status without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear timing guidance ('Use before resume when you need to inspect pending work') and implicitly contrasts with resume by noting it leaves the stack unchanged, but does not explicitly name alternatives or when-not-to-use conditions beyond that.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v0.1.0- First observed
checkpoint - First observed
context - First observed
guard - First observed
remember - First observed
resume - First observed
search - First observed
status - First observed
tasks
TDQS
Scored across 8 tools
Descriptions carefully cross-reference each other (checkpoint vs remember, context vs search, resume vs tasks) so boundaries are mostly clear. Minor potential overlap between checkpoint/remember and guard/status, but the descriptions resolve these distinctions well.
All names are single lowercase words with no camelCase/snake_case mixing, giving a uniform visual style. However, it mixes verbs (checkpoint, resume, search, guard, remember) with nouns (context, status, tasks), so the convention is not fully predictable.
Eight tools is well-scoped for a lightweight memory server, with each tool mapping to a distinct operation (write memory, write checkpoint, search, inspect, resume, diagnose). No tool feels redundant or bolted on.
Covers the core lifecycle: durable memory (remember/search), task state (checkpoint/context/resume/tasks), and diagnostics (status/guard). The one gap is the absence of an update/delete/forget operation for durable memories, which descriptions explicitly note they never remove, limiting correction of stale knowledge.
Maintenance
Related MCP Connectors
Persistent memory for AI agents across Claude, ChatGPT and any MCP client.
Long-term memory for AI coding agents: durable project facts, recalled by every MCP client.
Persistent AI memory shared across Claude, ChatGPT, coding agents, and compatible MCP clients.
Shared memory for coding agents. Stop re-explaining your codebase every session.
Related MCP Servers
- AlicenseAqualityAmaintenanceMCP-native, local-first memory for coding agents that turns real sessions into reusable decisions, gotchas, and domain knowledge.174,130 PyPI6MIT
- AlicenseBqualityBmaintenanceLocal-first persistent memory layer for MCP agents with hybrid search, file ingestion, and GDPR compliance.201,429 PyPI26Apache 2.0
- AlicenseNot gradedqualityBmaintenanceProvides a persistent, local-first memory for coding agents over MCP, enabling automatic recall and recording of past work, failures, and decisions to reduce repetition and token usage.MIT
- AlicenseNot gradedqualityAmaintenanceProvides persistent memory for AI coding agents via MCP, enabling agents to store and semantically recall facts, events, and lessons across sessions, all running locally without cloud dependencies.Apache 2.0