Linksee Memory
linksee-memory
Local-first agent memory MCP. A cross-agent brain for Claude Code, Cursor, and ChatGPT Desktop โ with a token-saving file diff cache that nobody else does.
v0.2.0 makes the package English-first for global launch: the bundled auto-invocation skill is now bilingual (EN + JP), session-extractor patterns cover common English keywords (
let's go,pivot,doesn't work,same error again, etc.), and the install CLI shows test examples in both languages. No API changes. See CHANGELOG.
๐ Landing page: linksee-site.vercel.app (includes non-developer onboarding for Claude Desktop / Cursor / Claude Code)
๐ฌ See it in action
Without linksee-memory โ Monday morning, new Claude session:
You: We deployed last week but it crashed. How did we fix it?
Claude: I don't have access to previous sessions. Can you describe
what happened and walk me through the problem?
[30 minutes of log-spelunking and re-explanation]With linksee-memory โ Same question, different outcome:
You: We deployed last week but it crashed. How did we fix it?
Claude: Let me check my caveats...
๐ง [caveat] NextAuth sessions invalidate when JWT_SECRET
rotates โ redeploy all affected projects in parallel.
(from session 2026-04-13, importance: 0.9)
Is this the deploy you're asking about? We hit it when
we rotated secrets mid-flow.
You: Yes, exactly. Let's not repeat that.That single caveat memory is what separates "flat fact storage" from "the agent actually remembers the WHY". linksee-memory stores it across six explicit layers so retrieval stays explainable.
Related MCP server: auxly-memory-cli
๐ง The 6-layer structure
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ ๐ฏ goal โ what the user is working toward โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ ๐งญ context โ why this, why now โ constraints, people โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ ๐ emotion โ user tone signals (frustration, etc.) โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ ๐ implementation โ how it was done (+ what failed) โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ โ ๏ธ caveat โ "never do this again" ยท auto-protected โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ ๐ฑ learning โ patterns distilled from cold memories โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ
โผ
Ranked recall via relevance ร heat ร momentum ร importance
Returns match_reasons explaining each hitEvery memory is tagged with exactly one layer. caveat-layer entries are protected from auto-forgetting. Cold low-importance memories get compressed into learning entries via consolidate().
What it does
Most "agent memory" services (Mem0, Letta, Zep) save a flat list of facts. Then the agent looks at "edited file X 30 times" and has no idea why. linksee-memory keeps the WHY.
It is a Model Context Protocol (MCP) server that gives any AI agent four superpowers:
Mem0 / Letta / Zep | Claude Code auto-memory | linksee-memory | |
Cross-agent | โณ (cloud) | โ Claude only | โ single SQLite file |
6-layer WHY structure | โ flat | โ flat markdown | โ goal / context / emotion / impl / caveat / learning |
File diff cache | โ | โ | โ AST-aware, 50-99% token savings on re-reads |
Active forgetting | โณ | โ | โ Ebbinghaus curve, caveat layer protected |
Local-first / private | โ | โ | โ |
Three pillars
Token savings via
read_smartโ sha256 + AST/heading/indent chunking. Re-reads return only diffs. Measured 86% saved on a typical TS file edit, 99% saved on unchanged re-reads.Cross-agent portability โ single SQLite file at
~/.linksee-memory/memory.db. Same brain for Claude Code, Cursor, ChatGPT Desktop.WHY-first structured memory โ six explicit layers (
goal/context/emotion/implementation/caveat/learning). Solves "flat fact memory is useless without goals".
Install
npm install -g linksee-memory
linksee-memory-import --help # bundled importer for Claude Code session historyOr use npx ad hoc:
npx linksee-memory # starts the MCP server on stdioThe default database lives at ~/.linksee-memory/memory.db. Override with the LINKSEE_MEMORY_DIR environment variable.
Register with Claude Code
claude mcp add -s user linksee -- npx -y linksee-memoryRestart Claude Code. Tools appear as mcp__linksee__remember, mcp__linksee__recall, mcp__linksee__recall_file, mcp__linksee__read_smart, mcp__linksee__forget, mcp__linksee__consolidate.
Recommended: install the skill (auto-invocation)
Installing the MCP alone doesn't teach Claude Code when to call recall / remember. The bundled skill fixes that:
npx -y linksee-memory-install-skillThis copies a SKILL.md to ~/.claude/skills/linksee-memory/. Claude Code auto-discovers it and fires the skill on phrases like "ๅใซโฆ", "ใพใๅใใจใฉใผ", "่ฆใใฆใใใฆ", new task starts, file edits, and so on โ no need to say "use linksee-memory".
Flags: --dry-run, --force, --help.
Optional: auto-capture every session (Stop hook)
Add to ~/.claude/settings.json to record every Claude Code session to your local brain automatically:
{
"hooks": {
"Stop": [
{
"matcher": "",
"hooks": [
{ "type": "command", "command": "npx -y linksee-memory-sync" }
]
}
]
}
}Each turn end takes ~100 ms. Failures are silent (Claude Code never blocks). Logs at ~/.linksee-memory/hook.log.
Tools
Tool | Purpose |
| Store memory in 1 of 6 layers for an entity. Rejects pasted assistant output / CI logs unless |
| FTS5 + heat ร momentum ร importance composite ranking with |
| Complete edit history of a file across all sessions, with per-edit user-intent context. |
| v0.1.0 Atomic edit of an existing memory. Preserves |
| v0.1.0 List what the memory knows about โ cheapest "what do I know?" primitive. Filter by |
| Diff-only file read. Returns full content on first read, ~50 tokens on unchanged re-reads, only changed chunks on real edits. |
| Explicit delete OR auto-sweep based on |
| Sleep-mode compression: cluster cold low-importance memories โ protected learning-layer summary. Supports |
CLI utilities
Command | Purpose |
| MCP server (stdio) |
| Claude Code Stop-hook entry point |
| Batch-import Claude Code session JSONL history |
| Install the Claude Code Skill that teaches the agent when to call recall/remember/read_smart |
| v0.1.0 Summary of the local DB (entity count / layer breakdown / top entities / top edited files). Add |
The 6 memory layers
Each entity (person / company / project / file / concept) can have memories across six layers. The layer encodes meaning, not category:
{
"goal": { "primary": "...", "sub_tasks": [], "deadline": "..." },
"context": { "why_now": "...", "triggering_event": "...", "when": "..." },
"emotion": { "temperature": "hot|warm|cold", "user_tone": "..." },
"implementation": {
"success": [{ "what": "...", "evidence": "..." }],
"failure": [{ "what": "...", "why_failed": "..." }]
},
"caveat": [{ "rule": "...", "reason": "...", "from_incident": "..." }],
"learning":[{ "at": "...", "learned": "...", "prior_belief": "..." }]
}caveatmemories are auto-protected from forgetting (pain lessons, never lost).goalmemories bypass decay while the goal is active.
Architecture
A single SQLite file (better-sqlite3 + FTS5 trigram tokenizer for JP/EN) contains five layers:
Layer 1 โ
entities(facts: people / companies / projects / concepts / files)Layer 2 โ
edges(associations, graph adjacency)Layer 3 โ
memories(6-layer structured meanings per entity)Layer 4 โ
events(time-series log for heat / momentum computation)Layer 5 โ
file_snapshots+session_file_edits(diff cache + conversationโfile linkage)
The conversationโfile linkage is the key. Every file edit captured by the Stop hook is stored alongside the user message that drove the edit. So recall_file("server.ts") returns "this file was edited 30 times across 3 days, and here are the actual user instructions that motivated each change".
Why the design choices
Local-first โ your conversation history is private. Nothing leaves your machine.
Single file โ
memory.dbis one portable artifact. Backup = file copy.MCP stdio โ works with every agent that speaks MCP, no plugins per host.
Reuses proven schemas โ
heat_score/momentum_scoreported from a production sales-intelligence codebase. Rule-based, no LLM dependency in the hot path.
Roadmap
โ Core 6 MCP tools (
remember/recall/recall_file/forget/consolidate/read_smart)โ Stop-hook auto-capture for Claude Code
โ JP/EN trigram FTS5
๐ง
PreToolUsehook to auto-interceptRead(zero-config token savings)๐ง Cursor + ChatGPT Desktop adapters
๐ฎ Vector search via
sqlite-veconce an embedding backend is chosen (Ollama / API / etc.)๐ฎ Optional anonymized telemetry โ MCP-quality intelligence layer
Comparison with Claude Code auto-memory
Claude Code ships a built-in memory feature at ~/.claude/projects/<path>/memory/*.md โ flat markdown notes for user preferences. linksee-memory complements it:
auto-memory = your scrapbook of "remember I prefer X"
linksee-memory = structured cross-agent brain with file diff cache and per-edit WHY
Use both.
Telemetry (opt-in, off by default)
linksee-memory ships with opt-in anonymous telemetry that helps us understand which MCP servers and workflows actually work in the wild. Nothing is sent unless you explicitly enable it. No conversation content, no file content, no entity names, no project paths โ ever.
Enable
export LINKSEE_TELEMETRY=basic # opt in
export LINKSEE_TELEMETRY=off # opt out (or just unset the variable)Exactly what gets sent (Level 1 contract)
After each Claude Code session ends, the Stop hook sends one POST to https://kansei-link-mcp-production.up.railway.app/api/telemetry/linksee containing only these fields:
Field | Example | What it is |
|
| Random UUID generated locally on first opt-in. Stored at |
|
| Package version |
|
| How many turns the session had |
|
| How long the session lasted |
|
| Counts only |
|
| Names of MCP servers configured (from |
|
| Percent distribution of file extensions touched |
| counts | Tool usage counters |
What is NEVER sent:
โ Conversation messages (user or assistant)
โ File contents
โ Entity names, project names, file paths, URLs
โ Memory-layer text (goal / context / emotion / impl / caveat / learning)
โ Authentication tokens, API keys, secrets
โ Your IP address (only a one-way hash for abuse detection)
Why we ask
Aggregated MCP-usage data helps the KanseiLink project rank which agent integrations actually work for real developers. If you're happy to contribute, LINKSEE_TELEMETRY=basic takes 1 second to set and helps the entire MCP ecosystem improve.
The full payload schema and validation logic is open-source โ read src/lib/telemetry.ts if you want to verify exactly what leaves your machine.
Pricing
Free forever.
linksee-memory is local-first and runs entirely on your machine. There is no hosted component you need to pay for. The SQLite DB lives in your home directory; backup = file copy.
No account, no credit card, no API key. Just install and use.
Troubleshooting
Verify the skill was installed:
ls ~/.claude/skills/linksee-memory/SKILL.mdIf absent, run
npx -y linksee-memory-install-skill.Restart Claude Code. Skills are indexed on session start.
Check that the MCP is registered under the name
linksee(the skill expectsmcp__linksee__*tool names):claude mcp list | grep linkseeIf it's registered as something else, either re-register or edit
~/.claude/skills/linksee-memory/SKILL.mdto match.
Check the hook log:
cat ~/.linksee-memory/hook.logRun a manual test:
echo '{"session_id":"test","transcript_path":"/path/to/some.jsonl"}' | npx linksee-memory-syncMake sure the
Stophook in~/.claude/settings.jsonpoints tonpx -y linksee-memory-sync(not the old-import).
v0.0.6+ fixed the entity detection bug that collapsed all memories into the session's starting cwd. To re-index existing history with correct project attribution, run:
npx linksee-memory-import --allThe importer is idempotent (wipes existing session data before re-inserting). Typical runtime: a few minutes for hundreds of sessions. Expect a dramatic improvement in recall precision afterward.
Reduce max_tokens:
recall({ query: "...", max_tokens: 800 }) // default is 2000Or narrow with entity_name and layer:
recall({ query: "...", entity_name: "my-project", layer: "caveat" })rm -rf ~/.linksee-memory # nuke everything; next run creates a fresh DBOr delete individual memories via the forget tool with a specific memory_id.
Run consolidate โ it clusters old cold memories into compressed learning-layer summaries:
consolidate({ scope: "all", min_age_days: 7 })Caveat and active-goal layers are always preserved. Consider scheduling a weekly run via cron / Task Scheduler.
FAQ
Three axes:
Local-first: those tools require cloud accounts and send your data to their servers. linksee-memory runs entirely on your machine โ one SQLite file, no network calls by default.
WHY-layered: they store flat facts or knowledge-graph nodes. linksee-memory has 6 explicit layers (
goal/context/emotion/implementation/caveat/learning) so retrieval returns structured reasoning, not just data.File diff cache:
read_smarttool saves 86โ99% of tokens on file re-reads via AST-aware chunking. None of the memory services do this โ it's a feature usually shipped in IDEs.
Claude Code's auto-memory is Claude-only (doesn't help if you switch to Cursor or ChatGPT Desktop) and stores flat markdown with no structure. linksee-memory is the same local-first principle but:
Works across Claude Code, Cursor, ChatGPT Desktop (shared SQLite)
Structured 6-layer format makes recall explainable
Provides explicit forget/consolidate primitives rather than the agent guessing
Yes โ see tools/bench-read-smart.ts in the repo. The read_smart tool:
Hashes file content on first read, returns full content + chunk metadata (AST/heading/indent boundaries).
On re-read with unchanged mtime+sha256, returns
~50 tokensof "unchanged" confirmation instead of re-sending the file.On real edits, returns only the changed chunks as full content + unchanged chunks as metadata-only references.
For a typical TypeScript file edit in an agentic loop, this cuts round-trip token costs by ~86%. On pure re-reads (user navigating back to a previously-read file), savings exceed 99%.
The default is no sync โ the SQLite file lives at ~/.linksee-memory/memory.db and stays there. If you want multi-machine sync, put that directory under Syncthing / iCloud Drive / Dropbox / Google Drive โ it's a single file, so any file-sync tool works. (Avoid simultaneous edits from two machines while the MCP server is running on both; SQLite's WAL mode handles single-writer well but multi-writer conflicts can corrupt.)
Two mechanisms:
Ebbinghaus forgetting: cold low-importance memories decay naturally, eligible for auto-forget sweeps.
caveatlayer and memories withimportance โฅ 0.9are always protected.consolidate(): compresses clusters of cold low-importance memories by entity into a singlelearning-layer summary, then deletes the originals. Run vialinksee-memory-consolidateCLI (or schedule weekly).
In practice a solo developer hits ~100MB after 6 months of heavy use. A year-old DB I tested with 80K memories still recalls in <10ms.
Yes โ any MCP-compatible client works:
Claude Code:
claude mcp add -s user linksee -- npx -y linksee-memoryClaude Desktop: add to
claude_desktop_config.json(see onboarding on the LP)Cursor: add to MCP settings in Cursor
ChatGPT Desktop: same pattern once MCP support ships
Custom agent: the MCP stdio protocol is documented at modelcontextprotocol.io
By default: zero network calls, zero telemetry. There's an optional Level-1 telemetry mode you can enable that sends anonymized aggregate metrics (tool call counts, error rates, latency percentiles โ never memory content, never file paths, never queries). The exact payload schema is documented in the Telemetry section and you see every byte before opting in.
After install, in a new Claude session ask: "Can you remember that I prefer TypeScript over JavaScript?" Claude should confirm it called mcp__linksee__remember and stored this. Then in a different session ask: "What languages do I prefer?" It should recall via mcp__linksee__recall and return the preference with match_reasons showing why.
Support
Issues & bug reports: github.com/michielinksee/linksee-memory/issues
Feature requests: open an issue with the
enhancementlabelSecurity concerns: see SECURITY.md if present, or file a private advisory on GitHub
Company: Synapse Arrows PTE. LTD. (Singapore)
Changelog
v0.2.0 โ English-first launch readiness (2026-04-20)
Prepares the package for a broader (primarily English-speaking) audience on Reddit, Hacker News, and Anthropic Discord. No breaking API changes.
Bilingualized
SKILL.md(auto-invocation skill). The bundled skill thatlinksee-memory-install-skillcopies into~/.claude/skills/linksee-memory/SKILL.mdwas Japanese-first; it is now English-primary with Japanese trigger phrases preserved inline. English speakers now get the skill firing on natural English phrases ("how did we solve this before?", "same error again", "remember this") in addition to the existing JP triggers.Install-skill CLI output is bilingual: example test phrases shown after installation include both English and Japanese.
Session-extractor EN coverage (
linksee-memory-import): expanded regex patterns for decisions, failures, and caveats so English Claude Code session logs get auto-tagged correctly. Additions includelet's go,pivot,switch to,settled on,approved,doesn't work,stuck,same error again,hit an error,debug,broke,revert.Clearer caveat-forget error hint: the previous message said "lower importance below 0.9 first, then forget" which was misleading โ caveat-layer memories are permanently protected regardless of importance. The hint now correctly distinguishes layer-protection from pin-protection.
README rework for launch readiness: added a "See it in action" before/after scenario, ASCII 6-layer diagram, MCP Official Registry + Glama score badges, landing-page link, and an 8-item FAQ covering questions that surface during public launches.
Internal: SKILL.md now documents pairing with KanseiLink skill as an English workflow example.
No code changes to the MCP protocol surface; all existing MCP clients continue to work unchanged.
v0.1.1 โ Pin threshold tweak (2026-04-19)
Based on real-world feedback that importance=0.95 memories were not
being treated as pinned despite intent.
Pin threshold lowered from
>= 1.0to>= 0.9. Memories withimportance >= 0.9are now exempt from the auto-forget sweep and surfacepinned: trueinrecallandrememberresponses. This matches the natural mental model ("0.9 = high importance = should survive cleanup") without requiring exact1.0.All existing memories with
importance >= 0.9(including older ones set to0.9or0.95) become pinned automatically โ no migration needed.Updated tool descriptions and error messages to reflect the new threshold.
v0.1.0 โ Major UX update (2026-04-18)
Based on one week of dogfooding, here's what changed:
New tools
update_memoryโ atomic edit with preservedmemory_id. Solves the "forget+remember breaks session_file_edits links" bug.list_entitiesโ fast "what do I know about?" primitive for session init. Supportskind/min_memoriesfilters and returns layer breakdown.npx linksee-memory-statsโ local DB summary CLI.
recall enhancements
match_reasonsarray on each memory: e.g.["content_match_fts", "heat:hot", "pinned"].score_breakdownwith per-dimension scores (relevance / heat / momentum / importance).Pagination via
offset/has_more/stopped_by.limitparameter (hard cap, complementsmax_tokensbudget).bandfilter to request only hot/warm/cold/frozen memories.mark_accessed=falsefor preview queries that shouldn't bump heat.Layer aliases:
decisionsโlearning,warningsโcaveat,howโimplementation, etc.Fix: opportunistic refresh of stale entity momentum scores. Entities recalled >1 h after last remember() no longer return stale momentum.
remember enhancements
Quality check: rejects pasted assistant output / CI logs / stack traces unless
force=true.importance=1.0now implicitly pins the memory (survives auto-forget).Layer aliases accepted.
forget changes
Pinned memories (importance=1.0) now preserved alongside caveat-layer memories.
Clear error response when attempting to delete a protected or missing memory.
dry-run now includes
sample_ids_to_drop.
consolidate changes
dry_run: truepreview mode โ reports cluster count + candidates without writing.
Infra
Fixed fresh-DB migration bug (was querying
metatable before it existed).Bumped to Node 20+ for structured language feature usage.
All changes are backward compatible โ existing integrations continue to work. Server.ts version banner now reports v0.1.0.
Older versions
See GitHub Releases.
License
MIT โ Synapse Arrows PTE. LTD.
Available Tools
8 toolsconsolidateA
Sleep-mode compression. Clusters cold low-importance memories by (entity, layer), summarizes each cluster into a single protected learning-layer entry, deletes originals, and runs a forget-sweep. Run at session end or on demand. Set dry_run=true to preview without writing.
| Name | Required | Description | Default |
|---|---|---|---|
| scope | No | session | |
| min_age_days | No | Override the default 7-day minimum age for clustering (set to 0 to consolidate everything immediately, useful right after a bulk import). | |
| dry_run | No | Preview what would be compressed without modifying the DB. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully carries the burden of disclosing behavior. It explains that the tool is destructive ('deletes originals') and runs a 'forget-sweep'. It mentions that results are 'protected learning-layer entries' and that dry_run previews without writing. While some terms like 'forget-sweep' are not further explained, the overall destructive nature and side effects are transparent. Score 4 for solid disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loading the primary action and listing key steps. Every sentence adds value: first defines the operation, second gives usage guidance. No fluff or redundancy. This is a model of concise yet informative description.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given there is no output schema, the description should explain what happens and what the output looks like. It covers the process (clustering, summarizing, deleting, forget-sweep) and mentions dry_run. However, it omits details on the 'forget-sweep' and what 'protected learning-layer entry' means. For a tool with 3 parameters and no output schema, this is reasonably complete, but leaves minor gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 67% (2 of 3 params have descriptions). The tool description adds context around 'cold low-importance memories' which helps interpret the min_age_days parameter. It also reinforces dry_run usage. The scope parameter's enum values are not elaborated in the description, but the tool's domain implies session vs all. Overall, the description adds meaningful context beyond the schema, scoring a 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's action: 'Sleep-mode compression. Clusters cold low-importance memories by (entity, layer), summarizes each cluster into a single protected learning-layer entry, deletes originals, and runs a forget-sweep.' It provides a specific verb (compress/cluster/summarize/delete) and resource (memories). This clearly distinguishes it from sibling tools like 'forget' (individual deletion) or 'remember' (saving).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly recommends when to use: 'Run at session end or on demand.' It also advises on dry_run usage. However, it does not explicitly state when not to use it or provide alternatives, such as using 'forget' for single-item deletion. The context is clear but lacks exclusionary guidance, scoring a 4.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
forgetA
Explicitly delete a memory by id, OR run auto-forgetting across all memories based on forgettingRisk (importance + heat + age). Caveat-layer, goal-layer, and pinned (importance>=0.9) memories are always preserved. Prefer update_memory for corrections โ forget is destructive.
| Name | Required | Description | Default |
|---|---|---|---|
| memory_id | No | ||
| dry_run | No | Report what would be deleted without actually deleting. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses that the tool is destructive, specifies preserved memory types (caveat-layer, goal-layer, pinned), and mentions the dry_run capability via schema. However, the description does not detail the exact forgettingRisk formula or that deletion is irreversible, but the dry_run parameter is well-described in schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences long, with the main action stated first. It is concise and free of unnecessary detail, efficiently communicating key information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (two modes, preservation logic, destructive nature), the description covers essential aspects. It explains when to use each mode, what is preserved, and the destructive intent. No output schema is provided, but the description does not need to discuss return values as the tool's effect is primary.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds semantic context for memory_id (explicit deletion) but does not explicitly mention dry_run. However, it enriches the understanding by explaining the two operational modes and preservation rules, which go beyond the schema's parameter descriptions. Schema coverage is 50%, but the description compensates with broader context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states two distinct modes: explicit deletion by ID and auto-forgetting based on forgettingRisk. It also specifies preservation of caveat-layer, goal-layer, and pinned memories. This distinguishes it from siblings like update_memory and remember.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly advises to prefer update_memory for corrections, indicating when not to use forget and providing an alternative. Implicitly guides when to use: for destructive deletion or auto-forgetting.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_entitiesA
List the entities currently known to this memory store, sorted by recent activity. Use at the start of a new session ("what do I know about?") before issuing specific recall queries. Cheaper than recall for the "give me an overview" question.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | Filter by entity kind. | |
| min_memories | No | Only include entities with at least N memories. Default 1. | |
| limit | No | Max entities to return. Default 30. | |
| offset | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries full burden. It discloses sorting behavior and states it is cheaper than recall, but does not mention any specific permissions or side effects. For a list tool, this is adequate but not fully transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, no redundancy. First sentence states purpose, second gives usage context, third adds cost comparison. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema, but description explains return is list of entities sorted by recent activity. Mentions cost comparison and usage context. Could detail return format more but sufficient for low complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema descriptions cover 75% of parameters with defaults and enum details. Description adds no additional meaning beyond the schema, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states 'List the entities currently known to this memory store, sorted by recent activity' with specific verb and resource. It distinguishes from sibling tools like recall by noting it is for overview before specific queries.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly advises use at session start for 'what do I know about?' before specific recall queries. Compares cost to recall tool, providing clear context for when to use this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
read_smartA
Read a file with diff-only caching. Returns: (1) full content + chunk metadata on first read, (2) "unchanged" + cached chunk list (~50 tokens) if mtime matches, (3) "unchanged_content" if mtime changed but sha256 matches (touched but not modified), (4) changed chunks with content + unchanged chunks as metadata-only if the file was truly modified. Use INSTEAD of Read for files you have read before โ saves 50%+ tokens on re-reads.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Absolute file path | |
| force | No | If true, return full content regardless of cache state |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description fully discloses the caching behavior and all four possible return states based on file modification status, offering complete transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is moderately concise and well-structured with numbered return cases. While informative, it could be slightly more compact without losing clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite lacking an output schema, the description thoroughly explains all four possible return types and the effect of the 'force' parameter, making it complete for an agent to understand the tool's behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Input schema has 100% coverage with descriptions for both parameters. The tool description does not add additional semantic value beyond what the schema already provides, so a baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool reads a file with diff-only caching and lists four distinct return scenarios. It explicitly distinguishes itself from an alternative 'Read' tool by advising when to use it instead.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states 'Use INSTEAD of Read for files you have read before โ saves 50%+ tokens on re-reads', providing clear guidance on when to use this tool vs. an alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recallA
Retrieve memories relevant to the current context using full-text search (BM25) + entity-name match, re-ranked by a composite score (relevance ร heat ร momentum ร importance). Returns only what fits in the token budget, with match_reasons explaining WHY each memory was returned. Opportunistically refreshes stale momentum scores for entities in the result set. Supports pagination via offset/has_more. Layer aliases accepted. Use at the start of any task that might involve prior work.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | What you want to remember (free-text, entity name, or FTS5 MATCH expression) | |
| entity_name | No | Optional โ narrow to a specific entity | |
| layer | No | Optional layer filter. Accepts aliases (decisions/warnings/how/etc.) as well as canonical names. | |
| band | No | Optional โ only return memories whose heat_band matches. | |
| max_tokens | No | Approx token budget. Default 2000. Either max_tokens or limit stops iteration (whichever fires first). | |
| limit | No | Optional hard cap on number of memories. Stops at min(max_tokens-budget, limit). | |
| offset | No | Skip this many top results (pagination). Use has_more from prior response to decide next offset. | |
| mark_accessed | No | Set false for preview / listing queries that should not bump heat. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, but the description fully covers behavioral aspects: search method (BM25+entity), re-ranking composite score, token budget, match_reasons, pagination, and side effect of refreshing stale momentum scores. This is highly transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four concise sentences, each adding significant information. Starts with core action, then details, then usage guidance. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers retrieval mechanism, ranking, token budget, match_reasons, side effect, pagination, layer aliases, and usage context. Lacks explicit description of return format (fields of memory objects), but given no output schema, it mentions match_reasons, which is key. Still fairly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline 3. The description adds value by explaining the composite score, pagination using has_more, layer alias acceptance, and the token budget mechanism, which enrich understanding beyond individual parameter descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it retrieves memories using full-text search and entity-name match, with specific re-ranking. It distinguishes from siblings by focusing on relevance to current context and mentions 'Use at the start of any task that might involve prior work', providing clear purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use ('at the start of any task that might involve prior work'). However, it does not mention when not to use or provide explicit alternatives among siblings. Still, the context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
recall_fileA
Get the COMPLETE edit history of a file across all sessions, with per-edit user-intent context. Returns: total edit count, daily breakdown, list of distinct user intents that drove the edits, and the linked memories. Use this when you need to understand WHY a file was modified historically โ far more accurate than recall() for file-centric questions because it queries session_file_edits (every physical edit) instead of summary memories.
| Name | Required | Description | Default |
|---|---|---|---|
| path_substring | Yes | Substring to match against file_path (e.g. "search-services.ts" or full absolute path) | |
| max_intents | No | Max distinct user-intent snippets to return. Default 10. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided, so description carries full burden. It discloses what data is returned (edit count, daily breakdown, intents, linked memories) and that it queries session_file_edits. However, it does not address potential issues like multiple file matches or performance implications.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences plus return list: efficient, front-loaded with key information, no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 2 params, no output schema, and no annotations, the description explains return values (edit count, daily breakdown, intents, linked memories) but lacks details on format, error handling, or behavior for multiple matches.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with both parameters already described. The description does not add significant new meaning beyond what the schema provides, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states verb 'Get' and resource 'edit history of a file', emphasizes 'COMPLETE' and 'per-edit user-intent context', and distinguishes from sibling recall() by specifying it queries session_file_edits.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells when to use: 'when you need to understand WHY a file was modified historically' and contrasts with recall(), providing clear guidance on tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rememberA
Store a memory about an entity (person/company/project/concept/file) in one of 6 layers: goal (WHY), context (WHY-THIS-NOW), emotion (USER tone), implementation (HOW โ success/failure), caveat (PAIN lesson, never forgotten), learning (GROWTH log). Use this when you discover non-obvious goals, unexpected failures, user preferences, or decisions worth preserving. Pasted assistant output or CI logs are rejected (use force=true only if you are sure).
| Name | Required | Description | Default |
|---|---|---|---|
| entity_name | Yes | Name of the entity this memory is about | |
| entity_kind | Yes | ||
| entity_key | No | Optional canonical key (email, domain, file path) | |
| layer | Yes | One of: goal / context / emotion / implementation / caveat / learning. Common aliases (why, decisions, warnings, how, ...) are accepted. | |
| content | Yes | The memory content (plain text or JSON) | |
| importance | No | 0.0-1.0. Set to 0.9 or higher to "pin" a memory (protects from forgetting even outside caveat layer). | |
| force | No | Bypass the paste-back/CI-log quality check. Only set when you are sure the content is original user or agent thought. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Even without annotations, the description fully discloses behavioral traits: memory is stored in one of six layers, importance can pin a memory, and pasted output is rejected unless force=true. All behavioral aspects are transparent beyond what the schema provides.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured paragraph that front-loads the core purpose, then details layers and usage. Every sentence adds value without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (7 params, 4 required, 6 layers, no annotations, no output schema), the description covers purpose, usage, layer semantics, quality check, and force flag. It is complete enough for correct selection and invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 86% schema coverage, the description adds significant meaning: explains each layer with its purpose (WHY, WHY-THIS-NOW, USER tone, HOW, PAIN lesson, GROWTH log), clarifies the importance field for pinning, and explains the force field bypassing quality checks. This goes beyond the schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool stores a memory about an entity, specifies six layers (goal, context, emotion, implementation, caveat, learning), and distinguishes itself from sibling tools like recall, consolidate, forget, etc. The verb 'Store a memory' is specific and the resource (entity with layers) is well-defined.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use ('discover non-obvious goals, unexpected failures, user preferences, or decisions worth preserving') and when not to use ('pasted assistant output or CI logs are rejected'), with an alternative (force=true). Provides clear context for tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
update_memoryA
Atomically edit an existing memory in-place. Preferred over forget+remember because it preserves memory_id, which matters for session_file_edits links and referential integrity. Use to correct facts, update deadlines in goal entries, refine caveats, or re-score importance. Caveat-layer memories can be updated but cannot have their protected flag removed.
| Name | Required | Description | Default |
|---|---|---|---|
| memory_id | Yes | The memory.id to update | |
| content | No | New content (plain text or JSON). If omitted, content is kept. | |
| layer | No | Move to a different layer (aliases accepted). If omitted, layer is kept. | |
| importance | No | New importance 0-1. Set to 0.9 or higher to pin. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Despite no annotations, description discloses atomicity, preservation of memory_id, and referential integrity. It mentions caveat-layer constraints. Missing details on permissions or error behaviors, but adequate for a mutation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences. First sentence states core purpose and key benefit. Second sentence provides use cases and a constraint. No redundancy or unnecessary details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers purpose, usage, parameter details, and key behavioral traits (atomicity, referential integrity, caveat-layer limitation). Without output schema, return value is unmentioned, but overall completeness is high given the tool's complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, but description adds value beyond schema: 'aliases accepted' for layer and 'Set to 0.9 or higher to pin' for importance. This enhances understanding beyond parameter names and types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states it atomically edits an existing memory in-place, specifying what actions it performs (update content, layer, importance). It explicitly distinguishes itself from the forget+remember alternative by highlighting preservation of memory_id and referential integrity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit guidance on when to use: for correcting facts, updating deadlines, refining caveats, re-scoring importance. It also states a limitation for caveat-layer memories regarding protected flags, implying when not to fully update.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v0.2.0- First observed
consolidate - First observed
forget - First observed
list_entities - First observed
read_smart - First observed
recall - First observed
recall_file - First observed
remember - First observed
update_memory
TDQS
Scored across 8 tools
Each tool has a clear, distinct purpose: memory storage (remember), editing (update_memory), retrieval (recall, recall_file, list_entities), deletion/consolidation (forget, consolidate), and file caching (read_smart). There is no overlap that would cause misselection.
Most tools follow a verb or verb_noun pattern (e.g., consolidate, forget, list_entities, recall, recall_file, remember, update_memory). 'read_smart' breaks this pattern with a verb_adjective form. Overall, the naming is mostly consistent with minor deviations.
With 8 tools, the count is well within the ideal range for a memory management server. Each tool serves a necessary function without being redundant or excessive.
The tool set covers the core CRUD operations for memories, along with consolidation and file-specific retrieval. A direct 'get_memory_by_id' tool is missing, but recall can retrieve specific memories, so the gap is minor and workable.
Maintenance
Related MCP Connectors
shared AI-context layer for teams โ persistent memory your agents search and update over MCP
Token-efficient MCP memory for Markdown vaults. Tiered search, GraphRAG, AI memories.
One memory, every AI. A shared, user-owned markdown memory your AI clients read and write over MCP.
- memnodeOAuthdev.memnode
Persistent, inspectable memory for AI agents with lineage, correction, and a hosted MCP endpoint.
Related MCP Servers
- AlicenseAqualityBmaintenanceGives your AI persistent memory across conversations. Stores facts automatically, finds them by meaning using hybrid search with query expansion, and organizes everything into topics without manual tagging.181MIT
- AlicenseNot gradedqualityAmaintenanceLocal-first, file-based memory layer for AI agents โ one shared Markdown vault across Claude, Codex, Gemini, Cursor and any MCP client. Provides read/write memory tools with an audit trail, per-agent trust levels, and Git sync; no cloud and no lock-in.2MIT
- FlicenseNot gradedqualityCmaintenanceLocal-first cross-agent memory for AI coding agents. Persistent, shared memory over MCP โ what you tell one agent can be recalled by another โ with all data stored in a single local SQLite file, no cloud and no API keys.-
- AlicenseNot gradedqualityDmaintenanceLocal-first AI memory layer with hybrid retrieval and brain-inspired namespaces. Enables agents to save, search, and manage memories directly via MCP tools.0MIT