agtmem-mcp
This server gives coding agents a local, file-based long-term memory over MCP, with search, read/write notes, bug tracking, candidate review, and codebase symbol lookup.
agtmem_search — hybrid keyword/substring search across notes, optionally filtered by type or scope, excluding superseded notes.
agtmem_read — fetch a full note by id, bounded by max_tokens.
agtmem_write — create or update notes (facts, decisions, projects, candidates, etc.), including replacing outdated notes via
supersedes.agtmem_append — safely append dated sections to existing notes.
agtmem_index — rebuild the SQLite search index from Markdown files; optionally scan a codebase for symbols.
agtmem_bug — record bugs with symptom, cause, and fix for future reuse.
agtmem_candidates — list unconfirmed candidate notes and promote them to accepted types.
agtmem_find — look up a symbol in the indexed codebase and get
file:linelocations.All data lives in plain Markdown files with SQLite as a disposable cache; no LLM calls, no network, zero runtime dependencies.
Provides the Hermes assistant with access to the agtmem memory tools, enabling it to search, record, and manage notes in a local Markdown/SQLite memory store.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@agtmem-mcpwhy is the index disposable?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
agtmem
Long-term memory for coding agents that you actually own.
Plain Markdown files on disk are the source of truth. A SQLite index is a disposable cache. Nothing here calls a model, and nothing here needs the network.
The problem
Most agent-memory products want to be the place your memory lives. That means your accumulated context — the decisions, the bug fixes, the reasons you chose one thing over another — ends up in someone else's database, behind someone else's pricing page, in someone else's format.
The failure mode is not a crash. It is that the free tier shrinks, or the pricing changes, or the company pivots, and now you are paying rent on your own notes. Or you switch tools and start over with an empty memory.
agtmem takes the opposite position: memory is a folder of text files.
Related MCP server: RepoRecall
The four rules
Everything else follows from these, and each one is enforced by a test.
The store is the contract; the index is a cache. Deleting the database must not be a loss.
rm ~/.agtmem/.index.sqliteand the next query rebuilds it from the Markdown. This is the acceptance test, and it is in the suite.The server never calls an LLM. Writing is file I/O. Retrieval is a SQL query. There is no per-recall token bill and no API key. Summarising is the agent's job, not the store's.
Zero runtime dependencies. Python standard library only, so this still runs on a bare interpreter in ten years. Asserted by the test suite, which parses every module and fails if any import is not in the standard library.
Every write goes through one code path.
index.index_note()refreshes a note and its supersession predecessor. Bypassing it is how a replaced note silently keeps showing up in search results — which is exactly the bug that motivated the rule.
Install
pip install -e .
agtmem initOr without installing:
python -m agtmem initRequires Python 3.10+ and a SQLite build with FTS5 (standard since 3.9;
agtmem doctor will tell you if yours lacks it).
Quick start
# Record a decision
agtmem add --title "Index is disposable" --type decision --scope myproject \
--body "Deleting the DB must not lose anything, so reindex is always safe."
# Record a bug in the three parts that make it reusable
agtmem bug --symptom "append hung for 15 seconds" \
--cause "the file lock was not re-entrant" \
--fix "per-process depth counter"
# Search it
agtmem search "why is the index disposable"
# Replace an outdated note (the old one is kept, marked superseded)
agtmem add --title "Index is disposable (revised)" \
--supersedes index-is-disposable --body "..."
# Map a codebase once, then look symbols up without reading files
agtmem anatomy . --scope myproject
agtmem find store_lock --signaturesHow it works
Three layers, each replaceable without touching the others.
~/.agtmem/
├── facts/ durable truths
├── decisions/ why something was done <- the most valuable category
├── bugs/ symptom / cause / fix
├── candidates/ unconfirmed, awaiting review
├── projects/ project overviews
├── anatomy/ codebase maps
├── sessions/ imported context-compaction summaries
├── log/ append-only journal
├── index.md hand-curated entry point (not a note)
└── .index.sqlite DELETE ME FREELYA note is a Markdown file with a deliberately restricted frontmatter block —
flat key: value lines only, no nesting, no block scalars, so it parses with
the standard library and stays comfortable to edit by hand:
---
id: index-is-disposable
title: Index is disposable
type: decision
scope: myproject
tags: design storage
status: active
origin: agent
captured: 2026-01-14
updated: 2026-01-14
---
Deleting the database must not lose anything.Retrieval is hybrid: SQLite FTS5 with unicode61 remove_diacritics=2 for
stemmed word matching, plus a trigram index for substrings that tokenisation
destroys (a file path, a hyphenated package name, an identifier). The two
result lists are fused with
Reciprocal Rank Fusion, so neither ranker needs score calibration against the
other.
Use it from an agent (MCP)
agtmem speaks MCP over stdio, so any MCP-capable agent can use it. Installing
the package creates agtmem-mcp, a no-argument executable:
pip install -e .
which agtmem-mcp # note the full path; most clients do not inherit PATHUse the full path to the executable in client configs. Many MCP clients spawn
servers with a minimal environment that does not include a virtualenv's bin or
Scripts directory, so a bare agtmem-mcp may resolve for you in a terminal and
still fail to start under the client.
AGTMEM_HOME is optional and defaults to ~/.agtmem. Set it only if you keep the
store somewhere else.
Most clients are satisfied with command alone. WorkBuddy is not — read the
next section before writing its config, or you will lose an evening to it.
WorkBuddy AI
Add the server to ~/.workbuddy-ai/mcp.json (note the filename — it is not
.mcp.json). On Windows the backslashes must be escaped:
{
"mcpServers": {
"agtmem": {
"command": "C:\\Users\\you\\envs\\default\\Scripts\\agtmem-mcp.exe",
"args": [],
"env": {
"AGTMEM_HOME": "C:\\Users\\you\\.agtmem",
"PYTHONIOENCODING": "utf-8"
},
"disabled": false
}
}
}
"args": []is load-bearing — omit the key and WorkBuddy exposes zero tools. The server still starts, still answerstools/listwith all eight tools in under a second, and still shows as enabled and trusted in the UI. It just never reaches the agent, and nothing in the logs says so.This is not hypothetical: it cost an evening here, and the same missing key had silently disabled an unrelated third-party server in the same config — which is what finally gave the pattern away.
The fix is free. WorkBuddy hashes
sha256(command + "|" + sorted(args).join(",") + "|" + sorted(env KEYS).join(",")), and a missingargsandargs: []serialise identically — so the hash is unchanged and the approval below survives. No second Trust click.
You must approve it once. WorkBuddy does not spawn a third-party MCP server just because it is in the config. Until you approve it, the connector shows as disabled with "This third-party MCP server requires your approval before connecting." Open the connector management page and click Trust on
agtmem. Restarting the app alone does not do it.
Two more consequences of that hash:
It covers the
envkey names, not their values. ChangingAGTMEM_HOME's value is free; adding or removing an env key means approving the server again.Every edit to
mcp.jsonneeds a restart. The file is read at startup only, so a correct fix applied to a running app changes nothing yet.
If the tools never appear
Four independent layers can each fail, and from the outside they all look the same: no tools, no error. Check them in this order.
Is it trusted?
~/.workbuddy-ai/mcp-approvals.jsonmust contain<configHash>::agtmem. Enabled and trusted are different things.Is the entry well-formed? Specifically, does it carry
"args": []?Is the server itself healthy? Rule this out first, by probing rather than by reading logs. WorkBuddy spawns stdio servers with only the variables in the server's
envblock — it does not inherit your shell — so reproduce exactly that: spawn the command with just those keys plusSystemRoot/windir/PATH, sendinitializeandtools/list, and count the tools. If that prints eight, the fault is in the client config, not here.Only then read the logs. They live in a dated directory, not
main.log:~/.workbuddy-ai/logs/<YYYY-MM-DD>/, whereskipping untrusted server "agtmem"means exactly what it says. Note that current builds log no connect line for stdio servers at all, so a missing line proves nothing either way.
The one reliable test is the tool index itself: have the agent look up
mcp__agtmem__agtmem_search. If that name does not resolve, the tools are not
loaded — whatever the UI and the logs suggest.
Hermes
Add a block under the top-level mcp_servers: key in ~/.hermes/config.yaml:
mcp_servers:
agtmem:
command: "/absolute/path/to/agtmem-mcp"
env:
AGTMEM_HOME: "~/.agtmem"
enabled: true
timeout: 120
connect_timeout: 60Then reload with /reload-mcp (or verify with hermes mcp test agtmem).
Claude Desktop, Cursor, and other JSON clients
{
"mcpServers": {
"agtmem": {
"command": "/absolute/path/to/agtmem-mcp"
}
}
}For clients that only accept command plus args, the module works too:
{
"mcpServers": {
"agtmem": {
"command": "python",
"args": ["-m", "agtmem", "mcp"]
}
}
}The eight tools
Tool | Purpose |
| search memory; returns snippets with ids |
| read one note, bounded by |
| create or update a note; |
| append to a note, concurrency-safe |
| rebuild the index; with |
| record symptom / cause / fix |
| list the candidate queue; |
| locate a symbol, get |
initialize returns an instructions field telling the agent to call
agtmem_search before starting work and to record conclusions with
agtmem_write. Clients that surface it will nudge the agent to use memory
without any prompt engineering on your side.
If you write your own MCP server
Two things that cost real debugging time here:
stdout carries JSON-RPC and nothing else. One stray
print()corrupts the stream and the client drops the connection. All human-readable output goes to stderr. The test suite asserts that every stdout line parses as JSON.Tool failures are results, not protocol errors. A missing note comes back as
isError: truewith the connection intact, not as a JSON-RPC error that kills the session.
Supersession instead of deletion
When knowledge changes, the old note is not deleted. It gets
status: superseded and a forward pointer, and the new note gets a backward
one:
agtmem add --title "New approach" --supersedes old-note --body "..."Default search skips superseded notes; --all includes them. You keep the
history of what changed and why, which is usually more valuable than the
current answer alone.
Raw sessions are not knowledge
type: session notes are excluded from search by default. Pass --sessions
(or sessions: true over MCP) to get them back.
A session is the transcript a note was distilled from — input, not knowledge.
It is also roughly ten times the size of a note (median 20 kB against 2 kB), and
BM25 rewards length: a transcript repeats every term the question uses, so it
outranks the note that actually answers it. The effect is not subtle. On a store
with 92 sessions among 275 notes, every one of the top ten results for
"How do I build the Android APK without gradlew?" was a raw session. With
sessions excluded, the first result is verbigem-android-build-invocation — the
note that answers the question.
The exclusion happens in SQL rather than after ranking. Filtering afterwards would let sessions consume the recall budget and hand back fewer results than asked for; the note has to surface even when a transcript would have outranked it.
Feeding it: distilling sessions
ingest-sessions imports raw session summaries. That is the input, not the
answer — a store holding 92 raw sessions is a store nobody reads. Turning them
into decisions/, facts/ and bugs/ is a two-stage job, and both stages
matter.
Stage 1 — distillation
Read a raw session and keep only what has lasting value: a decision, a fact about the system, a bug with symptom/cause/fix. Skip narrative, plans, and one-off exploration.
Measured density on a real corpus: 4.2 notes per session. Across 92 sessions that is ~390 notes — which is the trap. Distillation on its own swaps one problem ("too many raw sessions") for another ("too many notes").
Stage 2 — consolidation
Merge notes covering the same subject into one stronger note, and mark the
sources superseded. Nothing is lost: superseded files stay on disk and
search --all still finds them, they simply stop competing for rank with the
note that replaced them.
On the same corpus, 21 distilled notes collapsed to 12 — three Play Console notes into one, three Firebase identity notes into one, two admin-panel notes into one, and so on. A consolidation pass roughly halves the count without dropping a fact.
Cap the batch
Five to eight sessions per run. Not a performance limit — a review limit. 390 notes nobody reads is a worse deliverable than 21 notes somebody does.
Mark what has been processed
The store cannot do this for you, and the obvious place does not work: a custom
frontmatter key is silently dropped on the next save, because
Note.from_file() reads only the fixed FM_KEYS tuple and Note.render()
writes to_meta() back over it.
Use an append-only register note instead (log/session-distillation-audit).
Every run appends a section listing the session ids it consumed, and no entry
means not processed. It is greppable, human-readable, and survives a
serialization round-trip.
Where scope comes from
ingest-sessions derives a session's scope from the transcript's project
directory, not from the conversation: the directory is slugified, a leading
users/<name> or home/<name> is dropped, and the last two segments are kept.
A trailing session timestamp is stripped first — WorkBuddy names ad-hoc project
directories after the workspace plus the moment the session started, so
c-Users-milo-WorkBuddy AI-2026-09-04-11-43-59 yields workbuddy-ai and not
the clock reading 43-59.
Two consequences worth knowing:
A session run from a scratch workspace gets the scratch scope, not the project it was actually about. The directory cannot know what was discussed. If that matters, re-scope the note by hand —
scopeis a frontmatter field and nothing in the store depends on it for file layout (excepttype: project).Nothing validates a scope. A bad one is not an error, it is a new bucket, and
agtmem statswill list it next to your real projects as if it were one. Glance at that list occasionally.
Running it as a scheduled job
The server deliberately cannot do this — rule 2 above forbids LLM calls from
agtmem itself, and that is the right call. Distillation belongs to whatever
agent is already running, expressed as a scheduled prompt. The configuration
used here:
schedule | daily, 07:00 local |
cap | 8 sessions per run |
scope | one scope per run, in a fixed order; move to the next when the current one runs out |
consolidation | mandatory, in the same run |
register | append to |
The prompt hands the agent six steps: (1) read the register to learn what is
already done, (2) take the first scope in the list that still has unprocessed
sessions and pick at most 8 of them, (3) distill them into
decisions/facts/bugs, (4) consolidate duplicates and mark the sources
superseded, (5) append a register section, (6) run reindex and doctor.
One scope per run is deliberate. The consolidation step has to notice that two notes say the same thing, and that judgement is much easier inside a single project's vocabulary than across three of them. A run that drains a scope early is a short run, not a wasted one.
The list itself is just an ordered set of scope names, and the order should match whatever you care about. It only has to be explicit: "the largest remaining scope" sounds reasonable and drifts, because it depends on the agent noticing that the previous scope is empty.
Two things to get right if you copy it:
Pass every field when updating a note.
write_note(..., update=True)replacestype,scopeandtagsinstead of merging them, and the file moves to a different directory. An update that omits them silently resets the note's type and relocates it — which is how alog/note ends up infacts/.Cap the run. An uncapped job produces more notes than anyone will read, which is the exact failure this section exists to prevent.
Nothing in the store depends on the schedule. It is a convenience layer over
agtmem add and agtmem search: turn it off and the store keeps working.
Concurrency
Writes take an OS advisory lock (fcntl.flock / msvcrt.locking) and are
published atomically (mkstemp + os.replace).
The lock is re-entrant within a process, because
append_to()takes the lock and then callssave(), which takes it again.The lock is released by the kernel when a process dies, for any reason including
SIGKILL, so a crashed writer cannot wedge the store.The
.lockfile is never deleted and never written to. Deleting a file another process may be blocked on is a race that ends with two writers on two different inodes; writing to a byte range someone else has locked fails withEACCESon Windows.
The test suite runs twelve concurrent writers and asserts no update is lost.
Measuring whether it helps
A memory system you cannot measure is a liability, because it costs context on
every recall. agtmem eval reports R@5 and P@5 for both agtmem and a plain
grep over the same files, so the number means something:
agtmem eval --add "why is the index disposable => index-is-disposable"
agtmem eval --add-gap "a question no note answers"
agtmem evalA line beginning with ! records a known coverage gap: a question the store
cannot answer because no note was ever distilled for it. Those are counted
separately and excluded from R@5, because no amount of ranking work can return a
note that was never written — mixing them in would make a distillation gap look
like a retrieval failure and send you tuning the wrong component.
agtmem stats --usage reports how many tokens each retrieval actually injected.
Measured on a real corpus
Numbers below come from a 275-note store (183 distilled notes plus 92 raw session transcripts) with 20 scored cases. Ground truth was established by grepping the corpus for a distinctive phrase and confirming it occurs in exactly one active note — never by reading search results, which would make the eval score 1.0 by construction.
R@5 | P@5 | |
| 1.00 | 0.21 |
grep baseline | 0.05 | 0.01 |
P@5 looks low and is not a defect: most questions have exactly one right answer, so 0.2 is the best achievable score. It is reported anyway, because a metric that can only go up is not a metric.
Two rules that cost real debugging time, and are worth copying if you build your own set:
Never point ground truth at a raw session. It is input, not knowledge, and once sessions are excluded from search the case fails for a reason unrelated to ranking quality.
Never point it at a
supersedednote. Superseded notes are hidden by default, so the case is unanswerable by construction. Four of the first draft's targets had been superseded and had to be re-pointed at their successors.agtmem eval --addnow refuses both mistakes instead of writing the case.
Two layers, measured separately
An agent only ever reads distilled notes, so that is the layer the headline
number covers. But the same search can be pointed at the raw transcripts with
--sessions, and there the ranking used to fail badly: a session has a median
size of 21 160 B against 2 321 B for a note, so it contains every query term
simply by being a transcript of everything, and it won on length rather than on
relevance.
Counting term presence could not fix that, because both documents contain the
same terms. What fixed it was discounting coverage by length — see
COVERAGE_FREE_BYTES. Measured on the same 20 cases, with sessions left in the
pool:
before | after | |
correct note pushed out of the top 5 | 15/20 | 0/20 |
The discount has to be gentle at note scale, and that is the part worth knowing if you build something similar: it is one knob with two opposing requirements. Too strong, and it overrides better evidence — a 4.2 kB note matching four query terms must not outrank a 6.5 kB note matching five. Too weak, and a transcript starts winning again. The usable window measured out at 5 500–6 500 B, which is narrow. Sorting by term count first with length as a tie-break looks like the obvious fix and is wrong: it repairs the first case and re-breaks the second, because a transcript matches more distinct terms than a note does. | top result was a raw transcript | 17/20 | 0/20 |
What this eval does not measure
It now sits at R@5 = 1.00 on the knowledge layer, which means it has no headroom left and can no longer tell two good ranking strategies apart. The questions were written from each note's own vocabulary, so a query that paraphrases a note without sharing any of its words is still untested.
An earlier version of this table reported a gap between keyword queries (0.867) and natural-language ones (0.667). Re-measured on corrected ground truth, both phrasings score 1.00 — so that gap was at least partly an artifact of the old ground truth, which pointed at transcripts, and a transcript of everything is exactly the document a keyword query finds and a paraphrase misses.
The underlying limitation is unchanged: this is lexical retrieval, and lexical
retrieval is only as good as the vocabulary overlap between question and note.
Four alternatives were tried against the older set — proximity (NEAR) matching,
coverage as a score multiplier, title/tags weighted in BM25, and AND-first with
an OR fallback — and none beat coverage-first re-ranking. Paraphrase robustness
is what embeddings buy, and this project deliberately does not ship a model on
the hot path. If your queries are paraphrases rather than terms, you want a
vector index — and you should measure it, because the difference is not obvious
from the outside.
What it deliberately does not do
No embeddings, no vector search. FTS5 plus trigram covers thousands of notes without a model, a GPU, or a download.
No LLM calls. Summarising belongs to the agent — see Feeding it: distilling sessions for the workflow that does it, including the batch cap and the register note.
No proprietary database format. If SQLite disappeared tomorrow, the
.mdfiles are still yours.No cloud, no account, no sync. It is a folder. Sync it with whatever you already use.
No hooks, and no code that reads the store on your behalf. Deciding when to consult memory is the agent's job, not the store's — the store does not sit in the loop. Consumers that do sit in the loop live in
integrations/and are shipped separately from the core.
Integrations
Consumers of the store, kept outside the package: they may import agtmem,
nothing in agtmem/ may import them.
integrations/workbuddy/ — hits injected on every prompt
WorkBuddy spawns the CodeBuddy CLI, which has a hook system. This hook runs on
UserPromptSubmit, searches the store with the prompt's content words, and
appends the best hits to the context before the model sees the prompt — so a
question about why arrives with the relevant note already attached, instead of
depending on the agent deciding to go looking. It also emits a one-line reminder
on SessionStart.
It is the answer to the failure this project exists to fix: the store was being written to automatically and read never. What it is not is a relevance oracle — the README states plainly what the gate can and cannot do, with the measurements behind each decision.
Credits and prior art
This project borrowed a lot of ideas. The specific debts are listed in CREDITS.md — including which ideas came from OpenWolf, agentmemory, memanto/Moorcheh, Basic Memory, and the context-compaction behaviour of Claude-Code-style agents.
No code was copied; all of it is stdlib-only and written from scratch. Ideas and architecture are not copyrightable, which is precisely why the credits are specific rather than a vague "inspired by others".
How this was built
This project was designed and written in collaboration between a human (ihletru) and an AI coding agent. The division of labour was roughly:
The human set the direction, rejected the subscription-model alternatives, chose the design principles, and made the calls that mattered — including "delete the database must not be a loss" and "the server never calls an LLM".
The agent wrote essentially all of the code, the tests, and this documentation, and found and fixed the four concurrency and index bugs documented in docs/ARCHITECTURE.md.
It seems more useful to say that plainly than to pretend otherwise. Note also that copyright in AI-generated material is unsettled in several jurisdictions (US law, for instance, requires human authorship), which is one more reason the LICENSE names a human copyright holder.
License
MIT — do what you like with it.
Available Tools
8 toolsagtmem_appendA
Append a dated section to an existing note. Safe under concurrency.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| text | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries the full disclosure burden. It adds one useful behavioral guarantee, 'Safe under concurrency,' but does not explain mutation side effects, behavior when the note does not exist, or how dates are generated. The statement is not contradictory, but it is thin on operational detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one tight sentence with no filler; every part adds signal: operation, target, dated-section behavior, and concurrency safety. It is front-loaded and efficient, though short.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter append tool, the description gives the core operation and a relevant safety property, but is missing explicit parameter meaning, behavior for nonexistent notes, and return value expectations. There is enough to guess the intended use but not enough to cover all likely agent questions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description must therefore compensate for missing parameter documentation. The operation implies 'id' identifies the existing note and 'text' is the content of the appended section, but the description never explicitly explains either parameter or any formatting/constraint requirements. This is minimal compensation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Append'), a specific resource ('an existing note'), and a distinguishing behavior ('dated section'), which clearly separates it from siblings like agtmem_write, agtmem_read, or agtmem_search. Even without naming alternatives, the operation and target are unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for adding a dated section to an existing note rather than creating or replacing one, but it does not explicitly say when to prefer this over agtmem_write or mention any exclusions. An agent can infer the intended use case but is not given direct routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agtmem_bugA
Record a bug with the three parts that make it reusable: symptom, root cause, fix. Call this after solving anything non-obvious.
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | ||
| fix | Yes | what actually resolved it | |
| note | No | extra context, gotchas | |
| tags | No | ||
| cause | Yes | why it happened | |
| scope | No | global | |
| title | No | ||
| symptom | Yes | what was observed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It discloses the tool's purpose and the required structure (symptom, cause, fix), and implies a write/persist operation. However, it does not disclose whether the bug record is appended, overwritten, deduplicated, or how the 'id' and 'scope' fields affect behavior. For a write tool with no annotations, this is a moderate gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler. The core purpose and the trigger condition are front-loaded, and every word earns its place. The phrase 'the three parts that make it reusable' efficiently conveys the required structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a write tool with no annotations and no output schema, the description gives the essential trigger and required fields, but it lacks behavioral details like whether this is an insert or upsert, how 'scope' affects storage, and what happens on duplicate bugs. The sibling tools (search, read, append, write, index, candidates, find) suggest a memory system, but the description doesn't clarify how this bug record integrates with them. It's adequate but not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 50%, and the description names the three required parts (symptom, cause, fix) which maps to the schema's required fields. It adds the semantic framing that these three make the bug 'reusable', which is useful. However, it does not explain optional parameters like 'id', 'note', 'tags', 'scope', or 'title' beyond what the schema already provides. The description partially compensates for the coverage gap but not fully.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Record a bug') and resource ('the three parts that make it reusable: symptom, root cause, fix'), and it clearly distinguishes this from sibling tools by focusing on bug recording with a structured reusable format. It also names the trigger condition ('after solving anything non-obvious'), which helps an agent understand exactly what this tool is for.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit guidance on when to use the tool: 'Call this after solving anything non-obvious.' It does not explicitly name alternatives or exclusions, but the trigger condition is clear enough to route an agent to this tool versus the other agtmem_* siblings. A small deduction because it doesn't say when NOT to use it (e.g., for simple/obvious fixes).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agtmem_candidatesB
List unconfirmed notes (type='candidate') — inferences the agent made that a human has not validated. Pass promote= to accept one.
| Name | Required | Description | Default |
|---|---|---|---|
| to | No | fact | |
| promote | No | candidate id to accept |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must fully disclose behavior. It mentions listing and promoting, but does not clarify what happens when 'promote' is used (e.g., does it also return a list? does it mutate state?), nor does it explain the 'to' parameter that appears in the schema but is absent from the description. This incomplete disclosure is a significant gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no wasted words. The primary purpose is front-loaded, and the promote action is clearly appended. The structure is efficient and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no output schema, so the description must explain what the agent can expect from the call. It does not describe the return format, whether promote returns anything, or the behavior when 'to' is set. The unexplained 'to' parameter and lack of return value details make it incomplete for a tool with 2 parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 50% (only 'promote' has a description). The tool description adds meaning to 'promote' by explaining it accepts a candidate, but it completely ignores the 'to' parameter, leaving its purpose unexplained. The description does not compensate for the missing schema documentation of 'to'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'list' and the specific resource 'unconfirmed notes (type='candidate')', with a clear purpose: to show inferences a human hasn't validated. It distinguishes itself from sibling tools by focusing on the candidate type and the promote action, making it unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it (to list candidates) and mentions the promote action, but does not explicitly state when not to use it or name alternatives. It lacks exclusions or comparisons to sibling tools like search or read, leaving the agent to infer context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agtmem_findARead-only
Look up a symbol (function, class, method) in the codebase map and get file:line back. Cheaper and more precise than grepping or reading whole files.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| scope | No | ||
| symbol | Yes | ||
| max_tokens | No | ||
| signatures | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already signals this is a safe read operation. The description adds the behavioral trait that it is 'cheaper and more precise' than alternatives, which is useful context. However, it doesn't disclose what happens if the symbol is not found, whether it returns multiple matches, or how the limit parameter affects results. With annotations covering the safety profile, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, focused sentence that front-loads the core action and resource, then adds a comparative efficiency note. Every word earns its place, and it is appropriately sized for a simple lookup tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only lookup tool with a simple purpose, the description covers the main use case. However, with no output schema and 0% parameter coverage, the agent lacks details on return format, pagination, and the meaning of 'scope' and 'signatures'. The description is adequate for basic invocation but not fully complete for an agent to understand all behaviors and parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the five parameters. It explains the core 'symbol' parameter implicitly by describing the tool's purpose, but it doesn't clarify the meaning of 'limit', 'scope', 'max_tokens', or 'signatures'. The description adds some value by framing the tool as a symbol lookup, but it leaves the other parameters' semantics to be inferred from their names, which is a gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: looking up a symbol in the codebase map and returning file:line. It uses a specific verb ('Look up') and resource ('codebase map'), and the mention of being 'cheaper and more precise than grepping' distinguishes it from generic search approaches. It also differentiates from siblings like agtmem_search by focusing on exact symbol lookup rather than broader search.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool: when you need a symbol's location quickly and efficiently, as opposed to grepping or reading whole files. It doesn't explicitly name alternatives like agtmem_search or state when not to use it, but the efficiency comparison provides clear context. A more explicit mention of when to prefer agtmem_search over this tool would push it to 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agtmem_indexA
Rebuild the search index from the Markdown files. Always safe: the index is a cache and deleting it loses nothing. If path is given, also rescan that codebase into the symbol index and refresh its code map note.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | codebase root to scan | |
| scope | No | scope name for the scan |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It honestly discloses that the index is a cache, that deleting it loses nothing, and that path triggers additional rescan/refresh actions. This directly addresses potential destructive concerns without relying on annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with zero filler. The main action and safety guarantee are front-loaded, followed by the conditional behavior. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and no annotations, the description explains core behavior but omits what the tool returns and how `scope` affects the operation. It also lacks any guidance on when to use this over siblings. Adequate but has clear gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3. The description adds meaning beyond the schema by specifying that `path` triggers a codebase rescan and code map refresh. It does not elaborate on `scope`, but the added context for `path` elevates the score.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the operation with specific verbs and resources: 'Rebuild the search index from the Markdown files' and 'rescan that codebase into the symbol index and refresh its code map note.' This distinguishes it from sibling tools like agtmem_search (search) and agtmem_read/write (content access).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives context about safety ('Always safe') and conditional behavior ('If path is given'), but it does not explicitly say when to use this tool versus alternatives like agtmem_search or agtmem_find. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agtmem_readARead-only
Read one note in full by id. Bounded by max_tokens.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| max_tokens | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The readOnlyHint annotation already establishes the non-destructive nature, so the description's additional contribution is limited to the max_tokens bound. It discloses that output is limited, which is useful, but it does not explain truncation behavior beyond that or what happens when the note exceeds the bound. The 'in full' wording slightly muddies the disclosed truncation behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler. The core action is front-loaded, and the max_tokens caveat follows immediately. Every phrase earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-by-id tool with two parameters and a readOnlyHint, the description covers the essential invocation details. Minor gaps remain: there is no output schema, so the response format is not described, and behavior for missing ids or tokens is not mentioned. These are small omissions for a low-complexity tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description carries the burden of explaining parameters. It does name both parameters' roles: 'by id' clarifies the id parameter, and 'Bounded by max_tokens' gives a functional meaning to max_tokens. However, it does not explain what an id is, where it comes from, or how max_tokens interacts with note length beyond being a bound.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Read'), a precise resource ('one note... by id'), and a key constraint ('Bounded by max_tokens'). It is clear enough to distinguish from sibling retrieval tools like agtmem_search or agtmem_find, though it does not explicitly name them. The phrase 'in full' is slightly tensioned by the max_tokens bound, but the core action is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is implied: an agent should call this when it has a specific note id and wants that note's content, rather than searching or listing. There is no explicit when-to-use versus alternatives guidance, and no mention of when agtmem_search or agtmem_find would be a better choice. This is adequate but relies on inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agtmem_searchARead-only
Search long-term memory (hybrid FTS5 + trigram, RRF-merged). Returns short snippets with ids. Superseded notes are excluded. Use this first when resuming work on a project.
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | restrict to one note type | |
| limit | No | ||
| query | Yes | free text, any language | |
| scope | No | project scope, e.g. 'my-project' |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark readOnlyHint=true, so no extra safety disclosure is needed; the description adds useful behavior: RRF-merged hybrid search, snippet-only results, and exclusion of superseded notes. These details go beyond what annotations or schema state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences cover purpose, output, filtering, and usage context with no filler. Key behavioral facts are front-loaded before the usage recommendation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description supplies the essential return information ('short snippets with ids') and a key filter ('Superseded notes are excluded'). Combined with schema-described parameters and the read-only annotation, an agent has enough to call it correctly, though an explicit nod to alternatives like agtmem_find would complete the picture.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema descriptions already cover type, query, and scope, and limit has a default, so the parameter burden is largely carried by the schema. The description adds no parameter-level detail, but with roughly 75% coverage the gap is not critical; this is an adequate baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the operation ('Search long-term memory') and resource, augmented by the retrieval mechanism (hybrid FTS5 + trigram, RRF-merged). It also conveys the output form ('short snippets with ids'), so an agent can distinguish it from a full read tool. It does not explicitly differentiate from agtmem_find, which costs the fifth point.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use this first when resuming work on a project' gives concrete situational guidance and implies this is the entry point before reading or writing. It lacks an explicit when-not-to-use or named alternative, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agtmem_writeA
Create or update a note. Use type='decision' for choices and their rationale, 'fact' for durable truths, 'project' for project overviews, 'candidate' for unconfirmed guesses. Pass supersedes= to replace an outdated note — it is kept but marked superseded.
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | explicit id; defaults to a slug of title | |
| body | Yes | Markdown; ## headings become chunks | |
| tags | No | ||
| type | No | fact | |
| scope | No | global | |
| title | Yes | ||
| detail | No | one-line why/provenance | |
| origin | No | agent | |
| update | No | ||
| supersedes | No | id of the note this replaces |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full transparency burden. It discloses that superseded notes are kept but marked superseded, which is non-obvious. It implies mutation via 'update' and 'supersedes' but doesn't detail side effects like chunking behavior or scope/global implications. This is a reasonable disclosure for the context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense paragraph with highly relevant information, front-loading the purpose ('Create or update a note') and prioritizing usage guidance over form. Every sentence adds value: type semantics and supersedes behavior. It avoids fluff and is appropriately concise for a tool with 10 parameters.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 10 parameters and moderate schema coverage, this description is reasonably complete for a write tool without an output schema. It explains the critical 'type' and 'supersedes' parameters and provides usage scenarios. However, it omits behavior of 'update' flag and how 'scope' affects visibility, but these are less critical for typical use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 40%, so the description must compensate. It explains the 'type' parameter semantics (decision, fact, etc.) and supersedes, but doesn't cover many other parameters (id, tags, scope, update, origin, detail) that have minimal schema descriptions. The description adds some value but doesn't fully bridge the gap for all parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool creates or updates a note, with specific resource ('note') and verb ('Create or update'). It differentiates from siblings like agtmem_search and agtmem_read by focusing on write operations, and distinguishes note types and supersede behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides guidance on when to use each note type (decision, fact, project, candidate) and how to use supersedes to replace outdated notes. It implies write vs read alternatives (search/read for retrieving) but doesn't explicitly state when not to use this tool compared to append or bug, leaving a minor gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
8 tool updates
v0.1.0- First observed
agtmem_append - First observed
agtmem_bug - First observed
agtmem_candidates - First observed
agtmem_find - First observed
agtmem_index - First observed
agtmem_read - First observed
agtmem_search - First observed
agtmem_write
TDQS
Scored across 8 tools
Each tool has a distinct primary role: search/read/write/append cover note lifecycle, while index/bug/candidates/find are specialized. The only mild overlap is between agtmem_search and agtmem_find, but their different targets (notes vs. code symbols) are clearly explained.
All tools share the agtmem_ prefix and use lowercase snake_case, giving a predictable surface. Most names are verbs (search, read, append, write, index, find), but 'bug' and 'candidates' are nouns, which is a slight deviation from the verb-action pattern.
Eight tools is well-scoped for a memory server: core note operations, search, maintenance, and specialized record types are each covered without redundancy. The count feels appropriate and not bloated.
The surface covers create, read, update, search, append, and supersede, which is solid for a memory store. There is no explicit delete or general list-all operation, but superseding and search compensate for most practical workflows.
Maintenance
Related MCP Connectors
Shared memory for coding agents. Stop re-explaining your codebase every session.
Shared, governed long-term memory for AI agents across tools and sessions via MCP and REST.
Shared memory for coding agents and their teams: docs, epics, tasks and decisions over MCP.
1Agent-native notes, tasks, dev-docs, vaults, sync & handoffs. MCP + OpenAPI dual surface.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceProvides AI agents with persistent, local, and shareable project memory by storing decisions and code context in a searchable SQLite index, supporting keyword and semantic search via MCP.39 PyPI3MIT
- AlicenseNot gradedqualityBmaintenanceProvides persistent, local-first memory for coding agents with Markdown as the source of truth, exposed via CLI, loopback API, MCP, and Codex hooks for context retrieval and durable writes.MIT
- AlicenseAqualityBmaintenanceProvides cross-session persistent memory for coding agents via MCP tools to store, retrieve, and manage notes with hybrid keyword/semantic search and automatic deduplication.8211 npmMIT
- AlicenseNot gradedqualityBmaintenanceProvides a local markdown-based memory system for coding agents through MCP, enabling search, add, register, inventory, sync, and ingest operations across user and project memory. It gives any agent tool a durable, provider-agnostic shared memory stored in folders you own.2MIT