claude-engram
This server provides persistent memory, session intelligence, and code-analysis tools for Claude Code and any MCP client.
Health & lifecycle: Check status via
claude_engram_status; load deep context withsession_start; get a session summary withsession_end.Memory management: Store, recall, search, hybrid-search, dedupe, archive, restore, and consolidate memories; add/manage permanent rules; track and acknowledge mistakes.
Work tracking: Log mistakes and architectural decisions with
work.Scope guard: Declare, check, expand, and clear allowed files/patterns for multi-file tasks.
Loop detection: Record edits/tests and check whether it is safe to edit again to prevent death spirals.
Context protection: Save/restore checkpoints, verify task completion, create/retrieve handoffs, and register critical instructions.
Conventions: Add, retrieve, check, and remove project-specific coding conventions.
Output validation: Check code for fake/silent failures and validate results against expected formats/contents.
Codebase exploration: Semantic code search (
scout_search), LLM-based code analysis (scout_analyze), file summarization, dependency mapping, and impact analysis before refactors.Quality & audits: Check code for AI slop, validate against conventions, audit files in batch, and find similar bug patterns with regex.
Session mining: Search past conversations, find decisions, replay file discussions, identify struggles/errors/correlations, build timelines, get summaries/overview, predict edit context, cross-project insights, reflect on patterns, and reindex mining data.
Claude Engram
Persistent memory and session intelligence for Claude Code. Engram hooks into the session lifecycle and tracks mistakes, decisions, edits, tests and context on its own. It mines your full session history so past work comes back when it is relevant. It keeps long sessions safe across compaction. And it brackets Claude Code's /goal loop so an unattended run cannot poll all night, lose its record, or die silently.

Most of it runs through hooks with nothing to call. The rest is MCP tools, so any MCP client can use the tools, and Claude Code gets the hooks as well.
Contents: What runs on its own · Install · Day to day · The tools · Memory · Checkpoints and compaction · Unattended runs · Rules with detectors · Project defaults · Run report · Session mining · Configuration · Storage · Compatibility · Benchmarks
What runs on its own
Every row is a hook. You call nothing.
Moment | What engram does |
Session start | Prints the banner: the project's rules (inherited ones marked), its own past mistakes with a count of the pooled ones (every count names the project it was read for, the same store the prompt hook counts), the latest deliberate checkpoint with its |
Every prompt | Captures decisions from what you type ("let's use X", "switch to Y", "from now on always Z") by semantic scoring with a regex fallback, then a shape gate shared with the transcript miner: a declarative sentence with a deciding word, never a question, an acknowledgement, a count or a table row. Questions and requests for information are not captured. Delivers any staged nudge |
Before an edit | Injects the three memories most relevant to the file (a rule only when it names the file, an entry older than a month only when it names the full path), warns before a past mistake tied to that file (a code exception is never predicted for a non-code file), warns after three edits to the same file in one session (edit loop), verifies that a proposed import resolves against the per-project code index ( |
Before a read | Once per file per session: an orientation from the code index plus the file's most relevant memories ( |
Before a shell command | Matches the command against every rule that carries a detector and injects the rule before the command runs ( |
After an edit | Counts the edit for the loop warning and prints nothing. Files under |
After a shell command | Tracks test runs and says so only on the first result and on each flip, with a status that lasts the whole session (the output counts when the command names a test runner, or when nothing in the chain merely reads: |
After any batch of tool calls | Accounts the calls to the turn for stall detection and records detector matches on non-shell tools (path globs on edits, MCP tools by name) |
After a failed tool | Logs the error as a mistake, unless it was a failing test run. When the failure matches a recurring error from past sessions or a stored mistake, injects the past fix at once ("Deja vu: TypeError hit in 3 past sessions, fix: ...") |
After ExitPlanMode or a task update | Asks for the approved plan to be banked as a checkpoint with its steps pending. A task marked completed counts as a milestone claim |
Stop (end of a turn) | Saves the handoff (last message, files), reads the final message for a completion claim without a checkpoint behind it, judges the turn for stall detection, brings the |
Before compaction | Writes the automatic checkpoint (the floor) |
After compaction | States the pressure rhythm (heads-up, checkpoint, auto-compaction). The session-start banner that follows re-injects the rules, the mistakes and the checkpoint banked before the compaction |
An API failure | Records the failure with its type and, for a usage limit, the reset time. Sends an alert during a goal run |
A notification | A permission prompt, a needs-input or an idle notification during a goal run sends an alert |
Session end | Writes the session summary, the run report for a substantial session, the rotation plan, and spawns the background miner |
Subagents get none of the injections, to save their context, but their edits are still tracked so the parent session knows what changed. Every hook has a one to two second budget and fails silently past it. High-frequency hooks are served by a resident daemon with warm imports and fall back to in-process when it is down.
Related MCP server: io.github.420247jake/session-forge
Install
git clone https://github.com/20alexl/claude-engram.git
cd claude-engram
python -m venv venv
source venv/bin/activate # or venv\Scripts\activate on Windows
pip install -e . # Core
pip install -e ".[semantic]" # + embedding model for vector search and semantic scoring
python install.py # Hooks, MCP server launcher, /engram skill, migrationsinstall.py creates the storage directory, merges engram's hooks into ~/.claude/settings.json without touching other hooks, writes a .mcp.json in the repo for copying (and a launcher script as the fallback when no venv is active), pre-builds the decision template cache for semantic scoring, installs the /engram skill to ~/.claude/skills/engram/, and runs the data migrations. Ollama is optional. Without it the tools that use a local model degrade silently and everything else is unaffected.
Then, per project:
python install.py --setup /path/to/your/projectOr copy .mcp.json into the project root. That is the only per-project file. Hooks and the /engram skill are global. The CLAUDE.md in this repo is for people working on engram itself, and your projects do not need it.
Open the project in Claude Code and approve the claude-engram MCP server when prompted. If the project already has Claude Code history, the first session finds it and mines it in the background.
To update:
cd claude-engram
git pull
pip install -e ".[semantic]" # Reinstall if dependencies changed
python install.py # Re-run to update hooks and /engram skillHooks pick up code changes at once because the install is editable. Reconnect the MCP server with /mcp to reload it. Data migrations run on their own and are forward-only, idempotent and safe to downgrade across.
Day to day
Most of the time there is nothing to do. The few things worth doing on purpose:
Type
/engramwhen you want Claude to reach for the tools itself. Background tracking runs either way.Half remember something from weeks ago? Ask Claude to mine the sessions for it.
session_mine(search)reads everything you ever discussed, including sessions long gone from context.Something Claude must never forget goes in as a rule with
memory(add_rule). Rules in a project stay local. Rules at your workspace root cascade to every project under it.Before a compaction engram saves a checkpoint on its own, but a deliberate
context(checkpoint_save)with what you are doing and what is left resumes far cleaner. Engram nudges for one at the right moments (see Checkpoints and compaction).On return, ask what you said you would do this session.
session_mine(commitments)reads the live transcript for open loops, which the post-session index cannot see.When a plan should run without you, ask for
/engram run <goal>. Claude hands you the/goalline to type, and engram brackets the run (see Unattended runs).
The tools
Every tool takes project_path. Tools marked read-only never write anything.
Tool | Operations | What it is for |
|
| Store a discovery, a rule (with an optional detector), find memories by query, file or tags, manage mistakes, and keep the store tidy. |
|
| Record a mistake the auto-capture missed (description, file, how to avoid) or an architectural choice (decision, reason, alternatives) |
|
| Checkpoints and handoffs are one construct, a durable per-project ring of the last 20 deliberate saves. A save carries the task, current step, completed and pending steps, files, and optional handoff fields (summary, context needed, warnings). A restore takes |
|
| Everything derived from the transcripts. See Session mining |
|
| Declare the files and patterns a task may touch. Edits outside it are flagged before they happen |
|
| A file's imports and, with |
|
| Dependents, exported symbols at risk, suggested test targets and a risk level before a refactor |
|
| Semantic search over the codebase. Uses the local model when Ollama is up |
|
| Purpose, exports, dependencies and complexity from structural analysis |
|
| Audit files on disk for bugs, missing error handling, security issues and TODOs, or lint an inline snippet for long functions, vague names and deep nesting |
|
| Every occurrence of a regex bug pattern across the codebase, with context |
|
| Per-project coding conventions with categories and reasons, checked against code or a filename by pattern matching |
| Health: version, model, the scorer daemon's device, memory counts | |
| Rarely needed. The hooks run these on their own. Call them only for an explicit deep reload, a summary, or an impact check by hand |
Examples:
memory(operation="remember", content="The scorer must run on one pinned thread", project_path="/path")
memory(operation="add_rule", content="Never push without asking", reason="Owner decides what leaves the machine", project_path="/path")
memory(operation="hybrid_search", query="auth token refresh", project_path="/path")
memory(operation="list_mistakes", project_path="/path")
work(operation="log_decision", decision="Keep the ring per project", reason="Workspaces must not clobber each other", alternatives=["one global ring"])
context(operation="checkpoint_save", task_description="OAuth2 migration", current_step="Step 3: token refresh",
completed_steps=["Step 1: provider config", "Step 2: login flow"], pending_steps=["Step 3: token refresh", "Step 4: tests"],
files_involved=["auth.py", "oauth.py"], handoff_summary="60% done, refresh next",
handoff_warnings=["legacy auth.py is still used by mobile"], project_path="/path")
context(operation="checkpoint_restore", project_path="/path", index=1)
context(operation="checkpoint_list", project_path="/path")
session_mine(operation="search", query="why did we drop the cache", kind="decision", since="2026-04-01", project_path="/path")
session_mine(operation="commitments", project_path="/path")
deps_map(symbol="load_project_memory", project_path="/path")Memory
A correction typed once in one session is mined from its transcript and put in front of Claude before the next edit of the file it names.

Six categories:
Category | Purpose | Protected | Captured |
| Project rules that always apply | Never archived or decayed | By hand with |
| Curated insights synced from lesson files | Never archived; the source file owns the lifecycle | Opt-in, the miner syncs |
| Errors to avoid repeating | Manual ones never archive. Stale machine-written one-offs (three weeks old, never recurred, away from current work) auto-archive. All restorable | Failed tools, transcript mining, |
| Choices and their reasoning | No | Prompts, transcript mining, |
| Facts about the codebase | No | By hand, |
| Session notes | No | By hand; |
Two tiers. The hot tier (memory.json) holds rules, mistakes and recent memories and is what the hooks load on every tool call. The cold tier (archive.json) holds old inactive memories, searchable with archive_search and restorable by id, and is never on the hot path. Memories archive after 14 days without access (CLAUDE_ENGRAM_ARCHIVE_DAYS). A relevance of 8 or more exempts a memory from age archiving, which sits above every default (a manual remember is 5, the miner mints decisions at 7). Nothing is deleted without review: cleanup archives before it deletes.
Before every edit the hot memories are scored against the file: 35% file path match, 20% tag overlap, 20% recency (30-day decay), 15% importance, 10% access frequency, plus a bonus of 0.3 for rules, 0.25 for lessons and 0.2 for mistakes. Path matching is path-aware, so a shared basename across diverging paths is not a match, and generic names (__init__.py, index.js, README.md, main.py) need a full-path signal to score. The top three are injected. The outcome feedback loop applies a bounded multiplier (0.8 to 1.2) to the kinds of injection that precede passing tests.
In a workspace with several projects, memories are scoped to the sub-project of the file being edited, detected by markers like pyproject.toml, package.json, .git or CLAUDE.md. Rules and mistakes at the workspace root are inherited by every project under it, and memory(list_rules) marks the inherited ones. A session started inside a git worktree, or under node_modules, a virtualenv or a directory you list in non_project_dirs, belongs to the repository above it: its rules, its checkpoint ring, its patterns. Mined mistakes and decisions are filed under the registered project whose files they name, so list_mistakes, recall and the banner are per project.
Checkpoints and compaction
Hooks never see context usage, but the statusline does. Engram mirrors the statusline's token counts to a per-session file and computes the distance to the point where auto-compaction fires: CLAUDE_CODE_AUTO_COMPACT_WINDOW, else the --autocompact launch flag of a background job, else autoCompactWindow in settings (the env var, then the flag, then managed settings, then project and user settings, in Claude Code's own order), else the model default. Never the raw percentage.
Distance to the compaction point | Engram injects |
About 10% of the window out | A heads-up: finish the current step, start nothing long |
20K tokens above the auto-compaction trigger (10K on a 200K window) |
|
60 turns with neither a checkpoint nor a completed step | A fallback reminder |
Auto-compaction does not fire at the configured number. Claude Code keeps room for the model's output first, so the trigger sits about 32K tokens below the setting (a 750K setting compacted at 717,578). Engram places the last band above that measured trigger, not a fraction below the setting. Each band fires once per compaction cycle. After a compaction the banner restates the rhythm and shows the checkpoint banked before it.
Milestones are Claude's to call. When it judges a phase, a step or part of a plan done, the rule in the skill says checkpoint before saying so. Engram reads the final message at Stop, and a completion claim ("Phase 1 built", "step 3 done, next is X") with no deliberate checkpoint that turn gets one reminder quoting the sentence, but only when that turn also changed something (an edit, a commit, a delegated agent), never for a list bullet, and at most once an hour. A status report to you is not a close. Questions, negations and future tense never fire. A turn that banks state with memory(remember) and no checkpoint gets a different reminder: a remember stores a fact, the restore reads checkpoints only, and both together are fine. After ExitPlanMode it asks for the approved plan to be banked with its steps pending. Nothing is written for the model, and commits are never a trigger.
Every deliberate checkpoint records the commit the repo stood at and the active /goal, if any. A restore, whether the banner or context(checkpoint_restore), prints the goal and one staleness line: commits and files changed since the save, the checkpoint's own files among them, or that its commit is not in this history. A handoff carries the last session's framing, and that line is what to read before acting on it. A save with handoff content also writes a readable HANDOFF.md next to the store.
The nudges need a statusline that records the mirror. Either use engram's, which prints Fable 5.1 | ctx 660K/1000K | compact at 750K (90K left) | $1.25 | 5h 93% | proj (the usage window only on a Claude.ai plan):
"statusLine": {"type": "command", "command": "<venv python> -m claude_engram.hooks.context_pressure statusline"}or have your own script call claude_engram.hooks.context_pressure.record_statusline(data) with the JSON it received, or write the same record to ~/.claude_engram/sessions/<session_id>.ctx.json yourself (eight flat fields, listed in the module docstring; no engram import needed). Without a statusline engram says so at session start and only the turn cadence runs. The --autocompact launch flag is not in a hook's environment: for a background job (claude --bg) engram reads it from the job's saved launch flags (~/.claude/jobs/<id>/state.json); in an interactive session a window set only by the flag is invisible, so engram falls back to the model default and names the source it used.
The setting is a token count capped at the model's window. /autocompact 750k is 75% of a 1M model and the whole window of a 200K one, which leaves no room for the checkpoint call. Engram tells you once, at the first reading, when the number does not fit the model and names one that does. For long unattended runs, set the point yourself: CLAUDE_CODE_AUTO_COMPACT_WINDOW=750000 on a 1M model keeps turns cheaper and puts the checkpoint nudge a known distance below a number you chose.
On a Claude.ai plan the statusline also carries the 5-hour and 7-day usage windows, and engram mirrors them beside the token counts. At 90% of the 5-hour window (95% of the weekly) it says so once per window, with the reset time, and asks for a checkpoint and a park on a wakeup until the reset. A limit hit ends the run on an API error that no Stop hook sees and that Claude Code does not retry, which is why the warning comes first. The StopFailure hook records every API failure with its type and the reset time, and the run report shows it. API-key sessions get no windows and no nudge. A custom statusline mirrors the windows by copying rate_limits.five_hour and seven_day (used_percentage, resets_at) into the record as five_hour_pct, five_hour_resets_at, seven_day_pct and seven_day_resets_at.
Unattended runs
Engram does not invent a loop. /goal is the loop, and engram brackets it.
/goal <condition> # Claude Code's own loop; engram brackets it from the next stop
/goal clear # you end it; engram cannot
/engram run <goal> # Claude hands you the /goal line to type, then waits
/engram status [<session>] # strikes, halt, alerts, and the report once the run ended
/engram report [<session>] # the run report, summarized
/engram release <session> # lift a halt, on your word
python -m claude_engram.run --headless --goal "<goal>" --alert-command 'curl -s -d {message} https://ntfy.sh/<topic>' # cron and overnight onlyYou type /goal, in the session you are in. Nothing else sets a goal: no flag, setting, hook output or tool call, and Claude cannot type a slash command. When a plan should run unattended, Claude hands you the exact /goal line and waits. From the next stop engram sees the goal in the transcript and brackets it for the goal's life:
Rules with deny detectors refuse their commands instead of only showing the rule.
Three no-effect strikes halt every tool (see below), and alerts go out.
The goal is stamped into every checkpoint, and the compaction and budget nudges do their work.
The turn cap (150 by default,
goal_turn_cap) arms the same halt, since engram cannot end a goal loop, and the alert asks you to/goal clear.Engram never judges the goal. Claude Code's evaluator does, reading what Claude shows in the conversation, so write the goal as something Claude's own output can demonstrate.
The run ends when the evaluator says met (the goal clears itself), judges it impossible, or you clear it. Engram writes a manifest at the start (goal, the rules in scope with their detectors, the start commit), one alert at the end, and a run report with a "Goal run" section: the goal, turns against the cap, verdicts and how it ended. session_mine(run_status) shows the bracket's state at any time.
Stall detection
/goal stops its own loop after several turns with no tool use. The overnight failure that matters is the other one: a run that uses a tool every turn and changes nothing. Seen live, a goal loop ran nine turns of cat on a file that was never going to change. Engram judges a turn by its effect:
Turn | Counts as |
A file changed (an edit, a mutating shell command, or the git working tree moved), a test status flipped, a commit landed, work was delegated to an agent | Good |
Tools were used and none of that happened | No effect |
No tools at all, a park on a wait primitive (Monitor, ScheduleWakeup, a cron, a background task), or a test run that was not a repeat of the last one | Neutral |
Turns count, not time, so a four-hour foreground command is one turn. Verification is neutral. The same test command with the same verdict again is the re-run-and-hope pattern and counts as no effect.
Three consecutive no-effect turns are one strike. Strike 1 names the pattern and asks for a wait primitive instead of polling. Strike 2 re-injects the latest checkpoint and the rules and asks for a bearings check (task, last real change, blocker, different action). Strike 3 is the cap, and under a goal it becomes the halt. Strikes decay rather than reset: five consecutive good turns remove one. Every increment and decrement is an event with its turn number in the run report. Tune with CLAUDE_ENGRAM_STALL_TURNS, CLAUDE_ENGRAM_STALL_DECAY and CLAUDE_ENGRAM_STRIKE_CAP.
Halt and alerts
At the strike cap every tool call is denied with a reason the model sees, except PushNotification, ToolSearch, SendMessage and engram's checkpoint call, so the model can leave a record and a one-line notification, then stop. A subagent already running is never denied: it shares the session id, the halt is for the main loop, and it reports back with SendMessage. A deny ends the turn, and with no tool use the goal's own stall rule closes the loop. Engram never blocks a Stop of its own. A person lifts the halt with /engram release <session> or python -m claude_engram.hooks.stall release <session>, and strikes reset. python -m claude_engram.hooks.stall status <session> prints strikes, the halt and the alerts.
A halt, an API failure, a run waiting on a person (a permission prompt, needs input, idle), the launcher pausing for a usage limit, and the run ending each send one line under 200 characters through the command you configure: alert_command in .engram/config.json, CLAUDE_ENGRAM_ALERT_COMMAND, or --alert-command. {message} is replaced, or the message arrives on stdin. No service is built in. Every alert lands in the run report whether or not a command was configured. PushNotification is the model's tool, and on a headless run the desktop notification has nowhere to land, which is why the command exists.
Cron and overnight
Outside any session, the launcher runs claude -p with the same supervision. It writes the manifest, sets CLAUDE_CODE_AUTO_COMPACT_WINDOW to 75% of the model's window (--window-fraction, --context-window), a hard --max-turns (150), the park instruction (--no-park-hint drops it) and bypassPermissions because nobody is there to answer a prompt (--permission-mode). After a usage-limit exit it sleeps until the window resets, plus CLAUDE_ENGRAM_RESUME_SLACK seconds, and resumes the same session, up to --max-resumes (3) times. --model, --project, --prompt or --prompt-file for the task text after the goal, --session-id, and --dry-run to print the plan. Cost is the same as any session: a Claude.ai plan pays with its usage windows, an API key per token; the total_cost_usd a headless run prints is Claude Code's own client-side estimate.
A goal with nothing to do burns turns: a not-met goal re-prompts at once, so it starves /loop and cron (seen live: nine verdicts in two minutes, a cron fire lost). A goal idles while it waits on a Monitor, a wakeup or a background task, and the loop fires then. Tell the model to park on one of those instead of polling. From Git Bash on Windows set MSYS_NO_PATHCONV=1, or the leading /goal is rewritten into a filesystem path and Claude gets a plain prompt.
Rules with detectors
Rules are natural language and stay that way. A rule may carry a detector, hand-written and never inferred:
memory(add_rule, content="Never push without asking", detector={"tools": ["Bash"], "command": "\\bgit\\s+push\\b", "note": "push"})
memory(set_detector, memory_id=<rule id>, detector={"paths": ["design/**", "*.pem"]}) # {} clearsA detector has tools (tool names; empty means any), command (a regex against a shell command), paths (globs against an edited path), input (a regex against the JSON of the tool input), note (for the report) and optionally "unattended": "deny". With one, every matching call is recorded in the run report with the turn, an excerpt of the input and a verdict that is only what the hooks can see: unattended (bypass, auto or dontAsk mode, so no person approved the call), prompted (Claude Code's own permission prompt stood between the model and the call) or plan. A matching shell command also gets the rule injected before it runs. A rule without a detector is advisory, and the report says so per rule. Detector health is shown too, so a regex that stopped compiling is visible. No language model sits in this path.
"unattended": "deny" refuses the command during a goal run, with the rule as the reason, and the model is told to record what it needs approved, notify, and continue with allowed work. An attended session, even in bypass mode, is only shown the rule. The default pack ships deny detectors on destructive shell commands (recursive or forced deletes, hard resets, force-push, DROP and TRUNCATE, disk formats, kills), kill by image name, and anything that leaves the machine (push, pull request, publish, outbound POST). A project whose own rule already covers that ground adopts the pack detector if it has none. "compliance": false in .engram/config.json or CLAUDE_ENGRAM_COMPLIANCE=off turns the trail off.
Project defaults
Engram ships an opinion about how a project is kept. Everything is on by default and <project>/.engram/config.json turns each piece off:
{"rotation": "auto", "default_rules": false, "workflow_rules": true, "code_rules": true, "structure": true,
"compliance": true, "alert_command": "", "goal_turn_cap": 150,
"rotation_log_days": 30, "rotation_learn_days": 90, "rotation_learnings_days": 0, "rotation_learn_max_lines": 500}Structure: a project (a git repo, a package manifest or a CLAUDE.md) gets missing pieces scaffolded: a headered CLAUDE.md with Purpose, Testing and Structure, .learnings/ERRORS.md, .learnings/LEARNINGS.md and session-logs/. A multi-person layout (.learnings/<name>/, session-logs/<name>/) is left as it is.
Rules: three tiers seeded once into the project's memory on its first session. Universal: no destructive commands without asking, search first, quality over speed, prerequisites first, be direct, try before asking, private stays private, push back, follow the plan in order, never kill by image name, session maintenance, checkpoint when you judge a step done. Workflow: plan before code; proposed, open and decided are three things; every milestone has a gate before and a verdict after; delegate by size with a hard agent budget; verify before claiming done; numbers get a source; anything that leaves the machine is the owner's decision; write the learning when it happens. Code: one function, one purpose; names and comments that explain why; nothing left lying around; verify by quoting what it printed; never swallow an error on the decision path; same inputs, same outputs; the real thing over a stand-in; no secrets anywhere; performance from the start; both Windows and Linux; the fast path on purpose; never do work twice; smart over busy. The workflow and code text also land in the scaffolded CLAUDE.md. Any rule the project or an ancestor already has in substance is skipped, so your own rules win. memory(list_rules) shows them.
Rotation: nothing is deleted, ever. Session logs older than 30 days move to session-logs/archive/<month>/ with a monthly digest beside them. Dated ERRORS.md entries older than 90 days move to .learnings/archive/ERRORS-<year>.md. LEARNINGS.md holds patterns that do not age out, so it rotates only when over the cap. Either file over 500 lines sheds its oldest 30-day-plus entries until it fits. Undated and STANDING entries never move, and each trimmed file gets a one-line note under its title saying what moved and where. Per-person folders rotate inside themselves. By default SessionEnd only plans and SessionStart announces. session_mine(rotate, dry_run=false) applies, "rotation": "auto" applies at every session end, "rotation": false stops planning.
Run report
Every substantial session leaves one auditable record in the repo: <project>/.engram/runs/<date>-<session>.md plus a .json twin, written at SessionEnd or on demand with session_mine(run_report) or python -m claude_engram.run_report --session <id>. Add .engram/runs/ to the project's .gitignore, since the reports carry session ids, costs and local paths; .engram/config.json is the file worth tracking. Every line is hook-captured or read from the transcript, never self-reported by the model:
the
/goalcondition, every evaluator verdict with its reason, and the outcome; model, permission mode, branch, start and end commitwall time, turns, prompts, context at end and cost
every compaction with its trigger, before and after token sizes, and which checkpoint it restored
files touched with per-file edit counts; test runs, first and last status
errors grouped by signature, recurrences, and whether the miner already knew them
checkpoints written this session, deliberate versus automatic
stalls: turns with and without effect, every strike and decay with its turn number
rule matches with their verdicts, and detector health
alerts and API failures
what was not measured, listed rather than omitted
Session mining
The hooks capture what happens inside a session. The miner reads the transcripts for what they cannot see. After every session a background process indexes the session, extracts decisions, mistakes and approaches, and builds search embeddings. During a session a debounced live tick at turn end refreshes the same index, so session_mine(search) sees this session's earlier work. Recurring errors and struggles are attributed to sub-projects and filtered by what the last session touched, and errors quiet for 30 days drop out. The lessons bridge, opt-in through lessons_globs, syncs dated entries in curated markdown as protected lesson memories with code-index triggers, so a lesson naming a module surfaces when editing files that import it.
session_mine operations:
search(query, method=hybrid|semantic|keyword, kind=decision|next-step|error|narration, since, until): every past conversation, tool content includeddecisions(query): when and why a decision was made, from the transcripts and from the repository's own history (git log -Son the query and its most specific tokens, with the commit messages)replay(file_path): discussions about a file, followed by the commits that touched it;predict(file_path): the context an edit will needstruggles,errors,correlations(files always edited together),timeline,summaries,overview,status(index coverage),cross_projectreflect: which injection kinds precede passing tests, plus insights from recurring mistakes synthesized by the local modelcommitments: what you said you would do this session and whether it is done, from the live transcriptrun_report,run_status,rotate(dry_run),reindex(mode=post_session|bootstrap|full)
Claude Code files a transcript under the directory the session was started from. A session run from a workspace root indexes under the root, so the views for a sub-project (overview, timeline, struggles, errors, search, reflect) are the workspace's views, and the counts include every sibling project worked on from that root. Start the session inside a project when you want its history alone. Memory works the other way: each mined mistake and decision is filed under the registered project whose files it names, and one that names no file, or only files that cast no vote (relative traceback paths, files outside the root), goes where the session's own edits point.
If search quality degrades or after a big update:
python scripts/reindex.py "/path/to/your/workspace" --force # rebuild search index
python scripts/reindex.py "/path/to/your/workspace" --force --extract # also re-extract decisions/mistakesOr through MCP: session_mine(operation="reindex", mode="bootstrap"). On a large history this can exceed Claude Code's two-minute MCP call limit and continue in the background; the rebuild still finishes.
Configuration
All optional. The library book has the detail on each.
Variable | Default | Description |
|
| Storage location, also the test-isolation seam |
|
| Ollama model. Only |
|
| Ollama API URL |
|
| Timeout in seconds for local model calls |
|
| How long Ollama keeps the model loaded after a call ( |
|
| Embedding model (about 1.1GB scorer RAM). |
| model native | Matryoshka truncation dim. Stores are signature-stamped and rebuild on a model change |
| smart | Unset: the daemon stays on cpu and bulk jobs use a transient GPU worker that exits after the job. |
|
| Job size in texts that routes to the GPU worker |
|
| Rows per forward pass on the GPU (about 26 MiB per row) |
|
| Rows per forward pass in the resident daemon on the CPU. The daemon keeps the activation arena of its largest batch for life: 64 rows parked 1.2 GB more than 16 at the same speed |
|
| Scorer daemon idle timeout in seconds |
| unset | Set to run every hook in-process and never start the daemon (tests, benches) |
|
| Live mining tick interval in seconds. |
|
| Days until inactive memories archive |
|
| Prune session-search shards older than N days |
| unset | Mirror the last-read file path to this file (statusline integration) |
|
| Heads-up nudge this fraction of the window before the compaction point |
|
| Tokens Claude Code keeps below the setting before it compacts |
|
|
|
|
| Turns without a deliberate checkpoint or a completed step before the fallback reminder |
|
| Usage-window warning thresholds |
|
| No-effect turns per strike, good turns per decay, strikes to the cap |
|
| Turns under a |
| unset |
|
| unset | Shell command for alerts; |
|
| Seconds the launcher waits past a usage-window reset before resuming |
| on | Environment overrides for the project config switches ( |
| unset | Comma-separated directory names that are never a project (adds to |
| unset |
|
| unset | A file path; every git call the hooks make is logged there |
~/.claude_engram/config.json accepts embed_model, embed_dim, lessons_globs (a list of globs, for example ["docs/lessons/*.md"]; no default path) and non_project_dirs (directory names that are never a project of their own, added to the built-in node_modules, .venv, venv and __pycache__; a workspace with a .scratch/ convention lists it here). <project>/.engram/config.json holds the per-project switches listed under Project defaults.
Storage
Everything lives under ~/.claude_engram/ (CLAUDE_ENGRAM_DIR): manifest.json, projects/<hash>/ with memory.json, archive.json, embeddings.npy, session_index.json, extractions/ and the checkpoint ring, checkpoints/ for the per-task files, and sessions/<session_id>.json for each session's working state, so concurrent sessions never clobber each other. All writes are atomic (temp file, then replace). The scorer daemon serves the embedding model over localhost, stays on cpu with zero VRAM parked, runs every model call on one pinned thread, and exits after 30 minutes idle or when the package source changes. Ollama, when present, is used only by scout_search, memory(consolidate) and session_mine(reflect).
Compatibility
Platform | What works | Auto-capture |
Claude Code (CLI, desktop, VS Code, JetBrains) | Everything | Hooks and session mining |
Cursor, Windsurf, Continue.dev, Zed, any MCP client | MCP tools | No hooks |
Obsidian vaults | Everything, with a | With Claude Code |
Benchmarks
Retrieval (recall at k): LongMemEval 0.966 R@5 and 0.982 R@10 over 500 questions, ConvoMem 0.960 over 250 items, LoCoMo 0.649 R@10 over about 2k questions. About 43ms per query, 112ms cross-session over 7,310 chunks.
Product behavior, integration suites green: decision capture 97.8% precision, error auto-capture 100% recall, compaction survival 6/6, multi-project isolation 11/11, edit-loop detection 12/12, session mining 64/64, Obsidian vault compatibility 25/25.
Reproduce the retrieval numbers with tests/bench_longmemeval.py, tests/bench_convomem.py and tests/bench_locomo.py; the product suites are the other tests/bench_*.py scripts, listed in the library book's contributing chapter.
Documentation
The library book holds the design, the internals, the full usage guide, the API reference, the gotchas and the changelog. /engram is the quick reference that install.py installs. The CLAUDE.md in this repo is the developer's map of the hooks and modules.
demo/ holds the storyboard and the real hook output behind the two gifs. Regenerate it with venv/Scripts/python.exe demo/build_fixture.py, then vhs demo/engram.tape and vhs demo/memory.tape in WSL or Linux, then the 12 fps re-encode in demo/README.md.
License
MIT
Available Tools
21 toolsaudit_batchC
Audit multiple files for issues. Supports glob patterns.
| Name | Required | Description | Default |
|---|---|---|---|
| file_paths | Yes | ||
| min_severity | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Only mentions glob pattern support. No disclosure of read-only nature, permissions, or other behavioral traits. With no annotations, description should carry this burden but fails to.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, efficient and front-loaded. No unnecessary words, though slightly under-detailed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Missing return value description, behavior for empty paths, error cases, and meaning of min_severity. Given no output schema and no annotations, description is insufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Glob pattern support adds value for file_paths but min_severity is completely unexplained. With 0% schema coverage, description should compensate more.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states it audits multiple files for issues, with glob support. Distinguishes from sibling audit tools like code_quality_check by focusing on file batches.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives like code_quality_check or scout_search. Missing context about prerequisites or selection criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
claude_engram_statusB
Check Claude Engram health. Returns: status, model, memory stats.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full burden. It does not disclose whether the tool is read-only, expensive, or requires specific permissions. The return fields are listed but behavioral traits are absent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise: one sentence plus a short list of return items. Every word serves a purpose with no fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no parameters and no output schema, the description is moderately complete. It explains the purpose and return fields but lacks detail on the format or semantics of each return value, which an agent might need for interpretation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline is 4. The description adds meaning by stating what the tool returns, which compensates for the lack of parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'Check' and the resource 'Claude Engram health', and lists the return fields (status, model, memory stats). It is specific enough to distinguish from sibling tools like 'code_quality_check' or 'memory', though it does not explicitly differentiate itself from them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus its siblings or alternatives. With a list of 20 sibling tools, explicit usage context is missing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
code_pattern_checkC
Check code against stored conventions using LLM.
| Name | Required | Description | Default |
|---|---|---|---|
| project_path | Yes | ||
| code | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description bears full responsibility. It only mentions 'using LLM' but fails to disclose whether the tool is read-only, destructive, or requires permissions, nor does it explain what 'checking' entails (e.g., does it modify anything?).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, which is concise, but it sacrifices valuable detail. It does not earn its place because it leaves critical information out.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and minimal schema descriptions, the description is severely incomplete. It does not explain what the result of the check looks like or how the tool integrates with other tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not elaborate on the parameters 'project_path' or 'code'. The agent cannot infer what values are expected or how they relate to the checking process.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Check'), the resource ('code against stored conventions'), and the method ('using LLM'). This is specific and distinguishes it from sibling tools like 'code_quality_check'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like 'code_quality_check' or 'audit_batch'. There are no prerequisites, exclusions, or use-case hints.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
code_quality_checkC
Check code for AI slop: long functions, vague names, deep nesting.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | ||
| language | No | python |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description must fully convey behavioral traits. It only states the purpose but does not disclose whether the tool modifies code, requires specific permissions, or has any side effects. The word 'check' implies read-only, but this is not explicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no fluff, making it easy to scan. It is appropriately sized for the tool's simplicity, though it could benefit from slightly more detail in a structured format.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of output schema and annotations, the description is too minimal. It does not explain what the tool returns (e.g., a list of issues or a score), or any constraints like maximum code length. This leaves the agent guessing about the tool's full behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must add meaning. However, it provides no additional context for the 'code' or 'language' parameters beyond their names. No format, size limits, or examples are given.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it checks code for 'AI slop' including specific issues like long functions, vague names, and deep nesting. This verb+resource combination is specific and distinguishes it from siblings like 'code_pattern_check' which likely targets different patterns.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives such as 'code_pattern_check' or 'scout_analyze'. No prerequisites or context for invocation are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
contextA
Context protection for long tasks. Operations:
checkpoint_save: Save task state (task_description, current_step, completed_steps, pending_steps, files_involved)
checkpoint_restore: Restore last checkpoint (task_id optional)
checkpoint_list: List saved checkpoints
verify_completion: Claim task done + verify (task, evidence, verification_steps)
instruction_add: Register critical instruction (instruction, reason, importance)
instruction_reinforce: Get instructions to remember
handoff_create: Create session handoff (handoff_summary, pending_steps, handoff_context_needed, handoff_warnings)
handoff_get: Retrieve latest handoff document
| Name | Required | Description | Default |
|---|---|---|---|
| operation | Yes | Operation | |
| task_description | No | ||
| current_step | No | ||
| completed_steps | No | ||
| pending_steps | No | ||
| files_involved | No | ||
| task_id | No | For restore: specific checkpoint | |
| task | No | For verify: task to verify | |
| evidence | No | For verify: proof | |
| verification_steps | No | For verify: checks | |
| instruction | No | For instruction_add | |
| reason | No | ||
| importance | No | ||
| project_path | No | ||
| handoff_summary | No | For handoff_create: summary for next session | |
| next_steps | No | For handoff_create: what to do next | |
| handoff_context_needed | No | For handoff_create: context the next session needs | |
| handoff_warnings | No | For handoff_create: warnings for next session |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden for behavioral transparency. While it lists operations and their parameters, it does not disclose side effects, persistence behavior, or resource constraints, leaving significant gaps for an agent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured with a lead sentence followed by a bulleted list of operations. It is relatively concise given the number of operations, though some repetition (e.g., 'For verify') could be streamlined.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description provides a good overview of operations and their parameters, but it lacks details on output/return values, error handling, and prerequisites. For a complex tool with 18 parameters and no output schema, this is adequate but not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds meaningful context for many parameters by linking them to specific operations (e.g., 'For handoff_create: summary for next session'). This goes beyond the schema's brief descriptions, especially for the structured parameters like arrays, though some parameters like 'reason' lack context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose as 'Context protection for long tasks' and enumerates eight specific operations with brief descriptions, making it easy to understand what the tool does and how it differs from sibling tools like 'memory' or 'output'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage through operation names and parameter details (e.g., when to use checkpoint_save vs checkpoint_restore), but it does not explicitly state when to use this tool over alternatives or provide exclusion criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
conventionB
Project conventions. Operations:
add: Store rule (project_path, rule, category, reason, examples, importance)
get: Get rules (project_path, category)
check: Check code/filename (project_path, code_or_filename)
remove: Remove convention by matching text (project_path, rule)
| Name | Required | Description | Default |
|---|---|---|---|
| operation | Yes | Operation | |
| project_path | Yes | ||
| rule | No | ||
| category | No | ||
| reason | No | ||
| examples | No | ||
| importance | No | ||
| code_or_filename | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must convey behavioral traits. It mentions basic effects (store, retrieve, check, remove) but omits side effects, permissions, idempotency, error handling, or concurrency details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and uses a clear list format to separate operations. Every sentence contributes value; no redundant text. Slightly more structure (e.g., parameter roles) could improve readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 8 parameters, 4 operations, no annotations, and no output schema, the description is insufficient. It does not explain return values, error conditions, or default behavior for optional parameters, leaving the agent with significant unknowns.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 13%, so the description must compensate. It maps parameters to operations (e.g., rule, category for 'add'; rule for 'remove'), adding grouping information not in the schema. However, it still lacks detailed semantics for each parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as handling project conventions with four explicit operations (add, get, check, remove). Each operation is briefly described, and the tool is distinct from siblings like code_pattern_check or code_quality_check.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to use this tool versus alternatives. The description only lists operations without context on when each is appropriate or what prerequisites exist.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
deps_mapC
Map file dependencies. Shows imports and optionally reverse deps.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | ||
| include_reverse | No | ||
| project_root | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden but only states what it does vaguely. It does not disclose whether it modifies files, requires specific permissions, or handles missing files, which is insufficient for a dependency mapping tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, no redundancy, and front-loaded with the core purpose. Every sentence adds value without extra fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite clear intent, the description lacks specifics on output format, scope of dependency scanning, and the role of project_root. Without annotations or output schema, it leaves significant gaps for an agent to use effectively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description mentions 'imports' and 'reverse deps', which map to file_path and include_reverse, but does not explain project_root. With 0% schema coverage, this partially compensates but misses one parameter's meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool maps file dependencies and shows imports with optional reverse deps, using a specific verb and resource. However, it does not explicitly differentiate from sibling tools like code_pattern_check or impact_analyze.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives. Lacks context about prerequisites or scenarios where it is appropriate, leaving the agent to infer usage from the name and description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
file_summarizeC
Summarize file purpose. Modes: quick (pattern-based) or detailed (LLM).
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | ||
| mode | No | quick |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden of behavioral disclosure. It mentions two modes (quick/detailed) with brief explanations, but omits critical details: whether the tool is read-only, what happens if the file doesn't exist, how the output is structured, or any side effects. These gaps reduce transparency significantly.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise—only two sentences. The first sentence states the core purpose immediately, and the second adds crucial mode details. Every word earns its place with no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity (2 parameters, no output schema, no annotations), the description is incomplete. It lacks information about return values, error handling (e.g., file not found), and does not clarify if the tool modifies anything. An agent might need more context to use it reliably.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It explains the 'mode' parameter's enum values ('quick' is pattern-based, 'detailed' is LLM), adding meaning. However, it does not describe the required 'file_path' parameter beyond implying it's the file to summarize. Thus, partial compensation warrants a 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'summarize' and the resource 'file purpose', making the tool's function evident. It also distinguishes two modes, which adds specificity. However, it doesn't explicitly differentiate from sibling tools that might also deal with files, like 'convention' or 'scope', leaving slight ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. There is no mention of prerequisites, excluded scenarios, or related tools. The agent receives no context about appropriate usage, making it rely on inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_similar_issuesC
Search codebase for bug pattern (e.g., 'except:\s*pass').
| Name | Required | Description | Default |
|---|---|---|---|
| issue_pattern | Yes | Regex pattern | |
| project_path | Yes | ||
| file_extensions | No | ||
| exclude_paths | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It does not disclose behavioral traits such as read-only nature, potential performance impact, or whether results are limited to matches or include context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with a useful example. However, it could include more context without becoming verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 4 parameters with only 25% schema coverage and no output schema, the description should compensate. It falls short by not explaining return values, parameter usage, or behavioral expectations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only issue_pattern has a schema description ('Regex pattern'), and the tool description only provides an example pattern. No explanation for project_path, file_extensions, or exclude_paths, leaving their roles ambiguous.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description explicitly states the action ('Search codebase') and the resource ('bug pattern'), with a concrete example ('except:\s*pass'). This clearly distinguishes from siblings like code_pattern_check, which likely handles general patterns.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives like code_pattern_check or others. The description lacks context on prerequisites or scenarios where this tool is preferred.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
impact_analyzeC
Analyze change impact. Shows dependents, exports, risk level.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | ||
| project_root | Yes | ||
| proposed_changes | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, and the description only lists outputs but does not disclose side effects, mutability, authentication needs, rate limits, or output format.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Extremely concise at two sentences, but lacks crucial details. Still, no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (change impact analysis), the description omits output schema, interpretation of risk level, and limitations. Incomplete for practical use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, yet the description does not explain any parameter meanings or acceptable values for file_path, project_root, or proposed_changes.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool analyzes change impact and shows dependents, exports, and risk level. This distinguishes it from sibling tools like scout_analyze or deps_map.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No information on when to use this tool versus alternatives, or when not to use it. No prerequisites or context provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
loopA
Loop detection to prevent death spirals. Operations:
record_edit: Log file edit (file_path, description)
record_test: Log test result (passed, error_message)
check: Check if safe to edit (file_path)
status: Get edit counts and warnings
reset: Clear all loop tracking for a fresh start
| Name | Required | Description | Default |
|---|---|---|---|
| operation | Yes | Operation | |
| file_path | No | File being edited | |
| description | No | For record_edit: what changed | |
| passed | No | For record_test: did tests pass | |
| error_message | No | For record_test: error if failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden. It describes behaviors for each operation (log, check, clear) but lacks details on statefulness, side effects, or what 'death spirals' entails. Could be more explicit about mutability and persistence.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise: one sentence stating purpose followed by a clear bullet list of operations. No unnecessary words, and critical information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
While the operation list is helpful, the description omits expected return values or outputs for each operation (e.g., what does 'check' return? 'status' returns counts?). With no output schema, this gap is significant for an AI agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already describes all parameters with 100% coverage. The description adds operational context (e.g., which parameters apply to which operation) but does not significantly enhance understanding beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the purpose as 'Loop detection to prevent death spirals' and enumerates specific operations (record_edit, record_test, check, status, reset) with a brief explanation for each, making it distinct from sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage contexts via operations (e.g., record_edit when editing, check before editing), but does not explicitly state when to use this tool versus alternatives or provide any exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memoryC
Memory operations. Operations:
remember: Store a note (just content - category/relevance optional)
recall: Get all memories for project
forget: Clear project memories
search: Find by file/tags/query (file_path, tags, query, limit)
clusters: View grouped memories (cluster_id to expand)
cleanup: Dedupe/cluster/decay (dry_run, min_relevance, max_age_days)
consolidate: LLM-powered merge of related memories (tag, dry_run)
add_rule: Add permanent rule (content, reason) - never decays
list_rules: Get all rules for project
modify: Edit memory (memory_id, content, relevance, category)
delete: Remove single memory (memory_id)
batch_delete: Bulk delete by IDs (memory_ids) or by category. Rules/mistakes protected from category delete.
promote: Promote memory to rule (memory_id, reason)
recent: Get recent memories newest first (category, limit)
archive: Move old inactive memories to cold storage (dry_run to preview)
restore: Bring archived memory back to active (memory_id)
archive_search: Search archived memories (query, tags, limit)
archive_status: Show hot vs archived memory counts
hybrid_search: Semantic + keyword + scored search (query, file_path, tags, limit). Best retrieval.
embed_all: Generate AllMiniLM embeddings for all memories (enables hybrid_search)
| Name | Required | Description | Default |
|---|---|---|---|
| operation | Yes | Operation to perform | |
| project_path | Yes | Project directory | |
| content | No | For remember/add_rule/modify: content | |
| category | No | For remember/modify/batch_delete/recent: memory category | |
| relevance | No | For remember/modify: importance 1-10 | |
| file_path | No | For search: filter by file | |
| tags | No | For search: filter by tags | |
| query | No | For search: keyword search | |
| limit | No | For search/recent: max results | |
| cluster_id | No | For clusters: expand specific cluster | |
| tag | No | For consolidate: only consolidate memories with this tag | |
| dry_run | No | For cleanup/consolidate: preview only | |
| min_relevance | No | For cleanup: min to keep | |
| max_age_days | No | For cleanup: decay threshold | |
| memory_id | No | For modify/delete/promote: memory ID | |
| memory_ids | No | For batch_delete: list of memory IDs to delete | |
| reason | No | For add_rule/promote: why this rule |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It discloses destructive operations (forget, delete, etc.) and explains protections (rules/mistakes from category delete). However, it does not address authentication needs, rate limits, or failure behavior, leaving gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a long, unstructured list that is not concise. It could be organized into categories or groups. While front-loaded with 'Memory operations', it still reads as a wall of text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (20 operations, 17 parameters), the description is fairly complete in covering each operation's behavior. However, it lacks information about return values (no output schema) and error conditions, making it less than fully comprehensive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. The description adds minor context by associating parameters with specific operations, but largely repeats what the schema already states. It does not significantly deepen understanding beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly enumerates each operation with a brief explanation, making it specific about what the tool does across multiple memory management tasks. However, the lack of a title and the overwhelming list slightly reduce clarity. It is distinguishable from siblings by being a general memory tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus other sibling tools (e.g., context, convention). The description only lists operations without providing context on when each operation is appropriate or when to choose an alternative tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
outputC
Output validation. Operations:
validate_code: Check for fake/silent failures (code, context)
validate_result: Check output for fakes (output, expected_format, should_contain, should_not_contain)
| Name | Required | Description | Default |
|---|---|---|---|
| operation | Yes | Operation | |
| code | No | ||
| context | No | ||
| output | No | ||
| expected_format | No | ||
| should_contain | No | ||
| should_not_contain | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, and the description fails to disclose behavioral traits such as side effects, authorization needs, or return values. It is unclear whether validation failures raise errors or return results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and uses a clear bullet-point structure. It front-loads the purpose and efficiently lists operations. Minor improvement would be to add a return value note.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 7 parameters, no output schema, and no annotations, the description is insufficient. It does not explain return values, error handling, or full parameter details, leaving significant gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds meaning for parameters used in each operation (e.g., validate_code uses code and context), but many parameters remain unexplained (e.g., context, expected_format syntax). With only 14% schema description coverage, the description partially compensates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it performs output validation and lists specific operations with brief explanations. However, it does not differentiate from sibling tools like code_quality_check or code_pattern_check, which may overlap.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives, nor on choosing between validate_code and validate_result. The agent must infer usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pre_edit_checkB
Run BEFORE editing important files. Checks: past mistakes, loop risk, scope violations.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | File about to edit |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided; description does not disclose side effects, authorization needs, rate limits, or return behavior. Only lists what it checks, which is minimal transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero wasted words, key information front-loaded. Highly concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Missing output schema and description does not specify return value or format. For a check tool, agents need to know what to expect (e.g., boolean, list, or summary) for decision-making.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with a clear description for 'file_path.' The tool description adds no additional meaning beyond what the schema already provides, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states it is a pre-edit check tool that examines past mistakes, loop risk, and scope violations, distinguishing it from sibling tools like code_pattern_check or code_quality_check.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Run BEFORE editing important files,' providing clear timing for use. Does not specify exclusions or alternatives, but context is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scopeA
Scope guard for multi-file tasks. Operations:
declare: Set task scope (task_description, in_scope_files, in_scope_patterns)
check: Verify file is in scope (file_path)
expand: Add files to scope (files_to_add, reason)
status: Get violations
clear: Reset scope
| Name | Required | Description | Default |
|---|---|---|---|
| operation | Yes | Operation | |
| task_description | No | For declare: task being done | |
| in_scope_files | No | For declare: allowed files | |
| in_scope_patterns | No | For declare: glob patterns | |
| file_path | No | For check: file to verify | |
| files_to_add | No | For expand: files to add | |
| reason | No | For expand: why adding |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It explains the operations and their parameter associations but does not disclose behavioral details such as side effects (e.g., does 'clear' destroy data?), authorization needs, or concurrency behavior. The operations are transparent in what they do but not their full impact.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise: one introductory sentence and a bulleted list of operations. Every sentence is necessary and front-loaded with the core purpose. No redundant or verbose language.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (multiple operations with many parameters) and lack of output schema, the description adequately covers operations and parameter usage. However, it omits details about return values (e.g., what does 'status' return?), scope persistence, or how this interacts with sibling tools like 'session_start' or 'context'.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and each parameter already has a description indicating which operation it applies to. The description merely restates this mapping (e.g., 'For declare: task being done'), adding no new semantic meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool is a 'Scope guard for multi-file tasks' and lists five specific operations (declare, check, expand, status, clear), making the purpose unambiguous. It distinguishes itself from sibling tools like 'context' or 'loop' by focusing on file scoping.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for 'multi-file tasks' but provides no explicit guidance on when to use this tool versus alternatives among siblings (e.g., 'context', 'pre_edit_check'). No when-not-to-use or exclusion criteria are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scout_analyzeC
Analyze code with local LLM. Provide code and question.
| Name | Required | Description | Default |
|---|---|---|---|
| code | Yes | ||
| question | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses that analysis uses a local LLM, which is a key behavioral trait (privacy, dependency). However, it does not mention potential side effects, limitations, or authorization requirements. With no annotations, this is minimally adequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no wasted words. However, it could be more structured (e.g., bullet points) to improve readability.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema or annotations provided. The description does not mention return values, error handling, or usage context. For a simple tool, it leaves significant gaps in understanding what the tool produces.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description restates parameter names ('code and question') without adding meaning beyond the schema. With schema description coverage at 0%, the description fails to compensate by explaining constraints, formats, or examples.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'analyze' and resource 'code with local LLM'. It is specific but does not differentiate from siblings like code_pattern_check or code_quality_check, which also analyze code.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives. The description only instructs to provide code and question, lacking context for selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scout_searchB
Search codebase semantically. Returns findings with files, lines, connections.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | What to search for | |
| directory | Yes | Directory to search | |
| max_results | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It does not disclose whether the tool is read-only, any authentication needs, or performance implications. Only hints at output structure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Extremely concise: two short sentences that convey action and output. No wasted words, front-loaded with the core purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with 3 parameters and no output schema, the description covers the basic output shape but lacks explanation of 'connections' and 'semantic' search. Adequate but not fully informative.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67% with basic parameter descriptions. The description adds no additional meaning beyond what the schema provides. Baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states 'Search codebase semantically' with specific verb and resource. Also describes output as 'findings with files, lines, connections', making the tool's purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives like code_pattern_check or find_similar_issues. Usage context is merely implied by the description.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_endA
Optional. Shows session summary. All memories auto-save without this - just a nice recap.
| Name | Required | Description | Default |
|---|---|---|---|
| project_path | No | Project directory (optional) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description bears full responsibility. It transparently states the tool is optional, shows a summary, and doesn't affect memory saving. No negative traits like destruction or auth needs are mentioned, which is appropriate for a harmless tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise with one sentence that front-loads the key purpose ('Optional. Shows session summary.') and adds clarifying context. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple optional tool, the description covers the essential purpose and clarifies it's a recap. No output schema exists, but the description implies the output is a summary, which is sufficient. Lacks mention of the parameter, but given its optionality and schema coverage, it's adequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is one optional parameter 'project_path' with 100% schema description coverage. The tool description does not add any extra meaning beyond what the schema already provides, so baseline score of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it 'shows session summary', specifying the verb and resource. It distinguishes itself from sibling tools like 'session_start' and 'session_mine' by indicating it's a recap, not a start or mine.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description marks it as 'Optional' and clarifies that memories auto-save without it, implying it's for a recap only. While it doesn't explicitly state when to use or alternatives, the context is clear that it's not required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_mineA
Mine session history. Operations:
search: Search across past conversations (query, project_path, limit, method=hybrid|semantic|keyword)
decisions: Find when/why a decision was made (query, project_path)
replay: Find discussions about a file (file_path, project_path)
struggles: Files/areas with repeated difficulty (project_path)
errors: Recurring error patterns across sessions (project_path)
correlations: Files always edited together (project_path)
timeline: Project development timeline (project_path)
summaries: Auto-generated session summaries (project_path)
overview: High-level project stats (project_path)
status: Mining index coverage (project_path)
reindex: Trigger background re-indexing (project_path, mode=post_session|bootstrap|full)
predict: Predict context needed for a file edit (file_path, project_path)
cross_project: Patterns across all projects (no project_path needed)
reflect: LLM-powered analysis of mistakes, patterns, and decisions (project_path)
| Name | Required | Description | Default |
|---|---|---|---|
| operation | Yes | Operation to perform | |
| project_path | No | Project directory | |
| query | No | For search/decisions: search query | |
| file_path | No | For replay: file to find discussions about | |
| limit | No | Max results (default 10) | |
| method | No | For search: search method | |
| mode | No | For reindex: mining mode | |
| since | No | For search: filter after date (YYYY-MM-DD) | |
| until | No | For search: filter before date (YYYY-MM-DD) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description bears full responsibility. It describes what each operation does but fails to disclose behavioral traits such as side effects (e.g., reindex modifies state), authorization needs, or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is structured as a bulleted list within a paragraph, making it scannable. It front-loads the main purpose, but the list is lengthy (14 items) and could be more concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (14 operations, 9 parameters, no output schema), the description covers every operation and its associated parameters comprehensively, providing a complete picture of the tool's capabilities.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema coverage, the description adds value by grouping parameters per operation (e.g., 'for search/decisions: query') and clarifying which parameters apply to which operation, going beyond the schema's flat descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Mine session history' and enumerates 14 distinct operations with specific verbs (search, decisions, replay, etc.), correctly distinguishing from sibling tools like scout_search or memory.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description lists operations but does not provide explicit guidance on when to use this tool versus alternatives. Usage is implied through operation names, but no when-not-to or comparison to siblings is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_startA
Load full context: memories, checkpoints, decisions, memory health. Auto-cleans duplicates. Hook auto-starts basic session, but this gives deep context.
| Name | Required | Description | Default |
|---|---|---|---|
| project_path | Yes | Project directory path |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must fully disclose behavioral traits. It mentions 'auto-cleans duplicates' which implies mutation, but does not clarify whether the tool is read-only or modifies state. No disclosure of auth needs, rate limits, or side effects beyond cleanup.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core purpose, and uses clear language. It is efficient but could be more structured for skimming.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has no output schema, so the description should explain return values or side effects. It mentions auto-cleaning duplicates but does not describe what the tool returns or any prerequisites. Incomplete for a tool that loads context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for the single parameter 'project_path'. The description does not add any meaning beyond the schema's 'Project directory path' explanation. Baseline score of 3 applies as no extra value is provided.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool loads full context including memories, checkpoints, decisions, and memory health. It differentiates from a basic session start by noting the hook auto-starts basic session but this gives deep context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use this tool (when deep context is needed) versus the hook (basic session). It does not explicitly list when not to use or mention alternatives like session_mine, but provides sufficient guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workB
Work tracking. Operations:
log_mistake: Record error (description, file_path, how_to_avoid)
log_decision: Record choice (decision, reason, alternatives)
| Name | Required | Description | Default |
|---|---|---|---|
| operation | Yes | Operation | |
| description | No | For log_mistake: what went wrong | |
| file_path | No | For log_mistake: affected file | |
| how_to_avoid | No | For log_mistake: prevention | |
| decision | No | For log_decision: what was decided | |
| reason | No | For log_decision: why | |
| alternatives | No | For log_decision: other options |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description must bear the full burden. It states it 'records' data but provides no details on side effects, persistence, idempotency, or authentication requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise, using a bullet list format with no extraneous words. Every sentence serves a purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (7 parameters, no output schema, no annotations), the description provides an overview and parameter grouping but lacks information about return values, confirmation, or persistence behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 100% coverage with per-parameter descriptions. The description adds value by grouping parameters under each operation (log_mistake: description, file_path, how_to_avoid; log_decision: decision, reason, alternatives), clarifying which parameters belong to which operation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool is for 'Work tracking' and lists two operations (log_mistake, log_decision) with their purposes. This differentiates it from sibling tools, which do not mention logging mistakes or decisions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives or when not to use it. It only lists the operations without context on selection criteria.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
21 tool updates
v0.2.0- First observed
audit_batch - First observed
claude_engram_status - First observed
code_pattern_check - First observed
code_quality_check - First observed
context - First observed
convention - First observed
deps_map - First observed
file_summarize - First observed
find_similar_issues - First observed
impact_analyze - First observed
loop - First observed
memory - First observed
output - First observed
pre_edit_check - First observed
scope - First observed
scout_analyze - First observed
scout_search - First observed
session_end - First observed
session_mine - First observed
session_start - First observed
work
TDQS
Scored across 21 tools
Most tools have distinct purposes, but there is some overlap between code_pattern_check, code_quality_check, and convention.check, which all analyze code against rules. Additionally, scout_search and find_similar_issues both search code for patterns, though they target different use cases. Descriptions help differentiate, but an agent might occasionally misselect.
Tool names consistently use underscore_case with a verb_noun or noun_verb pattern (e.g., audit_batch, session_start, code_quality_check). However, a few names like claude_engram_status and find_similar_issues break the pattern slightly, and the nested operation prefixes (e.g., 'checkpoint_save' under 'context') could be more uniform.
With 21 tools, the count is on the higher side but still reasonable for a comprehensive developer assistant server. Each tool serves a distinct purpose, and the inclusion of nested operations (e.g., under 'memory' and 'session_mine') keeps the top-level list manageable. A slight reduction could improve navigability.
The tool set covers a wide range of features: memory management, session/handoff handling, code analysis, conventions, scope, output validation, and work tracking. It lacks direct file I/O tools, but that is likely handled by other servers. Overall, the surface is comprehensive for a context and memory management server, with only minor gaps like a dedicated planning or task decomposition tool.
Maintenance
Related MCP Connectors
Shared memory for coding agents. Stop re-explaining your codebase every session.
Persistent cross-session memory shared by Codex, Claude Code, ChatGPT, and other AI agents.
Persistent memory and cross-session learning for AI coding assistants (hosted remote MCP).
- EngramOAuthtools.engram
Memory for AI agent teams across tools, sessions, repositories, and teammates.
Related MCP Servers
AlicenseNot gradedqualityCmaintenanceProvides persistent memory for AI coding assistants, storing and retrieving architectural decisions, patterns, and solutions across sessions using semantic search, while also offering git integration for commit messages and code expertise mapping.MIT- AlicenseNot gradedqualityDmaintenanceNever start from zero. Persistent session intelligence for AI coding assistants.53 npm1MIT
- AlicenseAqualityDmaintenanceLong-term memory for AI coding assistants. Remembers context once and recalls it across sessions.722MIT
- AlicenseNot gradedqualityBmaintenancePersistent memory for AI coding tools that captures conversations, builds a searchable knowledge graph, and automatically injects relevant context into new prompts.6 npm244MIT