total-agent-memory
TAM gives AI coding agents persistent, local, searchable memory across sessions over MCP — storing knowledge and retrieving it with hybrid search, a knowledge graph and temporal reasoning.
Save and recall knowledge: store decisions, facts, solutions, lessons and conventions;
memory_recallsearches everything via a 6-stage pipeline (FTS5/BM25 → semantic embeddings → fuzzy → graph → optional CrossEncoder rerank → MMR) fused with reciprocal rank fusion.Tune retrieval: filter by project, branch, type, intent, topics, entities; choose detail levels (compact/summary/full/index), progressive-disclosure modes (search, index, timeline, context, evidence), graph expansion, contamination checks and citation-backed answers (
memory_answer).Manage records: update/supersede, soft or hard delete, version history (
memory_history), batched fetch by ID, tag search, dedup consolidation, export to JSON, stats and retention policy (memory_forget).Browse sessions: timeline views, session start/end summaries with highlights, pitfalls and next steps, plus extraction of pending session transcripts.
Knowledge graph and temporal facts: typed relations, entity resolution, concept search, graph traversal/stats, and time-aware facts (
kg_add_fact,kg_invalidate_fact,kg_at,kg_timeline,memory_temporal_query) with Allen interval and duration reasoning.Self-improvement loops: log errors, derive insights, promote behavioral rules, track patterns/effectiveness, save reflections and episodes, manage skills, and build optimal context or run reflection passes.
Workflows and code: learn/predict/track reusable workflows, ingest codebases into AST chunks, get file risk context before editing, and find analog solutions from other projects.
Reporting: deterministic day/week/month/all-time activity reports (decisions, fixes, errors, files, entities) and per-project Markdown wiki generation.
Operations and performance: fast-path save/search, search explainability, warmup, perf counters, FTS/embedding rebuilds, and built-in eval/benchmark harnesses (recall@k, LoCoMo/LongMemEval-style, temporal, contradiction, long-context tests).
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@total-agent-memoryRemember my preferred style for error handling in Python"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Persistent, local memory for AI coding agents: Claude Code, Codex CLI, Cursor and any MCP client.
total-agent-memory (TAM) is an open-source memory server for AI coding agents. Coding agents start every session without memory of earlier ones, so decisions, fixes and project conventions have to be explained again. TAM stores decisions, solutions, facts, errors and session summaries on your machine and returns them through the Model Context Protocol (MCP), so any MCP client can use it without code changes.
Each store is one directory built around a SQLite database. Recall combines
full-text BM25, dense embeddings computed locally, fuzzy matching and a
knowledge graph, and fuses the ranked lists with reciprocal rank fusion; an
optional cross-encoder can rerank the result. The default profile makes no LLM
call on write or search, so retrieval can be measured offline and a rerun gives
the same result. Facts can carry validity intervals (kg_add_fact, kg_at),
and a newer value of a single-valued fact can retire the older one.
TAM is for developers who use coding agents daily and for researchers who need a memory baseline they can run locally, inspect and change. The same package can also run as a team server: personal, department and company areas are separate stores behind one gateway, with roles, authorship and an audit trail (team server).
Installation
TAM needs Python 3.11 or newer. CI tests Python 3.11, 3.12 and 3.13 on Ubuntu, Windows and macOS.
From PyPI (use a virtual environment, or pipx / uvx):
pip install total-agent-memory # or: pipx install total-agent-memory
uvx total-agent-memory # run once without installingThis installs the total-agent-memory (alias tam) MCP server and the
tam-team, tam-remote and lookup-memory commands. The optional
cross-encoder reranker pulls in PyTorch and is an extra:
pip install "total-agent-memory[rerank]". The PostgreSQL backend of the team
server is the [postgres] extra.
Register it with your clients. Run tam setup at a terminal. The wizard
asks "Just me" or "Company server", detects Claude Code, Claude Desktop, Codex,
Cursor, Windsurf, Gemini CLI, Cline and OpenCode, and registers the server with
the ones you pick (setup wizard).
Other channels:
Channel | Command |
Docker (linux/amd64, linux/arm64) |
|
npx connector |
|
Claude Code plugin (server, skill and capture hooks) |
|
Plugin for Claude Code and Cowork, from the plugin repository (server and skill, no hooks; needs uv) |
|
The same plugin for Codex CLI |
|
Source checkout with IDE hooks and background services |
|
Use one of the two Claude Code plugins, not both. Per-platform details, WSL2, the IDE matrix, uninstalling and troubleshooting are in the installation guide.
Related MCP server: SharedMemory MCP Server
Quick start
Install the package and run
tam setup, or add the server to your client by hand. For Claude Code:claude mcp add memory -- total-agent-memoryFor clients that use an
mcpServersfile:{ "mcpServers": { "memory": { "command": "total-agent-memory" } } }If you installed into a virtual environment, use the full path to its
total-agent-memoryexecutable.Restart the client. Memory is stored in
~/.tam/unlessTAM_MEMORY_DIRpoints elsewhere.Ask the agent to remember something ("remember that we chose PostgreSQL for billing because of row-level security"). It calls
memory_save. In a later session, ask "which database did we choose for billing?"; it callsmemory_recalland gets the record back.
To try the server without an agent, this script starts it over stdio with the
MCP Python SDK (installed as a dependency), saves one record and recalls it.
The SDK passes only a few variables to the server by default, so the script
hands over the full environment, including TAM_MEMORY_DIR:
import asyncio
import os
from mcp import ClientSession, StdioServerParameters
from mcp.client.stdio import stdio_client
async def main() -> None:
server = StdioServerParameters(command="total-agent-memory", env=dict(os.environ))
async with stdio_client(server) as (read, write):
async with ClientSession(read, write) as session:
await session.initialize()
await session.call_tool("memory_save", {
"type": "decision",
"content": "Chose PostgreSQL over MySQL for the billing service",
"context": "WHY: row-level security per tenant",
"project": "demo",
})
result = await session.call_tool("memory_recall", {
"query": "which database for billing", "project": "demo", "limit": 3,
})
print(result.content[0].text)
asyncio.run(main())TAM_MEMORY_DIR="$(mktemp -d)" python quickstart.pyThe first run downloads the default embedding model. The output
is JSON with the saved decision under results.decision. More examples:
tools and interfaces.
Running the tests
The tests run from a source checkout. These commands mirror the CI workflows in .github/workflows:
git clone https://github.com/vbcherepanov/total-agent-memory.git
cd total-agent-memory
python3 -m venv .venv
.venv/bin/pip install -e . -r requirements-dev.txt
export FASTEMBED_CACHE_PATH="$PWD/.tam-models" TAM_MEMORY_DIR="$PWD/.tam-test-memory"
.venv/bin/python tests/smoke/prewarm_models.py # downloads the text and code embedding models (~850 MB) once
.venv/bin/python -m pytest tests -qThe tests need no API keys and no LLM; the full suite takes 15 to 20 minutes
on a laptop. Do not set MEMORY_LLM_ENABLED=false for the full suite: the
configuration tests check LLM auto-detection. Tests that need services you do not
have are skipped:
PostgreSQL (team server backend): install the extra with
.venv/bin/pip install -e ".[postgres]", then run.venv/bin/python -m pytest tests -m postgres --backend=postgresor--backend=both. WithoutTAM_TEST_PG_URLthe fixtures start a pgvector container through Docker. See CONTRIBUTING.md and .github/workflows/postgres.yml.Browser tests (
tests/browser, Playwright with Chromium, Firefox and WebKit) run in the image built from docker/Dockerfile.browser; see thebrowserjob in .github/workflows/smoke.yml.Installed-package smoke test:
python tests/smoke/installed_runtime.pychecks an installed wheel over local and remote MCP.
Reproducing the benchmarks
The retrieval benchmarks (LoCoMo, LongMemEval, BEAM) need no API key and run with scripts in benchmarks/ once the public datasets are downloaded. Results with the default profile, v13.0.0:
Benchmark | Metric | Result |
LongMemEval (470 questions) | R@5 (recall_any) | 95.1% |
LoCoMo (1,536 questions) | R@5 | 0.607 |
BEAM, 1M-token scale (625 probes) | R@5 | 0.448 |
Dataset locations, commands, end-to-end (LLM-judged) accuracy, negative controls, latency and the 14.x studies are in docs/benchmarks.md.
Documentation
Installation guide: all channels, per-platform setup, WSL2, IDE matrix, troubleshooting
Setup wizard:
tam setupand the team server's web wizardTools and interfaces: the 77 MCP tools, CLI, TypeScript SDK, dashboard
Configuration: environment variables, LLM providers, performance tuning
Team server quick start, with dashboard, PostgreSQL, backup, onboarding and reports
Benchmarks and comparison with other systems (April 2026 snapshot)
Citation
A paper describing TAM is under review at the Journal of Open Source Software (paper/paper.md). Until it is published, please cite the software and the preprint:
Cherepanov, V. total-agent-memory (version 14.7.0) [software]. https://github.com/vbcherepanov/total-agent-memory
Preprint: doi:10.5281/zenodo.23011523
Citation metadata for the software is in CITATION.cff.
Contributing
Issues, pull requests and benchmark reproductions are welcome. See CONTRIBUTING.md for the development setup, the rules for a pull request and the commit convention. Report security issues privately as described in SECURITY.md. Donations to support development: PayPal.
License
MIT. See LICENSE. Third-party licenses are listed in THIRD-PARTY-LICENSES.md.
Available Tools
77 toolsanalogizeBRead-onlyIdempotent
Find past solutions/lessons from OTHER projects whose feature set overlaps with the given problem text (Jaccard similarity).
| Name | Required | Description | Default |
|---|---|---|---|
| text | Yes | ||
| limit | No | ||
| min_score | No | ||
| only_types | No | ||
| exclude_project | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, and non-destructive behavior. The description adds useful context about the algorithm (Jaccard similarity) and the scope (other projects), which goes beyond the annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, clear, and well-structured sentence. It efficiently communicates the core purpose without unnecessary words or complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description gives a high-level purpose but lacks operational details. It does not explain the semantics of the parameters, what the output looks like, or any edge cases. An agent would likely be uncertain about how to properly configure the input beyond the single required 'text' parameter.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 5 parameters but zero description coverage. The description only mentions 'given problem text' and 'Jaccard similarity,' leaving the meaning and usage of limit, min_score, only_types, and exclude_project completely unexplained. The description fails to compensate for the lack of schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: finding past solutions/lessons from other projects using feature-set overlap and Jaccard similarity. It is specific and distinguishes this tool from sibling memory tools that focus on the current project's memory.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not provide any guidance on when to use this tool versus alternatives, nor does it mention any conditions or prerequisites. It lacks explicit direction for an agent to decide when analogize is the appropriate choice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
benchmarkARead-onlyIdempotent
Run the eval harness: recall_at_k, prevention_rate, latency percentiles. Loads scenarios from evals/scenarios/*.json by default.
| Name | Required | Description | Default |
|---|---|---|---|
| scenarios_path | No | Custom scenarios dir or file |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already establish read-only, idempotent, and non-destructive behavior, and the description does not contradict them. It adds useful context about loading scenarios and computing specific metrics, though it does not explicitly describe output or side effects beyond what the annotations imply.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences with no redundant or filler content. It packs the core action, metrics, and default loading behavior efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simple schema and clear annotations, the description is adequate for an agent to understand what the tool does and what inputs it accepts. It could be slightly more explicit about return values or output format, but no output schema exists and the intended evaluation metrics are named.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter has full schema coverage with a clear description ('Custom scenarios dir or file'). The tool description adds the useful default behavior (evals/scenarios/*.json), which clarifies how the parameter relates to the default when omitted.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb ('Run') and resource ('eval harness'), and itemizes the metrics it produces. It distinguishes this tool from the memory- and workflow-related siblings by focusing on evaluation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use for running evaluations and notes the default scenario path, but it does not explicitly state when to use this tool versus any alternative or provide conditions for when it should be avoided. There is no direct competing eval tool among siblings, but guidance is still minimal.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
classify_taskCRead-onlyIdempotent
v8.0: classify task into L1-L4 complexity + suggested phases.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | ||
| description | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate read-only, idempotent, non-destructive behavior, so the description does not need to restate these. It adds no additional side-effect information, but also does not contradict the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise and contains no irrelevant information. Every word adds meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, no parameter explanations, and no usage context. The description is too sparse to allow a user to confidently invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema provides no parameter descriptions and the tool description does not explain the meaning or expected format of 'project' or 'description'. Essential parameter semantics are entirely absent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the action (classify), the object (task), and the output (L1-L4 complexity plus suggested phases). However, it does not explicitly distinguish itself from sibling tools that might also handle task-related operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives, nor any prerequisites or context for invoking it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
file_contextARead-onlyIdempotent
BEFORE editing a file, call this to surface past errors, lessons, and related rules for that file path. Returns risk_score ∈ [0, 1].
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| limit | No | ||
| project | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate read-only, idempotent, non-destructive behavior. The description adds useful behavioral detail about surfacing past errors, lessons, related rules, and returning a risk score, with no contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One compact sentence conveys the action, timing, content returned, and risk score. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description partially compensates by mentioning risk_score and the categories of returned information, but it does not describe the full return structure. Adequate for basic invocation, but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has no field descriptions, and the description only clarifies 'path'. The 'limit' and 'project' parameters are left unexplained, so the description provides only partial compensation for the missing schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states a specific action and resource: call to surface past errors, lessons, and related rules for a file path, and returns a risk score. This clearly distinguishes it from generic memory recall tools and ties it to a pre-edit workflow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'BEFORE editing a file', giving a clear trigger for use. It does not name alternative tools or say when not to use it, but the primary use case is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ingest_codebaseB
Parse a file or directory into semantic AST chunks (functions, classes, methods) across 8 languages. Returns chunk count + sample.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | ||
| include | No | Extension allowlist e.g. ['.py','.go'] | |
| sample_limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate readOnlyHint=false and destructiveHint=false, but the description does not disclose any side effects (e.g., whether it writes to a database or modifies files). It only mentions parsing and returning data, leaving behavioral traits ambiguous and not adding clarity beyond the minimal annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, using two clear sentences without redundant information. It efficiently conveys the core action and output, maintaining a clean structure that is easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's simplicity and lack of an output schema, the description provides a basic understanding of the return (chunk count + sample) but omits details like the exact output format or error conditions. This is a gap for an agent that needs to interpret results reliably, so completeness is only moderate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has three parameters (path, include, sample_limit), but only 'include' has a description. The description does not elaborate on any parameter meanings, and with schema coverage at only 33% (low), the description fails to compensate, leaving the required 'path' and 'sample_limit' under-specified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: parsing files or directories into semantic AST chunks across 8 languages, and it explicitly mentions what it returns (chunk count + sample). This distinguishes it from sibling tools, which are memory-related, making the purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for code parsing but does not explicitly state when to use this tool versus alternatives. Since all siblings are memory/workflow tools, the context makes usage obvious, but no explicit guidance on when not to use it is provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kg_add_factB
Record a temporal fact assertion (subject, predicate, object). Supersedes any prior assertion with same (s,p) and different object — full history is preserved. Use for evolving architectural decisions.
| Name | Required | Description | Default |
|---|---|---|---|
| object | Yes | ||
| context | No | ||
| project | No | general | |
| subject | Yes | ||
| predicate | Yes | ||
| confidence | No | ||
| valid_from | No | ISO 8601 time the fact became true; default now. Back-dated facts close and are closed by their neighbours. | |
| invalidate_previous | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare safe-mutation flags, so the description carries real weight here: it discloses that a new assertion supersedes prior ones with the same (s,p) and different object, and that full history is retained. That is meaningful behavioral context beyond the structured fields and is consistent with destructiveHint=false, since history is preserved rather than erased.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences that front-load the core operation and then add the supersede rule and a usage cue. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An 8-parameter mutation tool with no output schema and near-zero schema coverage needs more than two sentences. The roles of confidence, context, project, and especially invalidate_previous (which governs the supersede behavior the description mentions) are never explained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 13% (just valid_from), leaving context, project, confidence, and invalidate_previous undocumented in both schema and description. The description clarifies the (s,p,o) semantics but does nothing for the remaining five parameters, so it falls well short of compensating for the coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Record') and resource ('temporal fact assertion') with the triple structure spelled out as (subject, predicate, object). It doesn't explicitly differentiate itself from the close sibling kg_invalidate_fact, but the temporal-assertion framing is clear enough to identify the operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use for evolving architectural decisions' gives a concrete usage scenario, and the supersede rule implies when re-asserting is appropriate. However, with siblings like kg_invalidate_fact and kg_at available, no alternative is named and no when-not guidance is given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kg_atBRead-onlyIdempotent
Point-in-time query: return fact assertions valid at timestamp (ISO 8601). Omit timestamp for currently-valid facts.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| object | No | ||
| project | No | ||
| subject | No | ||
| predicate | No | ||
| timestamp | No | ISO 8601 or omit for now |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and idempotentHint, so safety is covered. The description adds the key behavioral nuance of timestamp filtering, which is useful. However, it does not describe the return format, potential limits, or error conditions. Given that annotations cover the safety aspects, this is adequate but not exceptional.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no redundant words. It directly states the purpose and the optional behavior of the timestamp parameter. Every word serves a purpose, making it highly concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a query tool in a knowledge graph context, the description sufficiently explains how to use it: specify a timestamp for point-in-time valid facts, or omit for current ones. It distinguishes itself from timeline tools by focusing on a single time point. The unmentioned parameters are standard in such domains (subject, predicate, object) and likely understandable. A small deduction for not specifying the output shape, but overall it is complete enough for an agent to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only the timestamp parameter is described in the schema and referenced in the description. The other five parameters (limit, object, project, subject, predicate) have no descriptions and are not mentioned. With schema coverage at only 17%, the description fails to compensate for the missing parameter documentation, leaving most of the parameters underspecified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool returns fact assertions valid at a specific timestamp, or current facts if timestamp is omitted. This makes the core purpose unambiguous without needing to inspect the schema. However, it does not explicitly distinguish itself from sibling tools like kg_timeline, so it loses a point.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a clear condition for usage: include timestamp for point-in-time queries, omit for current facts. This gives practical guidance, but it does not mention when to prefer this over other query tools or mention any limitations (e.g., pagination, ordering). The guidance is implicit rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kg_invalidate_factBDestructive
Close a currently-valid fact assertion. History is retained.
| Name | Required | Description | Default |
|---|---|---|---|
| at | No | ISO 8601 time the fact stopped being true; default now | |
| object | Yes | ||
| reason | No | manually_invalidated | |
| project | No | general | |
| subject | Yes | ||
| predicate | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare destructiveHint=true with no read-only/idempotency guarantees, and the description usefully clarifies the nature of that destruction with 'History is retained' – a genuinely additive behavioral fact. It still omits whether closing an already-closed fact is a no-op and any permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short, front-loaded sentences that each carry information with no waste. Efficient, though extremely terse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter mutation tool with 17% schema coverage and no output schema, the description is too thin. It never explains reason/project semantics or the temporal 'at' behavior, leaving an agent under-informed for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 17% (just 'at'). The description supplies no meaning for subject, predicate, object, reason, or project, and does not compensate for the undocumented parameters. Only the bare concept of a 'fact assertion' is implied.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: 'Close a currently-valid fact assertion.' An agent can tell this is a write operation on an existing fact. However, it doesn't explicitly contrast with the obvious sibling kg_add_fact, so it isn't a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'currently-valid' qualifier hints that only valid facts should be closed, but there is no explicit when-to-use, when-not-to-use, or mention of alternatives like kg_add_fact. Guidance is left entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
kg_timelineBRead-onlyIdempotent
Full chronological history of assertions for a subject.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| project | No | ||
| subject | Yes | ||
| predicate | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds that results are full and chronological, but it does not disclose the default limit, whether invalidated assertions are included, or how predicate/project filters affect results. Read-only and idempotent behavior are covered by annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single concise sentence with no redundant or extraneous content; the core action and resource are front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema, the description conveys that a chronological history is returned, but omits filtering semantics and default limits, leaving some operational context implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With zero parameter descriptions in the schema, the text only adds meaning to 'subject' (the entity whose assertions are returned). limit, predicate, and project remain unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies the resource (chronological assertions for a subject) and implies retrieval via 'full history', distinguishing it from a point-in-time lookup like kg_at. It lacks an explicit verb but is clear in context.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given on when to prefer this tool over siblings such as kg_at or kg_add_fact, nor any mention of use cases or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
learn_errorC
Structured error capture: file, error, root_cause, fix, pattern. After N (default 3) errors share the same pattern, a prevention rule is auto-synthesized into the rules table.
| Name | Required | Description | Default |
|---|---|---|---|
| fix | Yes | ||
| file | Yes | ||
| error | Yes | ||
| pattern | Yes | ||
| project | No | general | |
| category | No | bug | |
| severity | No | medium | |
| root_cause | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses the key side effect of auto-synthesizing a prevention rule into the rules table, which aligns with the readOnlyHint=false annotation. However, it does not clarify whether repeated calls with the same pattern create duplicate rules or are deduplicated, and it omits any mention of idempotency or potential side effects on existing data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, concise sentence that efficiently communicates the core functionality and the auto-synthesis behavior. It avoids unnecessary detail, though it could be slightly more structured by separating the parameter list from the side-effect explanation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (8 parameters, side effects, pattern detection), the description provides insufficient context. It does not explain how the pattern is determined, what the rules table looks like, or the relationship with sibling tools like self_error_log or workflow_learn. The absence of an output schema and lack of examples further reduces completeness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description lists the required parameters (file, error, root_cause, fix, pattern) but does not explain their meaning or relationships, and the schema has no descriptions. Optional parameters like project, category, and severity are not mentioned at all, leaving their semantics unclear. The parameter names are self-explanatory but not fully defined, especially 'pattern' and how it differs from 'error'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states 'Structured error capture' and lists the key parameters, making the tool's purpose apparent. However, it does not explicitly name the verb 'record' or 'log', relying on the phrase 'error capture' to imply the action, which is a minor gap.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description mentions that after N errors with the same pattern, a prevention rule is auto-synthesized, hinting at when the tool's output becomes useful. But it does not explicitly state when to use this tool versus alternatives (e.g., self_error_log) or when not to use it, leaving usage guidance vague.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_intentsARead-onlyIdempotent
List recent user prompts from the intents table, newest first. Filter by project and/or session. Max 500 rows.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| project | No | ||
| session_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering safety. The description adds useful behavioral details: ordering (newest first) and a row cap (Max 500 rows). It also clarifies the data source (intents table). No contradictions with annotations, and the added traits are relevant to invocation expectations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, compact sentence that front-loads the primary purpose (List recent user prompts) and then efficiently adds ordering, filters, and a limit. Every word contributes value; there is no fluff or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple list tool with no output schema and annotations covering safety, the description is quite complete. It states the action, data source, sorting, filters, and a row limit. It does not describe the return format (e.g., fields of each prompt), but that is not critical for a straightforward read-only list. The absence of pagination details is mitigated by the limit parameter. Overall, an agent can call it correctly with the given information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With schema description coverage at 0%, the description must explain parameters. It explicitly covers 'project' and 'session_id' via 'Filter by project and/or session', giving meaning to those string fields. However, the 'limit' parameter is not clearly explained; the mention of 'Max 500 rows' appears to be a cap, but the relationship to the limit parameter (which defaults to 50) is ambiguous. Partial compensation for missing schema descriptions, but not complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (List), the resource (recent user prompts from the intents table), ordering (newest first), and optional filters. It is specific enough to distinguish from siblings like search_intents, which implies a search-oriented purpose. No ambiguity in what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides filtering options (by project and/or session) but does not explain when to choose this tool over alternatives such as search_intents, memory_timeline, or save_intent. There is no mention of when not to use it or any conditions that would make another tool more appropriate. Usage context is only implied by the action of listing recent prompts.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_answerA
Generate and verify a cited answer using the configured reasoning LLM. Evidence carries recording dates; when records about the same subject disagree, the latest one gives the current value and the answer names the value it replaced. First runs negative retrieval: a contradiction-seeking second search; a score >= 0.60 hands both sides to the reader and answers with a caveat (MEMORY_CONTRADICTION_POLICY=abstain refuses instead), 0.30-0.60 answers with a caveat (see negative); MEMORY_CONTRADICTION_SCORER=jev scores the pairs with TypeSafe's Jev instead of the LLM. Up to one missing-relation retrieval and eight LLM calls including inversion retry and bounded quote repair. Explicit project required. Citation offsets refer to returned evidence content. Ordinary recall remains local.
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | ||
| limit | No | ||
| query | Yes | ||
| branch | No | ||
| project | Yes | ||
| followup | No | ||
| max_bytes | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With only sparse annotations (readOnly false, destructive false), the description carries the full behavioral burden and does so thoroughly. It discloses the negative-retrieval flow, contradiction thresholds and policies, alternate scorer, retry and repair call limits, project requirement, citation offset semantics, and local-vs-remote distinction. There is no contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence carries operational information that an agent needs to invoke the tool correctly, and the purpose is front-loaded. No filler or redundant restatement of the tool name or schema is present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity, absence of an output schema, and sparse annotations, the description covers the core algorithm, policy branches, call budgets, and project requirement well. It is slightly incomplete on concrete return-value structure, but the cited-answer behavior is described in enough detail for an agent to understand what it will receive.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the 7 undocumented parameters, but it mentions only 'Explicit project required.' It does not explain query, type, limit, branch, followup, or max_bytes beyond what the raw schema already states; the citation-offset sentence is about output, not parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence states a specific verb and resource: 'Generate and verify a cited answer using the configured reasoning LLM.' It also differentiates from related recall tools by ending with 'Ordinary recall remains local,' so an agent can tell this tool apart from memory_recall and similar siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when this tool is appropriate: it is the LLM-backed, cited-answer generator, in contrast to local recall. It does not name sibling tools explicitly or say 'use X instead when Y,' so it stops short of a full when/when-not structure, but the usage context is genuinely clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_associateA
Associative recall — brain-like spreading activation through knowledge graph. Finds memories through concept resonance, not keyword search. In 'composition' mode, finds minimum set of memories covering all needed concepts.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | recall=find related, composition=build solution from parts | recall |
| query | Yes | Natural language query | |
| project | No | Filter by project | |
| max_results | No | ||
| min_coverage | No | Min coverage for composition mode (0.0-1.0) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations declare readOnlyHint=false, indicating the tool may have side effects, but the description frames it as a recall operation, which implies read-only behavior. It does not disclose any potential state modifications, performance implications, or other behavioral traits. The description adds no transparency beyond the functional modes, and the mismatch with the readOnlyHint is not directly contradictory but leaves the agent uncertain about side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core concept. Every clause adds value: the mechanism, the contrast with keyword search, and the special mode. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
While the description covers the tool's purpose and modes, it does not describe the return format or output structure. Since there is no output schema, the description should indicate what the tool returns (e.g., a list of memories, scores, or a composition result). It also omits any mention of parameter interactions or edge cases, such as how max_results and min_coverage behave together. These gaps leave the agent guessing about the response shape.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 80%, so most parameters are already documented. The description adds meaningful context for the 'mode' parameter by explaining the composition mode's goal (minimum set of memories covering concepts). It also clarifies that 'query' is natural language for concept resonance rather than keywords. This goes beyond the schema's brief descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool performs associative recall through a knowledge graph, explicitly contrasting it with keyword search. It also distinguishes the two modes (recall and composition), which differentiates it from siblings like memory_search_by_tag and memory_timeline. The verb 'finds' and resource 'memories' are specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains that this is for associative recall rather than keyword search, implying when to use it. It also describes the composition mode for building solutions from parts. However, it does not explicitly name alternative tools or provide clear when-not-to-use scenarios, only a general contrast with keyword search.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_conceptsCRead-onlyIdempotent
List or search concepts in the knowledge graph.
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | Filter by node type | |
| limit | No | ||
| query | No | Search concepts by name | |
| include_memories | No | Include linked knowledge records |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the description does not need to repeat these. However, the description adds no additional behavioral context (e.g., that it returns a list without side effects), which is acceptable but not enhancing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with no redundancy or fluff. It is efficiently structured, though it could benefit from a bit more detail without becoming verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the large number of sibling tools (many memory-related), the description is too sparse to provide adequate context. It does not hint at what 'concepts' specifically means, how this differs from other search tools, or any special behavior like returning linked memories when include_memories is true.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides descriptions for 3 of 4 parameters (type, query, include_memories), but the 'limit' parameter has no description, resulting in 75% coverage, below the 80% threshold. The description does not compensate by explaining any parameters, leaving the behavior of 'limit' ambiguous.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses specific verbs 'List or search' and a clear resource 'concepts in the knowledge graph,' which distinguishes it from most sibling tools. However, it does not fully differentiate from other search-related tools like memory_search_by_tag or memory_recall, so it is not a perfect 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives such as memory_recall, memory_search_by_tag, or memory_timeline. The description lacks any contextual cues to help an agent decide when to invoke this specific tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_consolidateA
Find and merge duplicate/similar knowledge records. Keeps the longest version, supersedes shorter duplicates. Reduces noise in recall results.
| Name | Required | Description | Default |
|---|---|---|---|
| dry_run | No | If true, only show what would be merged | |
| project | No | Consolidate only this project (optional) | |
| threshold | No | Similarity threshold 0.0-1.0 (default 0.75) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description transparently states that records are merged, the longest version is kept, and shorter duplicates are superseded. This clearly implies a mutating operation even without an explicit warning about side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured, covering the action, strategy, and benefit in three short sentences without any redundant wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is complete enough for a tool of this complexity, covering the operation, the merge strategy, and the result. It does not describe return values, but no output schema is expected and the dry_run parameter helps imply preview behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already provides complete parameter descriptions, including defaults and ranges. The description adds no further parameter-level meaning, so it meets the baseline but does not exceed it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the operation: finding and merging duplicate or similar knowledge records. It also explains the merging strategy and the overall benefit, leaving no ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the purpose and benefit, but it does not explicitly say when to use this tool versus alternatives like memory_delete or memory_search. Usage guidance is implied rather than directly stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_consolidate_statusARead-onlyIdempotent
v11.0 W2-G: report the consolidation daemon state — per-project last-run, active locks, recent activity. Use to verify the idle-project worker is making progress without interfering with active work.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint true, idempotentHint true, and destructiveHint false. Description aligns with these by saying it 'reports' and checks progress 'without interfering with active work,' adding useful behavioral context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Description is two short sentences with the main action front-loaded. It includes only essential detail and no filler, making it easy to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simple zero-parameter read-only nature of the tool, the description sufficiently covers what state is reported and why it is used. No output schema is present, so the listed output aspects (last-run, locks, recent activity) are enough context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Tool has no parameters, and the empty parameter schema fully documents this. There is nothing additional for the description to explain, so parameter semantics are complete.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description uses a clear verb ('report') and specifies the object (consolidation daemon state) plus concrete details like per-project last-run, active locks, and recent activity. It also states the intended purpose of verifying idle-project worker progress, making the tool's role unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Description explicitly says 'Use to verify the idle-project worker is making progress,' giving a clear use case. It does not explicitly name alternative tools or when-not-to-use conditions, but the context is sufficiently clear for this read-only status check.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_context_buildCRead-onlyIdempotent
Build optimal context for a query. Combines: spreading activation + knowledge graph + episodes + skills + self-model. The 'brain thinking' tool.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | What you need context for | |
| project | No | ||
| max_tokens | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already convey read-only, idempotent, non-destructive behavior, so the bar is lower. The description adds that it combines retrieval mechanisms, but does not explain actual processing steps, output format, or potential costs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded, with only minor filler in the 'brain thinking' metaphor. It could be more informative, but it does not waste words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the large sibling set and no output schema, the description leaves ambiguity about the return value, the meaning of 'context', parameter details, and how this differs from other memory tools. It is not complete enough for reliable selection.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33% (1 of 3 parameters described). 'query' has a terse description, while 'project' and 'max_tokens' are undocumented, and the description does not compensate for these gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a verb and resource ('build optimal context for a query') and hints at internal mechanisms, but 'optimal context' and 'brain thinking' are vague. With many similar memory/context sibling tools, it does not clearly define what distinct artifact it produces.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus memory_recall, memory_search_by_tag, memory_context, or other siblings. No conditions, alternatives, or situational examples are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_deleteADestructiveIdempotent
Delete a knowledge record. By default a soft delete: the record leaves search results and vectors but stays in the database. hard=true erases it for good, with its earlier versions, derived rows and its text in the raw call log (use for personal data or anything that must not be kept).
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Knowledge record ID to delete | |
| hard | No | Erase permanently instead of hiding; cannot be undone |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the destructiveHint/idempotentHint annotations by disclosing exactly what each mode touches: soft delete removes the record from search results and vectors while leaving it in the database, and hard delete erases earlier versions, derived rows, and the raw call-log text. This is precisely the kind of destructive-scope detail the annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences, front-loading the default safe behavior before the destructive option, with zero filler. Every clause carries information the agent needs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and annotations already covering the safety profile, the description supplies the remaining essentials for a 2-param mutation tool: default behavior, hard-mode consequences, and a use case. Nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds meaning beyond the schema's 'erase permanently instead of hiding': it enumerates the collateral (prior versions, derived rows, raw log text) and clarifies the default soft-delete behavior. That is genuine added value over the terse schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (delete) and resource (knowledge record), and crisply splits the operation into soft vs hard semantics. It does not, however, differentiate itself from the overlapping sibling memory_forget, so it stays at a clear-but-unrouted 4.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a concrete condition for the destructive path ('use for personal data or anything that must not be kept') and defines the default behavior, which is real usage context. It names no sibling alternative (e.g. memory_forget) and states no when-not, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_entity_resolveBIdempotent
v11.0 W1-F: resolve a mention to its canonical entity within a project+type. Cross-session coreference via name/alias index + embedding cosine. Returns canonical_id, matched_via, and is_new flag. Pronouns return -1.
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | Entity type: person, technology, project, company, ... | person |
| mention | Yes | ||
| project | No | general | |
| threshold | No | Cosine similarity threshold for embedding match. | |
| create_if_missing | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations indicate readOnlyHint=false and idempotentHint=true, and the description adds some context about pronoun behavior and matched_via. However, it does not disclose the potential side effect of creating a new entity when create_if_missing is true, which is a significant behavioral aspect not covered by annotations alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is mostly concise and well-structured, packing method, behavior, and return values into a short paragraph. The inclusion of 'v11.0 W1-F' at the start is extraneous and could confuse, but does not significantly harm clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the moderate complexity (5 parameters, one required) and no output schema, the description provides sufficient high-level context (return values, special behavior) but lacks detail on parameter semantics and side effects. It is adequate for a basic use case but not comprehensive for edge cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 40% (type and threshold have descriptions), and the description text does not elaborate on the undocumented parameters (mention, project, create_if_missing). Without additional explanation, an agent may not understand the full meaning of these parameters, especially 'mention' and 'create_if_missing' behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: resolving a mention to a canonical entity within a project+type, using cross-session coreference via name/alias index and embedding cosine. It also specifies the return values (canonical_id, matched_via, is_new) and the special case for pronouns returning -1, leaving no ambiguity about purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not provide any guidance on when to use this tool versus alternatives. It only explains the basic operation without context on ideal scenarios, prerequisites, or when another tool (e.g., memory_save or memory_search_by_tag) would be more appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_episode_recallBRead-onlyIdempotent
Find past episodes (experiences). Search by concepts, outcome, project, or impact.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | No | Search narrative text | |
| outcome | No | ||
| project | No | ||
| concepts | No | ||
| min_impact | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering the safety profile. The description adds no additional behavioral context such as result ordering, pagination, or what constitutes an 'episode'. It is consistent with annotations but provides no extra value beyond them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence that states the purpose and search dimensions with zero waste. It is appropriately concise for a straightforward search tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 6 parameters, no output schema, and no parameter descriptions, the description is too sparse. It fails to clarify what an 'episode' is, how impact is measured, what the default limit means, or what the return structure looks like. An agent calling this tool would lack critical information for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 17% (only the query parameter has a description). The tool description lists the searchable fields (concepts, outcome, project, impact) but does not explain their semantics or how they interact. Parameters like min_impact and limit are not elaborated, leaving agents to infer meaning from names alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool finds past episodes and lists the search dimensions (concepts, outcome, project, impact). It distinguishes from many sibling memory tools by focusing on episodes, though it does not explicitly name an alternative to differentiate from, leaving some ambiguity against memory_recall or memory_timeline.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives. It does not mention when to prefer it over memory_recall, memory_timeline, or memory_search_by_tag, nor any exclusions or prerequisites. The usage context is implied but not explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_episode_saveA
Save an episode — narrative of WHAT HAPPENED and HOW. Not just facts, but the journey: what was tried, what failed, what worked.
| Name | Required | Description | Default |
|---|---|---|---|
| outcome | Yes | ||
| project | No | general | |
| concepts | No | Key concepts involved | |
| narrative | Yes | 2-3 sentence narrative of what happened | |
| key_insight | No | The aha moment, if any | |
| impact_score | No | 0.0-1.0, how significant | |
| approaches_tried | No | ||
| frustration_signals | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already indicate this is not read-only, not idempotent, and not destructive, so the behavioral profile is partially known. The description adds that it saves an episode, but does not specify side effects like whether it updates an existing episode or always creates a new one. It does not contradict the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences, front-loading the action and then elaborating on the nuance. Every word contributes to the purpose, with no fluff or irrelevant detail. It is appropriately sized for a tool of this complexity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description gives the core purpose and hints at the type of content to include, but it does not explain the required parameters or the meaning of the outcome enum, nor does it address potential side effects or return values. Given the moderate complexity of 8 parameters, this leaves gaps for an agent trying to call it correctly. It is not fully self-contained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 50%, with outcome, project, approaches_tried, and frustration_signals lacking descriptions. The description does not clarify any of these parameters, instead focusing on the overall narrative concept. It fails to compensate for the missing schema details, so parameter semantics are weak.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (save) and the object (an episode) with a specific definition of what an episode is (narrative of what happened and how). It contrasts with just saving facts, which distinguishes it from memory_save. This is a clear, specific purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the tool is for recording narrative journeys, including what was tried, failed, and worked, which gives a clear use case. It contrasts with facts, suggesting use when a richer account is needed, but it does not explicitly name alternatives or conditions. This provides moderate guidance but falls short of explicit when-to-use instructions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_eval_contradictionsBRead-onlyIdempotent
v11.0 Phase 8: runs contradiction_detector against a labelled fixture. Requires balanced/deep mode (LLM). Returns {status: 'not_implemented', ...} if module is unavailable.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | fast | |
| fixture_path | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive, and the description adds useful behavioral details: it requires balanced/deep mode and returns {status: 'not_implemented', ...} if unavailable. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is one sentence, but the leading 'v11.0 Phase 8:' is unnecessary metadata that does not help an agent decide to invoke the tool. The rest is reasonably concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description explains purpose and a failure condition, but it omits what a successful run returns, leaves fixture_path unexplained, and does not clarify mode values beyond the contradictory requirement. It is not complete enough for reliable invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, and the description does not explain fixture_path. It mentions mode must be balanced/deep, which actually contradicts the schema default of 'fast', making mode semantics confusing rather than helpful.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool runs contradiction_detector against a labelled fixture, which distinguishes it from other memory_eval_* siblings. The version/phase prefix adds some noise, but the core purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states a prerequisite ('Requires balanced/deep mode') and a fallback behavior for missing module, but it does not explicitly explain when to choose this over the many sibling eval tools or how it fits among them.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_eval_entity_consistencyCRead-onlyIdempotent
v11.0 Phase 8: verifies entity_dedup canonicalization is stable across repeated saves of variant tag spellings.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | fast |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already indicate read-only, idempotent, and non-destructive behavior, so the bar is lower. The description's 'verifies' aligns with these annotations, but it does not add context about what happens on failure, whether a report is returned, or any side effects beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no extraneous words, and the main verb 'verifies' is front-loaded. It is compact and to the point, though the density of technical terms slightly reduces clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema or return value is described, which is important for an evaluation tool. The description also lacks context about how this fits into the broader memory evaluation workflow or what 'canonicalization stability' means in practice, making it incomplete for an agent operating in a complex domain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema fully documents the 'mode' parameter with an enum (fast/balanced/deep) and a default value, so coverage is high. The description adds no additional meaning or usage detail for the parameter, leaving the baseline score of 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific verb ('verifies') and a specific resource ('entity_dedup canonicalization'), and the phrase 'across repeated saves of variant tag spellings' narrows the scope. It is distinct from sibling evaluation tools by focusing on entity consistency, though the jargon is dense and not fully explained.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus the many sibling evaluation tools (e.g., memory_eval_long_context, memory_eval_temporal). There is no mention of use cases, prerequisites, or alternative selection criteria, leaving the agent without direction on when to invoke it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_eval_locomoBRead-onlyIdempotent
v11.0 Phase 8: run the LongMemEval-style recall+prevention scenario suite (loaded from evals/scenarios/) against the live store. Forces MEMORY_MODE=fast by default. Returns {scenarios_total, scenarios_passed, recall_at_5, recall_at_10, latency_ms, mode, llm_calls_during_eval, network_calls_during_eval}.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | fast | |
| limit | No | Cap how many scenarios to run. | |
| top_k | No | ||
| scenarios_path | No | Optional override path. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses the key behavioral detail of forcing MEMORY_MODE=fast, which is not covered by the annotations. Since the annotations already indicate read-only, idempotent, and non-destructive behavior, the description adds useful context without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence that concisely conveys the action, input source, default behavior, and return fields. There is no extraneous information, making it highly efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (4 parameters, no output schema), the description provides the return fields and default mode, which is helpful. However, it omits details about scenario format or output interpretation, leaving minor gaps for a fully autonomous agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 50% (only limit and scenarios_path have descriptions). The tool description does not clarify the meaning of mode or top_k, nor does it explain the return field semantics beyond listing them. This leaves significant ambiguity for the agent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: running a LongMemEval-style recall+prevention scenario suite against the live store. It differentiates from sibling eval tools by specifying the scenario type, though it does not explicitly name the alternative tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no explicit guidance on when to use this tool versus the many sibling eval tools (e.g., memory_eval_recall, memory_eval_temporal). It implies usage for LongMemEval-style scenarios but doesn't state conditions or alternatives directly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_eval_long_contextCRead-onlyIdempotent
v11.0 Phase 8: large-context recall scenario. Saves N records and queries them at the tail. Reuses eval_harness scenarios tagged 'long_context' if present.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | fast | |
| top_k | No | ||
| n_records | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description states that the tool 'Saves N records', which implies write/create behavior, but the annotations declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false. This is a direct contradiction and could mislead an agent about side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is brief and mostly to the point, but the opening 'v11.0 Phase 8' is version/phase noise that does not help an agent. The remaining sentences are functional but somewhat vague.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description lacks enough context for an agent to confidently invoke the tool: it does not define what 'large-context recall' means in practice, what output to expect, how mode affects behavior, or how this evaluation relates to the many sibling evaluation tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides no parameter descriptions, and the description only loosely maps 'N records' to n_records and 'queries them at the tail' to the evaluation behavior. The meanings of mode and top_k, and the effect of their defaults, are not explained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies a large-context recall scenario and states that it saves N records and queries them at the tail, which makes the core evaluation behavior clear. It is reasonably distinguishable from sibling memory_eval_* tools by the explicit 'long_context' tag mention.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives only a subtle hint about reusing eval_harness scenarios tagged 'long_context', but does not explicitly say when to choose this tool over sibling evaluation tools such as memory_eval_recall or memory_eval_locomo. No clear use-case boundaries are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_eval_recallCRead-onlyIdempotent
v11.0 Phase 8: generic recall benchmark on a dataset path or a small built-in fixture. Same payload shape as memory_eval_locomo.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | fast | |
| limit | No | ||
| top_k | No | ||
| dataset_path | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds little beyond the annotations. It mentions accepting a dataset path or built-in fixture but does not disclose expected side effects, return behavior, or how the benchmark is executed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is brief and to the point, with no redundant information. It efficiently conveys the core purpose in two sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is incomplete for a benchmark tool. It does not explain what the benchmark measures, what output to expect, how to interpret results, or any additional context needed to use the tool effectively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only dataset_path is implicitly referenced via 'dataset path or built-in fixture'. The parameters mode, limit, and top_k are not explained, and the schema provides no descriptions, leaving most parameters ambiguous.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly indicates a generic recall benchmark and distinguishes it from memory_eval_locomo by referencing the same payload shape. It names the resource (recall) and implies evaluation, though it could be more explicit about the exact action.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool over alternatives. The reference to memory_eval_locomo is helpful but does not explain selection criteria or use cases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_eval_temporalCRead-onlyIdempotent
v11.0 Phase 8: temporal recall using temporal_kg + temporal_filter. Returns {status: 'not_implemented', ...} when modules are missing.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | fast | |
| limit | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses that it returns a 'not_implemented' status when modules are missing, which is a useful behavioral detail. However, it does not describe any other side effects, return formats, or error conditions beyond that, and the read-only and idempotent annotations already cover safety.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that directly states the purpose and a key behavior. It is concise and avoids unnecessary fluff, though it does include version and phase information that might be considered extraneous.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simple schema and lack of output schema, the description provides minimal context. It does not explain what the tool actually returns (beyond the not_implemented case), how mode or limit influence results, or what temporal recall entails. This leaves significant gaps for an agent trying to decide whether to invoke it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema defines 'mode' with an enum and a default, and 'limit' as an integer, but the description does not explain their meanings or how they affect the operation. The enum values are self-explanatory to some degree, but 'limit' is left completely undefined.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states it performs 'temporal recall using temporal_kg + temporal_filter', which gives a general sense of the operation, but the verb 'temporal recall' is vague and it does not clearly distinguish from sibling memory_eval_* tools. It also mentions a version and phase, which adds context but not clarity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance is provided on when to use this tool versus alternatives. The mention of 'Phase 8' hints at a workflow, but there is no direct comparison to other memory_eval_* tools or any clear usage conditions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_explain_searchARead-onlyIdempotent
v11.0: same as memory_search_fast but returns a per-tier breakdown (fts/semantic/graph/fuzzy/hyde with raw scores, the merged RRF list, rerank_applied flag, embedding_space). Use to debug why a record did or didn't surface for a query.
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | all | |
| limit | No | ||
| query | Yes | ||
| project | No | ||
| embedding_space | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses the key output behaviors: per-tier breakdown with raw scores, merged RRF list, rerank_applied flag, and embedding_space. It does not contradict the read-only, idempotent, and non-destructive annotations, and adds meaningful detail about what the tool returns beyond what annotations imply.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that immediately conveys the relationship to memory_search_fast, the added output details, and the intended use case. No unnecessary words or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description explains the output structure and use case, which are important for a debugging tool. However, it lacks any explanation of the parameter semantics (especially type and project) and does not describe the expected behavior when parameters are omitted. Given the moderate complexity of the tool (5 parameters, no output schema), the description is only partially complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides zero parameter descriptions (0% coverage). The description only mentions embedding_space, leaving query, type, limit, and project unexplained. The parameter names are somewhat self-explanatory, but the enum values for type and the role of project are not described, so the description fails to compensate for the lack of schema documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: it is a variant of memory_search_fast that returns a per-tier breakdown for debugging search results. It explicitly names the sibling tool and the specific use case, leaving no ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says to use it 'to debug why a record did or didn't surface for a query,' which provides clear usage guidance. However, it does not explicitly state when not to use it (e.g., for normal search use memory_search_fast), though this is strongly implied by the 'same as' phrasing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_exportARead-onlyIdempotent
Export all knowledge as JSON for backup or migration. Includes knowledge, sessions, and relations.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Export only this project (optional) | |
| save_to_file | No | Save to <memory-dir>/backups/ (default true) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already convey read-only, idempotent, non-destructive behavior; description adds context by noting the export includes knowledge, sessions, and relations, and that save_to_file defaults to persisting to a backups directory.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences deliver action, scope, format, and purpose with no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter tool with no output schema, the description adequately covers what is exported and why; it could be slightly more explicit about the return shape when save_to_file is false.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers both parameters with descriptions; the tool description does not add much beyond the schema, so baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states a specific action (Export), a specific resource (all knowledge), a format (JSON), and a clear purpose (backup or migration), which distinguishes it from narrower memory retrieval tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly frames the tool for backup or migration scenarios, giving an agent clear context for when to choose it; no explicit alternative comparisons are needed given the broad export scope.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_extract_sessionA
Get pending session transcripts for knowledge extraction. Previous sessions are auto-captured on exit. Use action='list' to see pending, 'get' to read transcript, then save knowledge via memory_save, then 'complete' to mark as processed.
| Name | Required | Description | Default |
|---|---|---|---|
| chunk | No | Chunk number for large transcripts (0-based) | |
| action | Yes | list: show pending sessions. get: return transcript data. complete: mark as done. | |
| session_id | No | Session ID (required for 'get' and 'complete') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses that sessions are auto-captured on exit and that 'complete' marks them as processed, which are side effects. This is consistent with readOnlyHint=false. It doesn't mention any destructive actions, aligning with destructiveHint=false, but doesn't elaborate on state changes beyond marking processed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences. The first states the purpose, the second gives the workflow. There is no redundancy or unnecessary detail, making it easy to parse and act on.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 3 parameters and no output schema, the description covers the primary workflow and parameter usage. It doesn't mention output format, error handling, or chunking details, but given the simplicity and the presence of schema descriptions, it is sufficiently complete for an agent to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already covers parameter meanings (action enum, session_id required for get/complete, chunk for large transcripts). The description adds value by explaining the workflow order and how the parameters are used together, which goes beyond simple definitions. This is helpful, though not essential.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: retrieving pending session transcripts for knowledge extraction, with a specific verb and resource. It also mentions automatic capture on exit, providing context. However, it does not explicitly differentiate from sibling tools like memory_get or memory_recall, so it's not a perfect 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a clear workflow: use action='list' to see pending, 'get' to read, then save via memory_save, and 'complete' to mark processed. It mentions a subsequent step (memory_save) which gives usage context. It doesn't explicitly state when not to use this tool or compare it to alternatives, but the workflow is helpful.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_forgetADestructiveIdempotent
Apply retention policy: archive stale records (>180d, never recalled, low confidence), purge very old archived records (>365d). Keeps memory clean.
| Name | Required | Description | Default |
|---|---|---|---|
| dry_run | No | If true, only show what would be affected |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The tool is marked destructiveHint=true, and the description mentions 'purge', which implies deletion. However, it does not explicitly state that purged records are permanently removed, that archived records are moved to a different location, or that the operation might affect linked data. The annotations cover the core destructive nature, but the description adds limited behavioral detail beyond the obvious archive/purge actions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured, using exactly two sentences. The first states the primary action and conditions, the second summarizes the benefit. No unnecessary words or redundant explanations are present; every part contributes to understanding the tool's function.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
While the description is clear about the operation, it lacks broader context. It does not explain how this tool fits into the memory management lifecycle, nor does it mention what happens after archiving (e.g., whether archived records are still queryable). Given the existence of sibling tools like memory_delete and memory_consolidate, the description could benefit from clarifying the distinction between this automated retention task and those manual/alternative operations. The low overall complexity keeps this from being a serious gap, but the description is not fully complete on its own.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter, dry_run, has a clear description ('If true, only show what would be affected') that goes beyond the boolean type to explain its effect on the tool's behavior. This fully clarifies the parameter's meaning and how to use it, making the description self-sufficient for parameter understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's primary purpose: applying a retention policy that archives stale records and purges very old archived records. The verbs 'archive' and 'purge' are specific, and the condition thresholds (>180d, >365d) provide concrete scope. This is a clear, unambiguous statement of intent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not explicitly contrast this tool with alternatives like memory_delete or memory_consolidate. While it implies use for routine memory maintenance ('Keeps memory clean'), it lacks explicit guidance on when to choose this tool over a direct delete or when not to use it (e.g., if only a single record needs removal). The purpose is clear, but usage boundaries are not spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_getARead-onlyIdempotent
Batched fetch by ID — complement to memory_recall(mode='index'). Returns full content for ONLY the IDs the caller chose after inspecting an index. Typical 3-layer flow: recall(mode='index') → pick IDs → memory_get(ids=[...]).
| Name | Required | Description | Default |
|---|---|---|---|
| ids | Yes | Knowledge record IDs (max 50 per call; extras are silently dropped) | |
| detail | No | 'summary' truncates content to 150 chars, 'full' returns everything | full |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, covering safety. The description adds that it returns full content only for selected IDs, which clarifies output behavior but does not introduce additional side-effect or permission information. This is slightly above baseline because the description reinforces the non-destructive nature.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, using three sentences to convey purpose, relationship to sibling tool, and usage flow. No redundant or filler content; every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simple parameter set and absence of an output schema, the description fully equips an agent to decide when and how to use the tool. It provides the necessary context about the intended workflow and the tool's role within it, making it complete for this context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already provides 100% coverage for both parameters including the enum and default for 'detail', and the description of 'ids' mentions the maximum and silent dropping in the schema. The description itself does not add further parameter semantics beyond what the schema already states, so it stays at the baseline of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'fetch' and the resource 'memory by ID', and explicitly distinguishes it from the sibling tool memory_recall by describing it as a complement and specifying its role in a typical flow (after index inspection). This makes the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly names the alternative tool (memory_recall(mode='index')) and provides a clear usage flow: recall index first, then pick IDs, then call memory_get. This leaves no doubt about when to use this tool versus others.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_graphBRead-onlyIdempotent
Query the unified knowledge graph. Returns neighborhood of a node: connected rules, skills, memories, concepts, entities.
| Name | Required | Description | Default |
|---|---|---|---|
| node | Yes | Node name or ID to explore | |
| depth | No | Traversal depth (1-3) | |
| types | No | Filter by node types (rule, skill, concept, etc.) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate read-only, idempotent, and non-destructive behavior. The description's use of 'Query' aligns with these hints and adds no contradictory claims. It does not elaborate on edge cases like missing nodes or empty results, but given the annotation coverage, this is sufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that directly conveys the core functionality and return types. It is concise, well-structured, and contains no superfluous information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a query tool with a relatively simple output, the description adequately informs about the nature of results (connected rules, skills, memories, concepts, entities). It does not specify the exact output format (e.g., list vs. graph), but given the tool's simplicity and lack of an output schema, this is acceptable and does not leave major gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema provides descriptions for all three parameters (node, depth, types) with coverage of 100%. The tool description adds no additional semantic detail about these parameters, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's action ('Query the unified knowledge graph') and its primary output ('neighborhood of a node' with listed connected types). However, it does not distinguish itself from several closely related sibling tools (e.g., memory_graph_index, kg_at, or memory_recall), so its unique purpose is not fully explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus the many sibling tools that query or manipulate memory/knowledge graphs. There is no mention of appropriate contexts, prerequisites, or scenarios where this tool is preferred over alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_graph_indexAIdempotent
Reindex CLAUDE.md rules and skills into the knowledge graph. Run after modifying CLAUDE.md or adding new skills.
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | all |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate the tool is idempotent and not destructive. The description adds the maintenance context (run after changes) and clarifies that it updates the knowledge graph, but it doesn't disclose any side effects beyond reindexing. Given the annotation coverage, this is acceptable but minimal additional behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core action and followed by the trigger condition. There is no redundancy or filler; every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple tool with one optional parameter and no output schema, the description covers the action and when to run it. However, it omits any explanation of the parameter, which means an agent might not realize it can target subsets. Given the default 'all', a full reindex is achievable without parameters, but the omission limits flexibility.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, and the description does not mention the 'target' parameter or its possible values (all, claude_md, skills, rules). While the enum values are somewhat self-explanatory, the description fails to explain how they map to the reindex scope, leaving an agent to infer the parameter's role. This is a significant gap that the description should compensate for.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action: 'Reindex CLAUDE.md rules and skills into the knowledge graph.' It names the verb (reindex), the resource (CLAUDE.md rules and skills), and the destination (knowledge graph), and it includes a trigger condition. This distinguishes it from sibling reindex tools like memory_rebuild_fts and memory_rebuild_embeddings, which target different indexes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Run after modifying CLAUDE.md or adding new skills,' which is a clear, actionable trigger. It does not mention alternatives or exclusions, but the context is sufficient for an agent to know when this tool is appropriate, especially given the knowledge-graph specificity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_graph_statsARead-onlyIdempotent
Knowledge graph statistics: nodes, edges, communities, top concepts, health metrics.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already indicate read-only, non-destructive behavior, and the description does not contradict them. The mention of 'health metrics' hints at additional informational output, but no further behavioral traits are disclosed beyond what the annotations cover.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and directly lists the output categories. Every word contributes meaning, and the structure is clear without redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that no output schema is provided, the description still lists the key statistics returned. This is sufficient for a user to understand the tool's purpose and typical output, making the description complete in context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has no parameters, so the schema coverage is complete. The description adds no parameter-specific semantics, but none are needed; the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool provides knowledge graph statistics, listing specific elements: nodes, edges, communities, top concepts, and health metrics. This is distinct from sibling tools like memory_graph (which likely fetches the graph structure) and memory_concepts (which focuses on concepts).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description does not explicitly state when to use this tool versus alternatives. It merely describes what it returns, without noting situations where a user should prefer this over memory_graph or memory_concepts, or when it might be inappropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_historyBRead-onlyIdempotent
View version history for a knowledge record. Shows the chain of superseded versions (newest → oldest), enabling time-travel through knowledge evolution.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Knowledge record ID to get history for |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Given the annotations already declare the tool as read-only and non-destructive, the description adds useful behavioral context by specifying that it returns a chain of superseded versions ordered newest-to-oldest. This goes beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with no unnecessary words. It is well-structured, with the core action stated first and a clarifying detail second.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description provides essential context about the output: it shows a chain of superseded versions and the ordering (newest → oldest). This is sufficient for basic understanding, though it does not detail the fields of each version entry.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema covers 100% of parameters (only 'id') with a description that matches the tool's purpose. The description does not add any extra detail about the parameter, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('View') and resource ('version history for a knowledge record'), clearly indicating the tool's function. It does not explicitly name a sibling alternative, but the focus on superseded versions (newest → oldest) differentiates it from tools like memory_get or memory_timeline.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives. It does not mention scenarios or conditions that would make this the preferred choice over memory_get or memory_timeline.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_index_passagesB
Build local passage indexes before evidence searches. Repeat using next_after_id until remaining=0.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| project | Yes | ||
| after_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations are all false, so the description carries the burden of behavioral disclosure. It reveals that the tool is iterative (repeat until remaining=0) and that it builds indexes, implying a state-changing operation. However, it does not disclose side effects, whether it overwrites existing indexes, performance implications, or what 'remaining' refers to. The iterative hint is useful but incomplete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, both information-dense and front-loaded with the primary purpose. The repeat instruction is a necessary operational detail. No wasted words, though the second sentence could be slightly clearer about the parameter name (after_id vs next_after_id).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and no annotation support, the description gives the core workflow (build indexes, paginate with after_id) but omits what 'remaining' means, how results are returned, and what happens if the tool is called without prior setup. It is adequate for a simple paginated indexing tool but leaves gaps around return values and side effects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for all three parameters. It mentions 'next_after_id' in the repeat instruction, which maps to the after_id parameter, but it does not explain limit or project semantics. The description adds minimal value beyond the schema for parameter understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Build') and resource ('local passage indexes') and connects it to evidence searches, which distinguishes it from sibling tools like memory_recall or memory_search_fast. However, it doesn't explicitly name a sibling alternative or clarify what 'local' means relative to other indexing tools like memory_graph_index or memory_rebuild_fts.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear when-to-use signal: 'before evidence searches.' It also provides an explicit repeat instruction ('Repeat using next_after_id until remaining=0'), which is actionable usage guidance. It does not explicitly state when not to use it or name alternatives, but the context is sufficient for an agent to select it for pre-search indexing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_observeA
Save a lightweight observation (auto-capture). No dedup, no ChromaDB — fast and cheap. Use for tracking file changes, tool usage, and session activity. Observations auto-cleanup after 30 days.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | general | |
| summary | Yes | What happened (e.g. 'Modified auth controller') | |
| tool_name | Yes | Which tool triggered this (Write, Edit, Bash, etc.) | |
| files_affected | No | List of affected file paths | |
| observation_type | No | Type of observation | change |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses key behaviors beyond annotations: auto-capture, no dedup, no ChromaDB, fast/cheap, and 30-day auto-cleanup, which informs the agent about retention and performance characteristics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no fluff, main action and differentiators front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Sufficiently complete for a simple save tool; includes purpose, retention, and performance; no output schema needed, though could mention return value or side effects.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema already covers most parameter descriptions (80%); description adds no extra parameter-level detail, so baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states it saves lightweight observations for tracking file changes, tool usage, and session activity, and differentiates from heavier memory tools by noting absence of dedup and ChromaDB.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly identifies use cases (tracking file changes, tool usage, session activity) and implies alternative for lightweight, fast, cheap saves with 30-day auto-cleanup; could be more explicit about contrasting with memory_save but enough.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_perf_reportARead-onlyIdempotent
v11.0: dump in-process telemetry counters (search_total_ms, embed_ms, fts_ms, vector_ms, llm_calls, network_calls) plus persistent embedding_cache stats. Use to verify the fast hot path stays clean.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses that it dumps telemetry counters and embedding cache stats, adding context beyond the annotations which already declare read-only, idempotent, and non-destructive. It does not mention any side effects or limitations, but given the annotations cover safety, the description provides useful behavioral context about what data is returned. No contradictions with annotations are present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, well-structured sentence. It front-loads the action ('dump') and immediately lists the specific counters, followed by a clear usage note. The version prefix 'v11.0' is minor and does not detract. Every word contributes value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a report tool with no parameters and no output schema, the description covers the essential information: what data it dumps, the purpose, and the intended use case. It does not describe the exact output format, but that is not strictly necessary for an agent to decide to call it. The annotations cover safety, so the description is sufficiently complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and the schema coverage is 100% (vacuous). The description does not need to explain parameters. Per the calibration baseline for 0 params, a score of 4 is appropriate. No parameter information is missing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's action: 'dump in-process telemetry counters' and enumerates specific counter names, distinguishing it from sibling tools like memory_stats or memory_graph_stats that likely serve different metrics. The mention of 'persistent embedding_cache stats' adds specificity. The purpose is unambiguous and identifiable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit usage guidance: 'Use to verify the fast hot path stays clean.' This gives a clear when-to-use context. It does not explicitly mention alternatives or when not to use, but the stated purpose is sufficient for an agent to decide if this tool is appropriate compared to other reporting tools. Since there are no parameters, no further usage details are needed.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_rebuild_embeddingsADestructiveIdempotent
v11.0: re-encode every record (or every record in a given embedding space) and update the binary + float32 vectors. Idempotent. Pass embedding_space='code' to refresh only code rows after switching the code embedder. Returns {rebuilt: int, skipped: int}.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| project | No | ||
| batch_size | No | ||
| embedding_space | No | Optional: only re-encode rows in these spaces. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Description says 'update' and 're-encode', clearly indicating mutation, and adds idempotency info. Annotations already mark it destructive and not read-only, so no contradiction; slight lack of detail on potential data loss or performance impact.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Very concise: two sentences covering functionality, idempotency, an example, and the return shape. No superfluous words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Mentions the return shape ({rebuilt, skipped}) but leaves significant gaps: meaning of limit, project, batch_size is unexplained, and there is no warning about side effects or performance. Given the tool mutates records, more context is needed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only embedding_space has a description in the schema, and the description adds meaning for it. The other three parameters (limit, project, batch_size) remain undocumented, and the description does not compensate for the low 25% coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the action (re-encode and update vectors), the resource (records in an embedding space), and provides a concrete usage example (embedding_space='code').
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit scenario for using the tool (refreshing code rows after switching the embedder) and mentions idempotency, but does not contrast with alternative tools or state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_rebuild_ftsADestructiveIdempotent
v11.0: drop and rebuild the SQLite FTS5 virtual table from knowledge rows. Useful after migrations or content_type column changes that the FTS triggers didn't see. Returns {rebuilt: int}.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds details beyond the annotations by mentioning the return value '{rebuilt: int}' and the reason for use, while the destructive and idempotent nature is already captured in annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences, leading with the action and then providing context, with no extraneous detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers what the tool does, when to use it, and its return value, which is sufficient for a simple maintenance operation; a note on potential side effects (e.g., temporary unavailability) would make it fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are no parameters, so the schema is trivially covered; the description adds no parameter-specific information beyond what the empty schema implies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the specific action ('drop and rebuild the SQLite FTS5 virtual table') and the resource ('knowledge rows'), making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides a clear condition for when to use the tool ('after migrations or content_type changes that the FTS triggers didn't see'), though it does not explicitly contrast with alternative tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_recallARead-onlyIdempotent
Search ALL memory: decisions, solutions, facts, lessons from ALL past sessions. 6-stage pipeline: FTS5+BM25 → semantic → fuzzy → graph → (optional) CrossEncoder → (optional) MMR. Default: hybrid mode (BM25 + semantic + RRF). Use BEFORE starting any task. v11.0: routes to fast hot path when MEMORY_MODE=fast (default). Use memory_search_fast / memory_explain_search for explicit fast routing.
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Progressive-disclosure mode: 'search' (default) = normal results, 'index' = ultra-compact metadata only (id+title+score+type+project+created_at, ~40-60 tok/hit, no cognitive expansion, use memory_get(ids=...) to fetch full content), 'timeline' = chronological compact view; 'context' = source excerpts; 'evidence' = indexed passage search with bounded follow-up retrieval | search |
| type | No | all | |
| limit | No | ||
| query | Yes | What to search for | |
| branch | No | Filter by git branch (also includes branch-agnostic records) | |
| detail | No | Level of detail: 'compact' ~50 tokens/result (id+title+score), 'summary' truncates content to 150 chars, 'full' returns everything, 'auto' picks based on query complexity (paths/urls/code → full, short → compact). Ignored when mode!='search'. | full |
| fusion | No | Score fusion method: 'rrf' = Reciprocal Rank Fusion (better multi-tier ranking), 'legacy' = original additive scoring | rrf |
| intent | No | Filter by classified intent (question|procedural|fact|decision|problem|solution|incident|plan) | |
| rerank | No | Enable CrossEncoder re-ranking for higher precision (adds ~30ms latency) | |
| topics | No | Filter results to records tagged with any of these topics (from deep enrichment) | |
| diverse | No | Enable MMR diversity to reduce redundant results (useful for broad queries) | |
| project | No | Filter by project name | |
| entities | No | Filter by extracted entity names (technology/person/project, case-insensitive) | |
| neighbors | No | Timeline/context modes: records before/after each hit; context accepts 0–3. | |
| fill_budget | No | Context mode: search deeper (up to 100 hits) and keep whole hits in rank order until context_max_chars is used, instead of excerpting the top `limit` hits to fit. Neighbours default to 0 here; pass `neighbors` to keep each hit with the records around it. | |
| expand_budget | No | Max number of additional records to include via graph expansion | |
| decisions_only | No | Return only structured decisions (v8.0): type=decision AND tags contain 'structured'. Results include parsed schema payload under 'decision'. | |
| expand_context | No | Add graph-related records (1-hop neighbors via knowledge graph) as 'expansion' results | |
| missing_relation | No | Evidence mode: missing relation; subject must occur in the original question. | |
| context_max_chars | No | Context character budget; evidence mode uses the same value as a stricter UTF-8 byte budget. | |
| evidence_followup | No | Evidence mode: allow one search for an explicitly supplied missing_relation. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly/idempotent/non-destructive/openWorld=false, so the safety profile is covered. The description adds genuinely useful context beyond that: the 6-stage pipeline, the default hybrid mode (BM25 + semantic + RRF), and the MEMORY_MODE=fast hot-path routing behavior. It stops short of describing result volume or latency trade-offs of the optional stages.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with scope and the imperative use case, then pipeline details and routing. Five sentences is dense but each carries information; minor redundancy in listing pipeline stages that the mode parameters also imply.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 21-parameter, no-output-schema tool, the description supplies the search scope, default behavior, routing rules, and sibling alternatives that an agent needs to call it correctly. Return-shape specifics are delegated to the mode parameter's own enum documentation, which is appropriate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 90%, so the baseline is 3; the description adds meaning on top by explaining the default hybrid fusion mode and the environment-driven fast-path routing, which are not encoded in the schema. It doesn't walk through individual parameters, but the schema already does that thoroughly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (search) and resource scope (ALL memory: decisions, solutions, facts, lessons from ALL past sessions), and explicitly distinguishes itself from siblings by naming memory_search_fast and memory_explain_search. An agent can tell it apart from the fast-path tools without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear imperative ('Use BEFORE starting any task') and names the alternative tools for explicit fast routing when MEMORY_MODE=fast. However, it does not state when NOT to use this tool (e.g. narrow lookups better served by memory_get or memory_search_by_tag), leaving some routing decisions to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_recall_iterativeARead-onlyIdempotent
v11.0 W1-B: IRCoT-style iterative retrieval. Decomposes the query into sub-questions, retrieves per sub-question, and asks a planner LLM whether more retrieval is needed. Best for multi-hop questions. Returns unified evidence + provenance per iteration.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | ||
| project | No | ||
| llm_model | No | configured | |
| max_iters | No | ||
| k_per_iter | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, covering safety. The description adds useful behavioral detail about the internal process (decomposing, retrieving, asking a planner) and the return format ('unified evidence + provenance per iteration'). This goes beyond annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, with the main purpose stated up front. The version prefix 'v11.0 W1-B' is extraneous and could be removed, but it does not impair understanding. The core sentences are efficient and informative, though the parameter details are missing, which slightly reduces the structure's completeness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with 5 parameters and no output schema, the description is incomplete. It explains the high-level algorithm but does not describe how max_iters or k_per_iter affect behavior, what the project or llm_model parameters do, or the exact return structure beyond a vague 'unified evidence + provenance per iteration.' An agent cannot confidently set parameters or parse output without additional information.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain any of the 5 parameters (query, project, llm_model, max_iters, k_per_iter). It mentions retrieval and iteration but does not define how these parameters control behavior. Since the description carries the full burden for parameter semantics and fails to address them, the score is minimal.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: IRCoT-style iterative retrieval that decomposes queries into sub-questions, retrieves per sub-question, and uses a planner LLM. It explicitly notes it is 'Best for multi-hop questions,' which differentiates it from the simpler sibling 'memory_recall'. This gives a specific verb, resource, and scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear usage context by stating 'Best for multi-hop questions,' which signals when to use this tool over alternatives like memory_recall. It implies that for simpler, single-hop queries, other tools are preferable, but it does not explicitly name them or provide exclusions, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_reflect_nowB
Run reflection (the 'sleep' process). Consolidates knowledge, finds patterns, generates skill proposals, updates self-model.
| Name | Required | Description | Default |
|---|---|---|---|
| scope | No | quick=dedup only, full=digest+synthesize, weekly=deep analysis | full |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses that the tool updates the self-model and consolidates knowledge, which implies state changes, but it does not explicitly state side effects, reversibility, or what happens to data. Annotations indicate readOnlyHint=false and destructiveHint=false, which are consistent with the description but add no extra detail. The description goes slightly beyond the annotations by mentioning 'updates self-model' but does not fully elaborate on behavioral implications.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is exceptionally concise—two sentences that pack all essential information. It front-loads the primary purpose and then lists the key actions without any fluff. There is no redundant or irrelevant text, making it easy to parse and understand quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has only one parameter, no output schema, and a clear description of its actions, the context is largely complete. However, the description does not mention what the tool returns (e.g., success confirmation or a summary of reflection), and given the large sibling set, a note on when to prefer this over memory_consolidate would improve completeness. Still, the core functionality is adequately covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description covers the 'scope' parameter with a full enum and descriptions for each option (quick, full, weekly), so parameter semantics are already well-documented. The tool description does not add extra context beyond what the schema provides. With 100% schema coverage, the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function with a specific verb ('Run reflection') and resource ('the sleep process'), and lists concrete actions (consolidates, finds patterns, generates proposals, updates self-model). While it doesn't explicitly contrast with sibling tools like memory_consolidate or self_reflect, the specificity of the actions makes the purpose clear enough for an agent to understand what it does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus its siblings (e.g., memory_consolidate, self_reflect). It does not mention conditions, prerequisites, or alternatives. An agent would have to infer usage from the name and description alone, which is insufficient given the many overlapping memory and reflection tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_relateB
Create a typed relation between two knowledge records. Enriches graph expansion in Tier 4 search. Types: causal, solution, context, related, contradicts.
| Name | Required | Description | Default |
|---|---|---|---|
| type | Yes | Relation type | |
| to_id | Yes | Target knowledge record ID | |
| from_id | Yes | Source knowledge record ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate a non-read-only, non-idempotent, non-destructive operation. The description's 'Create' aligns with these, and it adds the graph expansion context. However, it does not disclose side effects like whether duplicates are prevented, ID existence checks, or return behavior, so it adds minimal value beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the core action. The types list is redundant with the schema but doesn't bloat it. No fluff; concise enough for quick scanning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the agent doesn't know what the tool returns. The description doesn't mention return value, error conditions, or prerequisites (e.g., IDs must exist). For a write operation of moderate complexity, this is adequate but not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% – each parameter has a description in the schema. The description repeats the enum types (causal, solution, etc.) but adds no new meaning beyond what the schema already provides. It doesn't explain how from_id/to_id relate semantically beyond 'source' and 'target'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Create a typed relation') and the resource ('between two knowledge records'), with a specific list of relation types. It distinguishes itself from sibling tools like memory_save or memory_update by focusing on relations, though it does not explicitly differentiate from kg_add_fact which might overlap.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a context hint ('Enriches graph expansion in Tier 4 search') but no explicit when-to-use vs alternatives. The purpose implies usage for linking records, but there is no guidance on when not to use it or which alternative to pick, so it's only implied.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_reportAIdempotent
Activity report for a project (or all projects) over a period: today (period=day), this week (week, ISO Monday-Sunday), this month (month), all time (all) or custom since/until dates; offset=-1 gives the previous day/week/month ('last week'). Sections: summary numbers with deltas against the previous equal period, key decisions with their WHY, solutions and fixes, errors with recurring patterns and lessons, open next steps and pitfalls from session summaries, most touched files, entities/technologies, and a day-by-day timeline. Every item carries source IDs (#id -> memory_get). Built deterministically from stored records, no LLM; include_llm_summary=true adds an optional paragraph from the configured LLM. Periods use the local timezone (or tz). save=true writes the Markdown to /reports//-.md. Use it when the user asks what happened, for a status/progress report, a weekly summary or a retrospective.
| Name | Required | Description | Default |
|---|---|---|---|
| tz | No | IANA timezone for period boundaries, e.g. Europe/Berlin; default: MEMORY_REPORT_TZ, TZ or the system zone | |
| save | No | Also write the Markdown under <memory dir>/reports/ | |
| limit | No | Maximum items per section | |
| since | No | custom only: first day YYYY-MM-DD (inclusive) or an ISO date-time | |
| until | No | custom only: last day YYYY-MM-DD (inclusive) or an ISO date-time (exclusive); default today | |
| format | No | markdown (readable) or json (structured) | markdown |
| offset | No | day/week/month only: 0 = current, -1 = previous (yesterday, last week, last month) | |
| period | No | day = today, week = this ISO week, month = this calendar month, all = since the first record, custom = since/until | week |
| project | No | Project name; omit for all projects | |
| include_llm_summary | No | Add an LLM-written paragraph (needs a configured LLM) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial context beyond the annotations: it is built deterministically from stored records with no LLM, include_llm_summary opts into an LLM paragraph, every item carries source IDs, periods use the local timezone, and save=true writes Markdown to a concrete path. This also explains why readOnlyHint is false (it can write files) while remaining non-destructive and idempotent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the purpose and period semantics, then sections, then save behavior and usage trigger. Dense but each clause carries information; the heavy semicolon chaining makes it slightly harder to scan than an ideal definition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter, no-output-schema tool, the description effectively documents the return content (section list, source IDs, deltas vs previous equal period) as well as side effects (save path) and timezone behavior. An agent has everything needed to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents tz, save, limit, since/until, format, offset, period and project. The description restates period/offset semantics that the schema already carries (day/week/month/all/custom, offset=-1 = previous period) without adding new syntax or format detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource+scope: 'Activity report for a project (or all projects) over a period,' then enumerates the exact sections produced (summary deltas, decisions, solutions, errors, next steps, files, entities, timeline). This is clearly distinguishable from memory_timeline, memory_stats and memory_perf_report without opening their schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the triggering intents: 'Use it when the user asks what happened, for a status/progress report, a weekly summary or a retrospective.' Clear context for use, but no exclusions or named alternative tools for overlapping cases (e.g. memory_timeline vs this report).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_saveA
Save knowledge explicitly. Types: decision (MUST include WHY in context), solution, lesson, fact, convention. Saving the same words again (case and punctuation aside) replaces the stored record, so it carries the latest date; any other text, including a changed value, is stored as a new record. v10: a quality gate scores the record before save; below-threshold records are rejected with a rejected_by_quality_gate: true response (override with MEMORY_QUALITY_GATE_ENABLED=false). Use importance to surface critical decisions at recall time (boosts the final RRF score). v11.0: routes to fast hot path when MEMORY_MODE=fast (default). Use memory_save_fast for explicit fast routing.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | No | ||
| type | Yes | ||
| coref | No | Opt into v10 coreference rewrite — expand pronouns ('after this it broke') into self-contained text using recent session history. Costs ~1s LLM round-trip; default off. | |
| branch | No | Git branch this knowledge relates to | |
| filter | No | Optional content filter (pytest|cargo|git_status|docker_ps|generic_logs). Trims noisy CLI output while preserving URLs/paths/code. | |
| content | Yes | The knowledge to save | |
| context | No | Additional context, WHY for decisions | |
| project | No | general | |
| agent_id | No | Optional Claude Code subagent ID (x-claude-code-agent-id header / OTEL agent_id attribute, v2.1.139+). Lets recall trace which subagent produced this knowledge. | |
| supersede | No | Retire active records of the same project and type that this one gives a new value for: same opening words, different trailing value ("X's citizenship is Argentina" -> "... is Armenia"). Use for single-valued facts only; "likes jazz" would retire "likes rock". The retired ids are returned as `superseded`. | |
| importance | No | Recall-time boost: critical x1.5, high x1.2, medium x1.0, low x0.8. Reserve `critical` for migration-blocking decisions and security incidents. | medium |
| source_format | No | Conversation preserves dialogue structure and bypasses automatic CLI filters. | auto |
| parent_agent_id | No | Optional parent agent ID (the dispatching Agent tool / parent span). Together with agent_id forms the subagent lineage tree. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses several non-obvious behaviors: identical saves replace the existing record, changed text creates a new record, a quality gate can reject saves with a specific response, importance affects recall ranking, and MEMORY_MODE=fast routes to a hot path. This goes far beyond the sparse annotations and gives the agent a realistic model of side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded with the core action and types, then adds behavioral details in a logical progression. The version labels (v10, v11.0) add a little noise, but each sentence contributes useful information for an agent deciding whether and how to call the tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 13 parameters and no output schema, the description covers the most decision-critical behaviors: required types, context requirements, replacement semantics, quality-gate rejection, and routing. It does not describe the normal success return shape, but the schema and the detailed parameter descriptions cover most of what an agent needs to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 77%, so the schema already explains most parameters. The description adds valuable semantics by tying the decision type to the context parameter ('decision MUST include WHY in context'), explaining the recall-time effect of importance, and mentioning the quality-gate override environment variable.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action and object ('Save knowledge explicitly'), then enumerates the supported content types, making it immediately clear what the tool does. It also distinguishes itself from the sibling memory_save_fast by noting that tool is for explicit fast routing, helping an agent tell them apart.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to save knowledge and names memory_save_fast as the alternative for explicit fast routing. It does not discuss when to prefer memory_update or memory_observe, so the guidance is clear but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_save_fastA
v11.0: same as memory_save but routes through the fast hot path (skip_quality=True, no LLM, no async-blocking). Use when you want to bypass the v10 quality gate without flipping the env flag.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | No | ||
| type | Yes | ||
| branch | No | ||
| filter | No | ||
| content | Yes | ||
| context | No | ||
| project | No | general | |
| agent_id | No | Optional Claude Code subagent ID (v2.1.139+) | |
| supersede | No | Same as memory_save.supersede: retire records this one gives a new value for. | |
| importance | No | medium | |
| source_format | No | auto | |
| parent_agent_id | No | Optional parent agent ID (the dispatching span) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description adds behavioral details like 'skip_quality=True, no LLM, no async-blocking' beyond the annotations, which is valuable. However, it doesn't describe return values, side effects, or permissions, and the annotations (readOnlyHint=false, etc.) indicate a write operation without full transparency on what the write entails.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, efficient sentence that front-loads the purpose and then gives the usage condition. No wasted words; it is both concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (12 parameters, no output schema), the description is only partially complete. It covers the key differentiator from memory_save and provides usage guidance, but lacks details on behavior, return values, and parameter semantics, making it insufficient for an agent unfamiliar with memory_save.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 12 parameters but only 25% have descriptions, and the tool description provides no additional parameter explanations. It merely refers to memory_save without elaborating on any of the parameters, leaving the agent to infer from the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly indicates this tool saves memory, and the 'fast hot path' clarifies it is a variant of memory_save. It distinguishes itself from siblings by explicitly referencing memory_save, though it doesn't fully describe the saving behavior on its own.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use this tool: 'when you want to bypass the v10 quality gate without flipping the env flag.' It also names the alternative (memory_save) implicitly, providing clear routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_search_by_tagARead-onlyIdempotent
Search knowledge by tag. Returns all active records with matching tag (partial match). Useful for categorical browsing.
| Name | Required | Description | Default |
|---|---|---|---|
| tag | Yes | Tag to search for (partial match) | |
| project | No | Filter by project (optional) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, idempotent, non-destructive behavior. The description adds useful context that results are limited to active records and that matching is partial, which goes beyond the annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with no unnecessary detail. The key behavior and intended use are communicated efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description explains what is returned ('all active records') and the matching behavior, which is sufficient for a simple search tool. It does not specify output structure or edge cases, but the absence of an output schema is offset by the clear return semantics.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already documents both parameters, including 'partial match' for tag and the optional project filter. The description does not add additional meaning beyond the schema, so it sits at the baseline for full schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action ('Search knowledge by tag') and a clear result ('Returns all active records with matching tag'). The partial-match behavior and categorical-browsing use case make the tool's purpose immediately identifiable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies use for categorical browsing but does not explicitly distinguish when to use this tool over sibling tools like memory_search_fast or memory_recall. There is no direct when-to-use or when-not-to-use guidance relative to alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_search_fastARead-onlyIdempotent
v11.0: like memory_recall but with rerank=False, diverse=False forced. Deterministic fast path — zero LLM, FastEmbed-only.
| Name | Required | Description | Default |
|---|---|---|---|
| type | No | all | |
| limit | No | ||
| query | Yes | ||
| branch | No | ||
| detail | No | full | |
| fusion | No | rrf | |
| project | No | ||
| embedding_space | No | Filter to one or more embedding spaces (text|code|log|config). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses that it forces rerank=False and diverse=False, uses no LLM, and is deterministic, which is consistent with readOnly and idempotent annotations. It does not mention output shape or side effects, but the read-only behavior is already covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Very short and front-loaded; two sentences convey the key distinction without padding. The version prefix is minor but does not detract from clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
As a standalone description, it depends heavily on the unnamed memory_recall tool and omits return shape, result ordering, and parameter semantics. Given no output schema and low schema coverage, an agent would need additional context to use it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only embedding_space is described in the schema; query, limit, type, branch, detail, fusion, and project have no parameter-level explanation and the description does not clarify them beyond the memory_recall reference. With 8 parameters and 13% schema coverage, most parameter semantics remain implicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly identifies a memory search tool and distinguishes it from memory_recall by forcing rerank=False and diverse=False for a deterministic fast path. The exact search semantics rely on familiarity with memory_recall, but the name and 'like memory_recall' anchor the purpose.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says when to use it: deterministic fast path, zero LLM, FastEmbed-only, in contrast to memory_recall. It does not spell out all trade-offs or when not to use, but the deterministic/no-LLM cues are practical guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_self_assessBRead-onlyIdempotent
Self-assessment: how competent am I in given domains? Shows level, confidence, blind spots.
| Name | Required | Description | Default |
|---|---|---|---|
| concepts | No | Domains/concepts to assess competency for | |
| full_report | No | Return full self-model report |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and non-destructive behavior. The description adds that it shows level, confidence, and blind spots, which is useful but does not detail whether it only reads existing state or computes new assessments. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely short and front-loaded. The core purpose appears in the first sentence, and the additional output details are concise. No redundant or irrelevant wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only introspection tool with no output schema and annotations covering safety, the description is essentially complete. It explains what the tool does and what it returns at a high level, though it could mention the absence of side effects more explicitly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema descriptions cover both concepts and full_report, so coverage is complete. The tool description adds only minimal meaning to 'concepts' via 'given domains' and nothing extra about full_report. This meets the baseline for schema-covered parameters but does not enrich them.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the action (self-assessment) and the target (competency in given domains), plus the kind of output (level, confidence, blind spots). It does not explicitly name a resource like 'self model', but the intent is unambiguous enough.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus alternatives such as self_reflect, self_insight, or memory_recall. It lacks context for selecting this self-assessment over related introspection tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_skill_getCRead-onlyIdempotent
Find skills matching a trigger. Skills are learned procedures — HOW to do things.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Get skill by exact name | |
| trigger | No | Natural language trigger to match | |
| list_all | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint and idempotentHint. The description is consistent with these but adds no further behavioral context (e.g., return format, behavior on no match, or fuzzy matching semantics).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, consisting of two short sentences. It defines the key concept (skills) without unnecessary verbosity, though it could benefit from a structured parameter breakdown.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description omits critical context: it does not specify what the tool returns (e.g., a list of skill names or full objects), how parameters interact, or when to prefer this over siblings like memory_skill_update or memory_search_by_tag. With no output schema, this incompleteness is especially problematic.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description only clarifies the 'trigger' parameter by mentioning 'matching a trigger.' It does not explain 'name' (e.g., exact match) or 'list_all' (e.g., return all skills), which is a significant gap given the schema lacks any parameter descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the tool 'Find skills matching a trigger' and defines skills as 'learned procedures — HOW to do things.' This gives a clear verb and resource, though it doesn't explicitly distinguish it from sibling tools like memory_get or memory_search_by_tag.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives, nor on how the parameters (name vs trigger vs list_all) should be chosen. The description implies trigger-based search but leaves parameter semantics entirely implicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_skill_updateA
Record skill usage or refine a skill. Updates success rate and metrics.
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | ||
| success | Yes | Was the skill application successful? | |
| skill_id | Yes | Skill ID | |
| new_steps | No | Additional steps to add | |
| new_anti_pattern | No | Anti-pattern learned from failure |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description explicitly states that it 'Updates success rate and metrics', making the side effect transparent. Annotations reinforce non-readonly/non-destructive behavior, but details on whether updates append or overwrite are not disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise—two sentences—with no redundancy. It directly states the action and effect without extraneous detail, fitting the tool's simple scope.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's straightforward nature and no output schema, the description covers essential aspects: what it does and what it affects. The requirement of skill_id and success is implicit. It does not discuss edge cases or relationships with other skill-related tools, but this is not critical for a basic update operation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Of the 5 parameters, 4 have descriptive text (skill_id, success, new_steps, new_anti_pattern) with clear meanings. The 'notes' parameter lacks a description but appears self-explanatory. 80% coverage with clear field names supports effective use.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Record skill usage or refine a skill. Updates success rate and metrics.' It distinguishes itself from generic memory save/update tools by explicitly targeting skill usage and metrics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage (when recording or refining a skill) but does not explicitly state when to prefer this over siblings like memory_save or memory_update. It lacks explicit when-not-to-use guidance, though the phrase 'skill usage' provides some context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_statsARead-onlyIdempotent
Memory statistics with health metrics: sessions, knowledge by type/project, retention zones (active/archived/consolidated), stale records, storage size, config.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and destructiveHint=false, and the description does not contradict them. However, the description adds no additional behavioral disclosure beyond what annotations provide, so the bar is lower but still met.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that conveys the tool's purpose and enumerates the output categories without unnecessary verbosity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description lists the main output categories, giving a good sense of what the tool returns. Since there is no output schema, this enumeration partially compensates, but it lacks details about the output format or structure.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes no parameters, so the baseline score of 4 applies. There is nothing to document, and the description does not need to explain parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the tool provides memory statistics and health metrics, listing the specific categories (sessions, knowledge by type/project, retention zones, stale records, storage size, config), which distinguishes it from retrieval or modification tools in the sibling list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Does not explicitly state when to use this tool versus other stats-oriented siblings like memory_graph_stats or memory_perf_report. The purpose is implied but lacks explicit conditions or alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_temporal_queryBRead-onlyIdempotent
v11.0 W1-C: deterministic temporal reasoning — Allen interval relations, duration arithmetic (days/weeks/months/years), and natural-language date normalization (en + ru). Pass op=relation|duration_between|normalize.
| Name | Required | Description | Default |
|---|---|---|---|
| a | No | ISO datetime — duration_between | |
| b | No | ||
| op | Yes | ||
| lang | No | auto | |
| a_end | No | ||
| b_end | No | ||
| anchor | No | ISO datetime anchor for relative phrases | |
| phrase | No | Natural-language date — normalize | |
| a_start | No | ISO datetime — relation only | |
| b_start | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already state readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds that the tool performs deterministic temporal reasoning and which operations are available, giving useful behavioral context without contradicting the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense sentence and mostly efficient. The leading version string 'v11.0 W1-C' adds noise and could be removed, but the core capabilities are conveyed without excessive verbosity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and the description does not describe the return format or provide examples. Given the parameter complexity and the lack of an output schema, the description is not self-sufficient for an agent to confidently invoke the tool and interpret results.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 40%, and several parameters (b, b_end, b_start, lang, op) have no or minimal description. The description mentions op modes but does not map each mode to its required parameters, leaving significant ambiguity about how to populate the interval fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the resource as deterministic temporal reasoning and enumerates the three operation modes (relation, duration_between, normalize), giving a clear sense of the tool's purpose. However, it does not explicitly contrast with sibling tools like memory_timeline, so differentiation is mostly implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It instructs the caller to pass op=relation|duration_between|normalize, which is a direct usage cue. But it does not explain when each operation should be chosen over alternatives such as memory_timeline or other memory tools, leaving selection partly to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_timelineBRead-onlyIdempotent
Browse session history. sessions_ago=N for 'N sessions ago', session_number=1 for first session, date_from/date_to for date ranges.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | No | ||
| date_to | No | YYYY-MM-DD | |
| project | No | ||
| date_from | No | YYYY-MM-DD | |
| sessions_ago | No | ||
| session_number | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description is consistent with the annotations (readOnlyHint=true, destructiveHint=false, idempotentHint=true) and does not contradict them. It adds minimal behavioral detail beyond stating it 'browses' session history, but given the simple read-only nature implied by annotations, the description provides adequate transparency without contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, tightly packed sentence with no redundancy. It clearly conveys the core purpose and the meaning of several parameters without any fluff, making it highly concise and well-structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description provides the essential information for basic usage: what the tool does and how to specify certain filtering parameters. However, it omits details about the output format, how to interpret results, and when to choose this tool over similar ones. Given the absence of an output schema, the description is minimally adequate but leaves some gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 29% (2 out of 7 parameters have descriptions). The description text explains the meaning of three parameters (sessions_ago, session_number, date_from/date_to) but leaves limit, query, and project unexplained. Since schema coverage is low and the description compensates only partially, the semantics are insufficient for full understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose as browsing session history, with a specific verb ('browse') and resource ('session history'). It also explains the meaning of key parameters, which helps clarify the intended use. However, it does not explicitly differentiate itself from sibling tools like memory_history or kg_timeline, so it lacks explicit sibling distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides specific guidance on how to use certain parameters (sessions_ago, session_number, date_from/date_to), which is useful. However, it does not mention when to use this tool over alternatives, nor does it provide any prerequisites or context about the expected input format or output. Thus, usage guidance is partial.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_updateADestructive
Update existing knowledge. Replaces the record id, or the best match for find (within project when given): supersedes it and creates a new version.
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | Exact record to replace; takes precedence over find | |
| find | No | Search query to find the old knowledge (when no id) | |
| reason | No | Why updating | |
| project | No | Limit the find search to this project | |
| new_content | Yes | New content to replace with |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=false. The description adds meaningful context beyond that: the operation is non-destructive in the sense that it supersedes and versions rather than erasing, which an agent needs to understand before calling a 'destructive' tool. It does not cover auth requirements or whether old versions remain retrievable.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence with the core mechanism front-loaded and zero filler; everything after the first clause sharpens selection between id and find.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 5 parameters, full schema coverage, and safety annotations present, the definition covers what an agent needs to invoke it. The only gap is what happens to the superseded record (retrievable via memory_history?), which the description leaves implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description restates the id-over-find precedence, which the schema already documents per-parameter, and adds no format or constraint detail beyond that.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (update) plus resource (existing knowledge) and explains the mechanism: replaces a record and supersedes it with a new version. This clearly distinguishes it from memory_save (create) and memory_delete (remove) among the many memory_* siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives useful selection context for its own inputs (id takes precedence over find; project narrows find), but says nothing about when to choose this tool over memory_save, memory_delete, or memory_consolidate. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_warmupAIdempotent
v11.0: pre-load FastEmbed model and open the vector store, so the first save/search after process start doesn't pay model-load latency.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses the primary side effects: pre-loading a model and opening the vector store. It does not mention potential errors or whether prior setup is required, but the idempotentHint annotation aligns with the described warming behavior and there is no contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, compact sentence that conveys the action, the resource involved, and the reason for the action. Every word contributes meaning, and no unnecessary detail is included.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple, parameterless warm-up utility, the description provides the essential context: what is loaded, what is opened, and why it matters. It could mention failure modes or prerequisites, but the current level is complete enough for the tool's simplicity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are no parameters, so parameter-level semantics are trivially covered. The description adds no parameter information because none exists; a baseline score of 4 is appropriate for a zero-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: pre-load the FastEmbed model and open the vector store. It also explains the intended benefit—avoiding model-load latency on the first save/search—which is specific and actionable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool: after process start and before the first save/search. It does not explicitly mention alternatives, but with zero parameters and a focused warm-up purpose, the guidance is sufficiently clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
memory_wiki_generateAIdempotent
v10 — Render the per-project wiki digest (top decisions, active solutions, conventions, recent changes) as Markdown. Pass project to refresh one wiki, omit it to refresh all active projects. Files land in /wikis/.md and are deterministic (no LLM call).
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Project to refresh (omit for all) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses side effects: files land in <MEMORY_DIR>/wikis/<project>.md, and states determinism. The 'refresh' wording implies overwriting and writing to disk, consistent with readOnlyHint=false. IdempotentHint=true aligns with deterministic behavior. Some details about whether existing files are overwritten or other memory structures change are not explicit, but the core behavior is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences with no filler. All key information—purpose, parameter usage, output location, and determinism—is packed efficiently and logically ordered.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers purpose, parameter semantics, side effects, and determinism. It does not specify the exact Markdown structure or define 'active projects', but given the tool's simplicity and the presence of sibling tools for detailed queries, this is sufficient for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description ('Project to refresh (omit for all)') is clear, and the tool description reinforces it with concrete behavior ('refresh one wiki' vs 'all active projects'). This adds a bit of context beyond the schema, clarifying the optional parameter's effect on scope.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (render) and resource (per-project wiki digest), and enumerates the digest contents (top decisions, active solutions, conventions, recent changes). This clearly distinguishes the tool from memory-related siblings without needing to name them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit parameter usage: pass `project` to refresh one wiki, omit it to refresh all active projects. It also notes the tool is deterministic (no LLM call), helping an agent decide when to use it. It doesn't explicitly mention when-not alternatives, but the purpose is sufficiently unique.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
phase_transitionC
v8.0: advance a task to the next phase.
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | ||
| task_id | Yes | ||
| artifacts | No | ||
| new_phase | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The word 'advance' implies a state change, which aligns with readOnlyHint=false, but no additional behavioral traits are disclosed. It does not mention side effects, return values, or other outcomes beyond the transition itself.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, succinct sentence with no redundancy. It front-loads the verb and object, making the core operation immediately clear.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is extremely sparse. It omits details about phase definitions, the role of notes and artifacts, whether the transition is reversible, and what the output or result might be. Given the complexity of the sibling toolset, this is insufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
None of the parameters (task_id, new_phase, notes, artifacts) are described. Their names hint at their purpose, but there is no explicit semantic detail about expected values, formats, or how they affect the transition.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('advance') and the object ('a task to the next phase'), but lacks context about what a 'phase' is or how this action differs from other task-related tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like task_create or task_phases_list. The description gives no situational context or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
rule_set_phaseAIdempotent
Attach or remove a phase scope on a rule (v8.0 lazy rule loading). Tag-based: manages 'phase:' on the rule's tags. phase=null clears the phase tag (rule becomes core — applies to every phase). Valid phases: van, plan, creative, build, reflect, archive.
| Name | Required | Description | Default |
|---|---|---|---|
| phase | No | Phase name or null to clear. | |
| rule_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already indicate readOnly=false, idempotent=true, and destructive=false, and the description adds useful detail about null clearing the tag and making the rule core. It does not disclose potential side effects beyond tag modification, but transparency is adequate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and information-dense, covering purpose, mechanism, null semantics, and valid values in a single sentence. No unnecessary words or redundant structure are present.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description provides enough context to call the tool correctly: what it does, how tags are managed, what null means, and which phase values are valid. No output schema exists, but the outcome is clear from the described behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The phase parameter is well explained with valid enum values and null behavior, and rule_id is self-explanatory as an integer identifier for a rule. The description adds meaning beyond the schema, though rule_id itself lacks an explicit schema description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool attaches or removes a phase scope on a rule, identifies the tag-based mechanism ('phase:<X>'), and lists valid phases. It is immediately distinguishable from sibling tools by naming the specific rule-tag operation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains what the tool does and how null works, but it does not explicitly state when to prefer this over sibling phase-related tools, nor does it mention when not to use it. Usage guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_decisionA
v8.0: save a structured architectural decision (options + criteria matrix + rationale + discarded). Adds structured tag and a JSON blob in context. Use for Creative-phase outputs; plain type=decision memory_save still works.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | No | ||
| title | Yes | Short decision title | |
| options | Yes | Options considered: [{name, pros[], cons[], unknowns[]}, ...] | |
| project | No | ||
| selected | Yes | Chosen option name (must be in options) | |
| discarded | No | Option names rejected (subset of options - {selected}) | |
| rationale | Yes | Why this option was chosen | |
| criteria_matrix | Yes | criterion -> {option_name: rating 0-5} |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Description mentions adding a structured tag and JSON blob, going beyond annotations which only indicate non-read-only, non-destructive, and non-idempotent. It does not specify whether it creates new entries or updates existing ones, but the 'Adds' phrasing implies creation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences that are directly relevant, with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (8 params, nested objects), the description adequately covers the core purpose and usage context. It does not describe return values, but no output schema exists, so that is not a gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers 75% of parameters with descriptions. The tool description reiterates the main required components but does not clarify the missing 'project' or 'tags' parameters beyond what the schema lacks. No additional per-parameter semantics are provided.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the tool saves a structured architectural decision with specific components (options, criteria matrix, rationale, discarded). Also notes it adds a structured tag and JSON blob, distinguishing it from plain memory_save.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly directs usage for Creative-phase outputs and mentions that plain type=decision memory_save still works as an alternative, providing clear when-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_intentA
Persist one user prompt into the intents table (same source as the UserPromptSubmit hook). Use when programmatically seeding intents — the hook covers normal interactive usage. Dedupes same prompt within 5 min per session.
| Name | Required | Description | Default |
|---|---|---|---|
| prompt | Yes | User prompt text as submitted | |
| project | No | Project slug | |
| session_id | No | Session id (defaults to current MCP session) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses that the tool deduplicates the same prompt within 5 minutes per session, which is useful behavioral context not present in annotations. It also clarifies the relationship to the hook, though it does not detail side effects like return behavior or error cases.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two concise sentences that convey purpose, use case, and deduplication behavior without unnecessary detail. It is well-structured and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple insert-style tool with no output schema, the description covers the essential context: what is persisted, the relationship to the hook, when to use it, and deduplication behavior. No critical information is missing for an agent to decide to call it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema already describes all three parameters with full coverage, including the default behavior of session_id. The description adds only minimal additional parameter context, such as dedupe behavior involving prompt and session, so it stays at the baseline for schema-covered parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool persists a user prompt into the intents table and explicitly names the same source as the UserPromptSubmit hook. It distinguishes itself from normal interactive usage and from sibling tools like save_decision by focusing on programmatic intent seeding.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says to use this tool when programmatically seeding intents and notes that the hook covers normal interactive usage. This gives clear guidance on when to use the tool versus relying on the existing hook.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_intentsARead-onlyIdempotent
Substring search over user prompts (LIKE). Returns newest match first. Useful for 'what did I ask about X' without mining transcripts.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| query | Yes | Substring to match in prompt text | |
| project | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, non-destructive. Description adds behavioral detail: it uses LIKE substring matching and returns newest matches first. This goes beyond the annotations and helps the agent understand result ordering.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no redundant words. It front-loads the core functionality and immediately gives a use case.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers purpose, ordering, and a practical use case. Given the tool's simplicity and the presence of read-only annotations, this is sufficient for an agent to decide when to use it. It does not need to explain output format since there is no output schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only the query parameter has a description in the schema; limit and project are undocumented. The tool description does not add any explanation for limit or project, leaving these parameters ambiguous. With 33% schema coverage, the description should have compensated, but it does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States exactly what it does: substring search over user prompts using LIKE. Also specifies result ordering (newest first). The verb 'search' is specific and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a clear use case ('what did I ask about X') and contrasts with a manual alternative (mining transcripts). Does not explicitly name sibling tools that might be more suitable for other types of search, so the guidance is good but not fully explicit about exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
self_error_logA
Log an error/failure for pattern analysis. Call AUTOMATICALLY when: bash command fails, wrong assumption discovered, API returns error, config issue found, loop detected, or any mistake occurs. System detects patterns (3+ same category) and suggests insights.
| Name | Required | Description | Default |
|---|---|---|---|
| fix | No | How it was fixed (empty if unresolved) | |
| tags | No | ||
| context | No | What was being done when error occurred | |
| project | No | general | |
| category | Yes | Error category for pattern grouping | |
| severity | No | medium | |
| description | Yes | What went wrong: symptom, expectation vs reality |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description mentions pattern detection and insight suggestion, which implies side effects beyond simple logging. It does not explicitly state persistence or non-idempotency, but the annotations already cover these aspects appropriately.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, front-loading the core purpose and then listing triggers efficiently. No redundant information, and the structure is logical.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a logging tool with no output schema, the description adequately covers the purpose, triggers, and expected behavior. It could benefit from an example, but the essentials are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description does not add any parameter-specific guidance beyond what the schema already provides. With schema coverage at 57%, several parameters lack clarity, and the description fails to compensate for this gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: logging errors/failures for pattern analysis. It specifies the exact scenarios in which to call it, making its role unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit triggers for automatic invocation (bash failure, wrong assumption, API error, config issue, loop, or any mistake), leaving no ambiguity about when to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
self_insightA
Manage insights from error patterns (ExpeL-style). Actions: add (create, importance=2), upvote (+1), downvote (-1, auto-archive at 0), edit, list, promote (to rule when importance>=5 AND confidence>=0.8). Call 'add' when pattern detected. Call 'upvote' when insight confirmed again.
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | Insight ID (for upvote/downvote/edit/promote) | |
| tags | No | ||
| action | Yes | ||
| content | No | Insight text (for add/edit) | |
| context | No | ||
| project | No | general | |
| category | No | Error category (for add) | |
| source_error_ids | No | Error IDs that spawned this (for add) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses important side effects such as downvote auto-archiving at 0 and the promote threshold. Annotations are not contradicted, and the description adds useful behavioral context beyond the flags.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured, using a clear action list and parenthetical modifiers. Every sentence adds useful information without unnecessary fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the multi-action nature and 8 parameters, the description covers the core action semantics and thresholds well. It does not mention return values or output format, but the lack of an output schema reduces the impact of that omission.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is about 50%; the description adds meaning for action-specific behavior and thresholds but does not elaborate on parameters like tags, context, project, or source_error_ids beyond their schema descriptions. Some parameter semantics are left to inference.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly describes managing insights from error patterns and enumerates the concrete actions (add, upvote, downvote, edit, list, promote). This makes the tool's scope obvious and distinct from sibling memory/self tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit trigger conditions for adding and upvoting insights ('Call add when pattern detected', 'Call upvote when insight confirmed again'), giving clear contextual usage. It does not give when-not-to-use guidance for every action, but the main intent is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
self_patternsARead-onlyIdempotent
Analyze error patterns and self-improvement stats. Views: error_patterns (frequency, repeating 3+), insight_candidates (ready for promotion), rule_effectiveness (success rates, stale rules), improvement_trend (weekly errors), full_report (all). Call periodically to track improvement.
| Name | Required | Description | Default |
|---|---|---|---|
| days | No | ||
| view | No | full_report | |
| project | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only and idempotent behavior. The description adds value by detailing what data is analyzed (error patterns, insight candidates, rule effectiveness, trends) and the periodic cadence. No contradiction; the description enriches the annotation-provided safety profile.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero fluff. The main purpose is front-loaded, and the view list is compactly presented. Every word contributes to the tool's functionality.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the core functionality and views well, but omits details about the 'project' parameter and does not describe the output format. Given the tool has 3 optional parameters and no output schema, these gaps make it adequate but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema coverage, the description must explain parameters. It fully explains the 'view' enum values, but does not explain 'days' (only hints via 'weekly errors') or 'project' at all. It adds meaning for the primary parameter but leaves two parameters under-documented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states a specific verb ('Analyze') and resource ('error patterns and self-improvement stats'), then enumerates five distinct views that define exactly what analyses are available. This clearly separates it from sibling tools like self_error_log (logging) and self_reflect (general reflection).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description suggests periodic invocation ('Call periodically to track improvement'), implying a monitoring use case. It does not explicitly name alternatives or exclusion criteria, but the view list implicitly differentiates from other self-* tools. This is clear context without formal when-not guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
self_reflectA
Save a verbal self-reflection (Reflexion pattern). Call after completing a task or encountering difficulty. NOT for errors (use self_error_log). For meta-observations about strategy, approach effectiveness, process improvements.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | No | ||
| outcome | No | success | |
| project | No | general | |
| reflection | Yes | What went well, what to improve, what to do differently | |
| task_summary | Yes | Brief description of what was done |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description says 'Save' and references 'Reflexion pattern' but goes no deeper into side effects — where the reflection is stored, whether it appends to memory_timeline, or whether repeated calls create duplicate entries. Annotations provide only readOnlyHint/destructiveHint/idempotentHint flags with no additional guidance, so the description carries most of the burden and only partially discharges it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with verb+resource ('Save a verbal self-reflection'). Negative guidance and content scope are packed efficiently without redundancy; every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Covers the core essentials — what to save, when to call, what content belongs, what to avoid. Does not state where the reflection is stored (e.g., retrievable via memory_get or visible in memory_timeline) nor clarify the role of outcome/project/tags, leaving minor gaps for an agent operating within the rich sibling tool set. With no output schema, nothing about return values is owed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema covers only 2 of 5 parameters (reflection, task_summary); tags, outcome, and project have no schema description. The tool description's 'NOT for errors...' and 'meta-observations about strategy...' lines clarify the intended content of reflection, but the uncovered parameters remain unexplained, so the description only partially compensates for the 40% coverage gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States the verb ('Save'), the resource ('verbal self-reflection'), the pattern ('Reflexion pattern'), and the invocation trigger ('after completing a task or encountering difficulty'). The 'NOT for errors (use self_error_log)' line explicitly differentiates it from the closest sibling, making purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Names the explicit alternative (self_error_log) and the condition that selects it ('NOT for errors'), plus the content scope ('meta-observations about strategy, approach effectiveness, process improvements') and trigger ('after completing a task or encountering difficulty'). The agent can decide without guessing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
self_rulesB
Manage behavioral rules (SOUL). Rules are promoted insights that shape agent behavior. Actions: list, fire (record relevance), rate (success=true/false), suspend, activate, retire, add_manual. Auto-suspend: success_rate < 0.2 after 10+ fires.
| Name | Required | Description | Default |
|---|---|---|---|
| id | No | Rule ID (for fire/rate/suspend/activate/retire) | |
| tags | No | ||
| scope | No | global | project:<name> | category:<name> | global |
| action | Yes | ||
| content | No | Rule text (for add_manual) | |
| project | No | general | |
| success | No | For rate: was rule helpful? | |
| category | No | Category (for add_manual) | |
| priority | No | 1-10 |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses key behavioral details beyond the annotations, such as the auto-suspend rule (success_rate < 0.2 after 10+ fires) and clarifies the side effects of actions like 'fire (record relevance)' and 'rate (success=true/false)'. It does not mention potential destructive consequences of 'suspend' or 'retire', but the annotations already set destructiveHint to false, so no conflict.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is brief and information-dense. It front-loads the core purpose ('Manage behavioral rules') and then lists actions and the auto-suspend rule without unnecessary verbiage. Every sentence adds value.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (multiple actions, a project/category scope, auto-suspend logic), the description covers the main behaviors but omits details about expected outputs, error conditions, or how scope/project/category interact. It is sufficient for a basic understanding but leaves some operational gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description does not add meaning to the parameters beyond what the schema already provides. Several parameters (tags, project) lack schema descriptions, and the description does not compensate by explaining them. It only describes actions, not the fields needed to perform them, so parameter semantics are not enhanced.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool manages behavioral rules (SOUL) and enumerates the supported actions (list, fire, rate, suspend, activate, retire, add_manual), giving a clear sense of purpose. However, it does not explicitly differentiate this from similar sibling tools like self_rules_context or rule_set_phase, so it is slightly less than a perfect 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description lists actions but provides no guidance on when to use this tool versus alternatives (e.g., when to use self_rules_context or rule_set_phase). There is no explicit 'use this when' or 'instead of' guidance, leaving the selection largely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
self_rules_contextARead-onlyIdempotent
Get active behavioral rules for current session. Call at SESSION START to load rules. Returns rules filtered by project and scope. v8.0: pass phase to lazy-load rules relevant to current task phase — core rules (no phase tag) + rules tagged phase:. Cuts prompt tokens ~70%. After task completion, rate rules: self_rules(action='rate', id=X, success=true/false).
| Name | Required | Description | Default |
|---|---|---|---|
| phase | No | Optional: lazy-load only rules relevant to this phase (core + phase-specific). Omit to get all rules. | |
| project | No | general | |
| categories | No | Error categories relevant to current task |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare readOnly, idempotent, and non-destructive behavior, and the description is fully consistent with these—it describes only retrieving rules. The description does add some output-related context (filtered by project and scope) but does not go beyond annotation coverage in terms of side effects, auth, or rate limits. Since the annotations carry the main safety information, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured and front-loaded, starting with the core purpose and then providing usage and feature details. Each sentence adds relevant information (purpose, usage timing, output behavior, lazy-loading feature, and follow-up action). While slightly verbose with the v8.0 note and rating instructions, every sentence earns its place, so it merits a 4.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description must explain what the function returns. It only states 'Returns rules filtered by project and scope,' which is vague—it does not specify the format, structure, or whether the output is a list, string, or JSON. This incomplete information about the return value leaves ambiguity for the agent, making the description insufficiently complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 67% (phase and categories have descriptions, but project does not). The description adds meaningful context for the phase parameter by explaining lazy-loading and the tagging scheme, but it only vaguely references project and scope without defining them. The coverage is below the high threshold, so the description should compensate, but it only partially does so, leading to a 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's primary purpose: 'Get active behavioral rules for current session.' It uses a specific verb (Get) and resource (active behavioral rules), and immediately identifies the intended usage context (session start). This makes the tool's function unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'Call at SESSION START to load rules,' providing a clear when-to-use instruction. It also explains the lazy-loading behavior with the phase parameter and suggests rating rules after completion. However, it does not mention any alternatives or when not to use this tool compared to sibling tools, so it's not a perfect 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_endA
End-of-session capture: summary + highlights + pitfalls + next_steps so the next session can resume cleanly. Set auto_compress=true to have the LLM generate the missing summary/next_steps/pitfalls from stored session artifacts (or from an optional transcript).
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | general | |
| summary | No | ||
| pitfalls | No | ||
| highlights | No | ||
| next_steps | No | ||
| session_id | Yes | ||
| transcript | No | ||
| auto_compress | No | ||
| open_questions | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Discloses the auto_compress behavior, explaining that setting it to true generates missing fields from stored artifacts or transcript. However, it does not mention side effects such as overwriting existing session data, and annotations provide no safety hints (all false), so the description carries more burden but still leaves some ambiguity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is succinct, consisting of two sentences with no redundancy. It front-loads the core purpose and then adds the key behavioral detail about auto_compress, maintaining efficiency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Provides enough to understand the core action and the auto_compress feature, but lacks crucial context: it does not state whether any fields are required when auto_compress is false, nor how this tool fits relative to sibling session/memory tools. This could lead to incorrect usage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description explicitly mentions several parameters (summary, highlights, pitfalls, next_steps, transcript, auto_compress) and explains auto_compress. However, it does not clarify the purpose or usage of required session_id, project, or open_questions. Given low schema coverage (0%), it partially compensates but leaves gaps.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states it is an 'End-of-session capture' with specific fields (summary, highlights, pitfalls, next_steps), giving a clear verb and resource. It implies a specific usage context but does not explicitly distinguish from sibling tools like memory_save or session_init.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'End-of-session capture' and 'so the next session can resume cleanly' provide a temporal context, but the description does not explicitly mention when to use this over alternatives or when not to use it. It lacks direct guidance against using other memory tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
session_initC
At session start: return the most recent unconsumed end-of-session summary with highlights / pitfalls / next_steps.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | general | |
| mark_consumed | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description does not disclose the side effect of marking the summary as consumed when mark_consumed is true, nor does it mention any other effects like authentication or rate limits. The annotations indicate non-destructive, but the state change is not explained.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, one sentence, with no fluff. It states the essence directly. However, it could benefit from a bit more detail on the parameters, but conciseness itself is acceptable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, and the description does not describe the expected return format or additional context. The agent is left with insufficient information about how to interpret the result or handle the parameters. Minor context is provided by the mention of highlights/pitfalls/next_steps, but overall it's incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description does not explain the 'project' or 'mark_consumed' parameters, leaving the agent to infer their meaning from the schema alone, which has no descriptions. Since schema coverage is 0%, the description should have compensated but did not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: to return the most recent unconsumed end-of-session summary, with a specific list of contents. However, it doesn't explicitly differentiate from similar tools like session_end or memory_recall, but the name and context (session start) provide reasonable clarity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides a temporal trigger ('At session start') but offers no guidance on when to use this tool versus alternatives such as memory_recall or session_end. No conditions, prerequisites, or exclusions are mentioned.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
task_createC
v8.0: start a task in van phase (auto-classifies level if missing).
| Name | Required | Description | Default |
|---|---|---|---|
| level | No | ||
| task_id | Yes | ||
| description | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses the auto-classification behavior when the level is missing, which goes beyond the annotation flags (all false). However, it does not mention side effects, permissions, error modes, or what happens in the 'van' phase—leaving significant behavioral aspects unexplained.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence and technically concise, but the 'v8.0:' version prefix adds noise without value. The cryptic 'van' term and lack of elaboration make it feel under-specified rather than efficiently structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is too terse to provide sufficient context. It does not explain what the 'van' phase means, what 'level' classification entails, or what the expected outcome is. With no output schema and minimal context, the tool is not adequately described for correct usage.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has three parameters (task_id, description, level) with no descriptions. The description only references 'level' in relation to auto-classification, and does not explain what task_id or description represent. With 0% parameter coverage, the description only minimally compensates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('start a task') and a condition ('auto-classifies level if missing'), which makes the core purpose clear. However, the phrase 'van phase' is cryptic and may confuse, though it is still distinguishable from sibling tools like classify_task or phase_transition.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. The mention of auto-classification gives an implicit hint, but there is no explicit direction or context for choosing this over similar sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
task_phases_listBRead-onlyIdempotent
v8.0: list all phases of a task in chronological order.
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, so the safety profile is covered. The description adds the chronological ordering behavior, which is useful, but doesn't disclose other details like whether all phases (including completed) are returned or if there's any pagination.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One concise sentence with the key action and ordering up front. The 'v8.0:' version prefix adds minor noise but doesn't significantly detract.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only list operation with one parameter and annotations covering safety, the description is mostly adequate. However, it omits any return format details (e.g., phase objects with status/timestamps) and provides no usage context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explicitly explain the task_id parameter. It only implies it through 'of a task', but doesn't state it's required or specify format. The parameter is simple and self-explanatory, but the description fails to formally map it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('list'), resource ('phases of a task'), and adds ordering ('chronological'). This clearly distinguishes it from task creation/transition tools, though it doesn't explicitly contrast with a sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives. It doesn't mention it's a read-only inspection query or that it should be used to track task progress rather than modify phases.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workflow_learnB
Record a learned workflow (named sequence of steps) for future reuse.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | ||
| steps | Yes | ||
| context | No | ||
| project | No | general | |
| description | No | ||
| trigger_pattern | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The word 'Record' clearly indicates a state-changing mutation, consistent with the annotations (readOnlyHint=false). However, the description does not go beyond that to explain persistence, overwrite behavior, idempotency, or any side effects beyond the basic act of recording.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, tight sentence with no redundant or misleading wording. It efficiently conveys the core purpose and even defines the key concept of a workflow.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has six parameters, nested objects, and no output schema, but the description only covers the two most obvious parameters. It does not mention expected output, errors, prerequisites, or how this relates to sibling workflow and memory tools. A more complete description would address these gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description clarifies that 'name' is the workflow's name and 'steps' are the sequence, but it leaves all other parameters (context, project, description, trigger_pattern) unexplained. With zero schema descriptions, this is a significant gap; the context object and trigger_pattern semantics are especially unclear.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Record'), the resource ('learned workflow'), and defines it as a 'named sequence of steps' for future reuse. This makes the tool's purpose unambiguous and distinct from typical memory-saving tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a use case ('for future reuse') but does not explicitly state when to prefer this tool over siblings like workflow_track, workflow_predict, or memory_save. No guidance is given on when not to use it or how it differs from alternative recording/tracking tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workflow_predictARead-onlyIdempotent
Predict outcome (success probability, avg duration) for a workflow by id OR by trigger keyword. Uses Laplace-smoothed success rate.
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | ||
| trigger | No | ||
| workflow_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only and idempotent behavior. The description adds meaningful transparency by revealing the Laplace-smoothed success rate and the nature of the prediction (probability and duration), which is not captured in annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Extremely concise and well-structured. Two sentences contain all essential information without any filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with three parameters and no output schema, the description provides enough context to understand the core functionality and invocation modes. It lacks details on parameter requirements (e.g., project) and output format, but these are partially compensated by the simplicity of the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
No parameter descriptions exist in the schema. The description hints that workflow_id and trigger are alternative identifiers, but project is completely unexplained. Coverage is insufficient to fully understand parameter roles and constraints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Clearly states the action ('Predict outcome'), the resource ('workflow'), the specific outputs ('success probability, avg duration'), and the two identification methods ('by id OR by trigger keyword'). No ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides limited guidance on when to use this tool versus alternatives. It mentions two identification modes but does not explain when to prefer one or when to use this tool over similar siblings like workflow_learn or workflow_track.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
workflow_trackA
Record a workflow execution outcome. Outcome ∈ {success|failure|partial|aborted}. Aggregates update automatically.
| Name | Required | Description | Default |
|---|---|---|---|
| notes | No | ||
| outcome | Yes | ||
| duration_ms | No | ||
| workflow_id | Yes | ||
| error_details | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description mentions that 'Aggregates update automatically,' which hints at side effects. It does not contradict the annotations (readOnlyHint false, etc.), but it could be more explicit about whether the operation is append-only or modifies existing records.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is extremely concise and to the point, using a compact notation for the outcome enum. Every word adds value, and there is no redundant information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple logging tool, the description provides the essential purpose and a key side effect. However, it lacks details on parameter semantics and any caveats (e.g., whether the workflow must already exist). This leaves some gaps for an agent unfamiliar with the tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description explains the 'outcome' parameter partially by listing its enum values, but provides no explanation for 'workflow_id', 'notes', 'duration_ms', or 'error_details'. With 5 parameters and 0% schema description coverage, the description does not adequately clarify the meaning or usage of most parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action ('Record') and the object ('workflow execution outcome'), and enumerates the valid outcome values. It is immediately obvious what the tool does and how it differs from sibling workflow tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the core action but does not explicitly state when to use this tool versus alternatives like workflow_learn or workflow_predict. The context signals and sibling names imply it is for logging outcomes, but explicit guidance is absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v14.7.0- Changed
kg_add_fact1 field changed- added
Input schema / properties / valid_fromAdded value: +{ + "description": "ISO 8601 time the fact became true; default now. Back-dated facts close and are closed by their neighbours.", + "type": "string" +}
- Changed
kg_invalidate_fact1 field changed- added
Input schema / properties / atAdded value: +{ + "description": "ISO 8601 time the fact stopped being true; default now", + "type": "string" +}
- Changed
memory_delete1 field changed- added
Input schema / properties / hardAdded value: +{ + "default": false, + "description": "Erase permanently instead of hiding; cannot be undone", + "type": "boolean" +}
- Changed
memory_recall1 field changed- added
Input schema / properties / fill_budgetAdded value: +{ + "default": false, + "description": "Context mode: search deeper (up to 100 hits) and keep whole hits in rank order until context_max_chars is used, instead of excerpting the top `limit` hits to fit. Neighbours default to 0 here; pass `neighbors` to keep each hit with the records around it.", + "type": "boolean" +}
- Added
memory_report - Changed
memory_update4 fields changed- changed
Input schema / properties / find / descriptionPrevious value: -"Search query to find the old knowledge"New value: +"Search query to find the old knowledge (when no id)" - added
Input schema / properties / idAdded value: +{ + "description": "Exact record to replace; takes precedence over find", + "type": "integer" +} - added
Input schema / properties / projectAdded value: +{ + "description": "Limit the find search to this project", + "type": "string" +} - changed
Input schema / requiredPrevious value: -[ - "find", - "new_content" -]New value: +[ + "new_content" +]
2 tool updates
v14.5.0- Changed
memory_save1 field changed- added
Input schema / properties / supersedeAdded value: +{ + "default": false, + "description": "Retire active records of the same project and type that this one gives a new value for: same opening words, different trailing value (\"X's citizenship is Argentina\" -> \"... is Armenia\"). Use for single-valued facts only; \"likes jazz\" would retire \"likes rock\". The retired ids are returned as `superseded`.", + "type": "boolean" +}
- Changed
memory_save_fast1 field changed- added
Input schema / properties / supersedeAdded value: +{ + "default": false, + "description": "Same as memory_save.supersede: retire records this one gives a new value for.", + "type": "boolean" +}
6 tool updates
v14.0.0- Added
memory_answer - Added
memory_index_passages - Changed
memory_recall6 fields changed- added
Input schema / properties / context_max_charsAdded value: +{ + "default": 24000, + "description": "Context character budget; evidence mode uses the same value as a stricter UTF-8 byte budget.", + "minimum": 512, + "type": "integer" +} - added
Input schema / properties / evidence_followupAdded value: +{ + "default": true, + "description": "Evidence mode: allow one search for an explicitly supplied missing_relation.", + "type": "boolean" +} - added
Input schema / properties / missing_relationAdded value: +{ + "additionalProperties": false, + "description": "Evidence mode: missing relation; subject must occur in the original question.", + "properties": { + "relation": { + "type": "string" + }, + "subject": { + "type": "string" + }, + "time": { + "type": "string" + } + }, + "required": [ + "subject", + "relation" + ], + "type": "object" +} - changed
Input schema / properties / mode / descriptionPrevious value: -"Progressive-disclosure mode: 'search' (default) = normal results, 'index' = ultra-compact metadata only (id+title+score+type+project+created_at, ~40-60 tok/hit, no cognitive expansion, use memory_get(ids=...) to fetch full content), 'timeline' = top-K hits expanded with ±neighbors from same session (chronological)"New value: +"Progressive-disclosure mode: 'search' (default) = normal results, 'index' = ultra-compact metadata only (id+title+score+type+project+created_at, ~40-60 tok/hit, no cognitive expansion, use memory_get(ids=...) to fetch full content), 'timeline' = chronological compact view; 'context' = source excerpts; 'evidence' = indexed passage search with bounded follow-up retrieval" - changed
Input schema / properties / mode / enumPrevious value: -[ - "search", - "index", - "timeline" -]New value: +[ + "search", + "index", + "timeline", + "context", + "evidence" +] - changed
Input schema / properties / neighbors / descriptionPrevious value: -"Timeline mode only: how many records before/after each hit to include."New value: +"Timeline/context modes: records before/after each hit; context accepts 0–3."
- Changed
memory_recall_iterative6 fields changed- changed
Input schema / properties / k_per_iter / defaultPrevious value: -5New value: +10 - added
Input schema / properties / k_per_iter / maximumAdded value: +50 - added
Input schema / properties / k_per_iter / minimumAdded value: +1 - changed
Input schema / properties / llm_model / defaultPrevious value: -"haiku"New value: +"configured" - added
Input schema / properties / max_iters / maximumAdded value: +12 - added
Input schema / properties / max_iters / minimumAdded value: +1
- Changed
memory_save1 field changed- added
Input schema / properties / source_formatAdded value: +{ + "default": "auto", + "description": "Conversation preserves dialogue structure and bypasses automatic CLI filters.", + "enum": [ + "auto", + "conversation" + ], + "type": "string" +}
- Changed
memory_save_fast1 field changed- added
Input schema / properties / source_formatAdded value: +{ + "default": "auto", + "enum": [ + "auto", + "conversation" + ], + "type": "string" +}
74 tool updates
v0.1.0- First observed
analogize - First observed
benchmark - First observed
classify_task - First observed
file_context - First observed
ingest_codebase - First observed
kg_add_fact - First observed
kg_at - First observed
kg_invalidate_fact - First observed
kg_timeline - First observed
learn_error - First observed
list_intents - First observed
memory_associate - First observed
memory_concepts - First observed
memory_consolidate - First observed
memory_consolidate_status - First observed
memory_context_build - First observed
memory_delete - First observed
memory_entity_resolve - First observed
memory_episode_recall - First observed
memory_episode_save - First observed
memory_eval_contradictions - First observed
memory_eval_entity_consistency - First observed
memory_eval_locomo - First observed
memory_eval_long_context - First observed
memory_eval_recall - First observed
memory_eval_temporal - First observed
memory_explain_search - First observed
memory_export - First observed
memory_extract_session - First observed
memory_forget - First observed
memory_get - First observed
memory_graph - First observed
memory_graph_index - First observed
memory_graph_stats - First observed
memory_history - First observed
memory_observe - First observed
memory_perf_report - First observed
memory_rebuild_embeddings - First observed
memory_rebuild_fts - First observed
memory_recall - First observed
memory_recall_iterative - First observed
memory_reflect_now - First observed
memory_relate - First observed
memory_save - First observed
memory_save_fast - First observed
memory_search_by_tag - First observed
memory_search_fast - First observed
memory_self_assess - First observed
memory_skill_get - First observed
memory_skill_update - First observed
memory_stats - First observed
memory_temporal_query - First observed
memory_timeline - First observed
memory_update - First observed
memory_warmup - First observed
memory_wiki_generate - First observed
phase_transition - First observed
rule_set_phase - First observed
save_decision - First observed
save_intent - First observed
search_intents - First observed
self_error_log - First observed
self_insight - First observed
self_patterns - First observed
self_reflect - First observed
self_rules - First observed
self_rules_context - First observed
session_end - First observed
session_init - First observed
task_create - First observed
task_phases_list - First observed
workflow_learn - First observed
workflow_predict - First observed
workflow_track
TDQS
Scored across 77 tools
Many retrieval tools overlap heavily (memory_recall, memory_search_fast, memory_associate, memory_recall_iterative, memory_context_build, memory_answer, memory_temporal_query) and save variants blur (memory_save vs memory_save_fast vs memory_observe vs memory_episode_save). Descriptions add nuance, but an agent can easily misselect among these clusters.
Nearly all names use snake_case and many carry domain prefixes (memory_, kg_, workflow_, self_), but conventions are mixed: prefix_noun (memory_save), verb_noun (save_intent, classify_task), and bare verb/noun (analogize, benchmark, file_context). Still readable, but not a predictable pattern throughout.
77 tools is far beyond a well-scoped memory server; many are eval, debug, perf, or internal admin utilities that could be consolidated. This creates severe surface bloat and high selection cost.
Core lifecycle is well covered: save/update/delete/recall/history/export/forget/consolidate plus graph, episodes, skills, rules, and sessions. Minor gaps exist (e.g., no explicit import counterpart to export, some tools are eval/admin rather than user-facing memory operations), but core workflows have no dead ends.
Maintenance
Related MCP Connectors
Persistent memory for AI agents across Claude, ChatGPT and any MCP client.
Persistent AI memory shared across Claude, ChatGPT, coding agents, and compatible MCP clients.
Persistent memory for Claude Code and Cursor. Stop re-explaining your project every session.
- mcpOAuthai.butlerbrain
Persistent memory for AI assistants. Save once; recall from Claude, ChatGPT, or any MCP client.
Related MCP Servers
- AlicenseNot gradedqualityAmaintenanceEnables AI coding agents to maintain persistent, cross-session memory of codebase architecture, naming conventions, and decisions through MCP tools. Eliminates repetitive project re-explanation by automatically injecting stored context into every session with local-first SQLite storage and optional team sharing capabilities.4MIT
- AlicenseAqualityDmaintenanceGives Claude Code, Claude Desktop, Cursor, VS Code Copilot, and other MCP-compatible tools persistent memory.1853 npm1MIT
- AlicenseNot gradedqualityBmaintenancePersistent memory for AI coding agents that stores and recalls preferences, decisions, and conventions via semantic similarity, with zero cloud dependencies and plug-and-play MCP integration for Claude Code.Apache 2.0
- AlicenseAqualityAmaintenancePersistent memory for AI coding agents, storing decisions, bug fixes, conventions, and discoveries in a local SQLite database and automatically recalling them when relevant. Works with Claude Code, Codex, Cursor, Gemini CLI, and other MCP-compatible agents.222MIT