@contextq/mcp
OfficialClick on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@@contextq/mcpsearch my ContextQ memories for notes on the auth migration"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
@contextq/mcp
MCP server for ContextQ -- exposes the ContextQ knowledge-management API (89 tools: save, search, ingest, goal graphs, agent sessions, relays, and more) as Model Context Protocol tools. A curated ~24-tool default set loads at connection to keep the token cost of tools/list low; the rest load on demand or via CONTEXT_MCP_TOOL_PROFILE=full -- see below.
npx -y @contextq/mcpClient configuration
Two environment variables are required in every client:
Variable | Description |
| Base URL of your ContextQ server (e.g. |
| API key sent as |
Optional:
Variable | Description |
|
|
Setup paths: Claude Code and Claude Desktop have automated setup via the contextq init CLI command. Cursor, Windsurf, and Cline require manual config file editing — see docs/mcp-setup.md for the full reference.
Claude Desktop
Add to claude_desktop_config.json:
{
"mcpServers": {
"contextq": {
"command": "npx",
"args": ["-y", "@contextq/mcp"],
"env": {
"CONTEXT_API_KEY": "sk_live_YOUR_API_KEY",
"CONTEXT_API_URL": "https://ctx.example.com"
}
}
}
}Claude Code
claude mcp add contextq \
-e CONTEXT_API_KEY=sk_live_YOUR_API_KEY \
-e CONTEXT_API_URL=https://ctx.example.com \
-- npx -y @contextq/mcpCursor
Add to your Cursor MCP config (.cursor/mcp.json or Settings > MCP):
{
"mcpServers": {
"contextq": {
"command": "npx",
"args": ["-y", "@contextq/mcp"],
"env": {
"CONTEXT_API_KEY": "sk_live_YOUR_API_KEY",
"CONTEXT_API_URL": "https://ctx.example.com"
}
}
}
}Windsurf
STATUS (2026-06-02): Windsurf was rebranded as Devin Desktop and Cascade was end-of-lifed (2026-07-01). If you have an existing Windsurf install, the configuration below still applies, but new installations should use Devin Desktop instead. Devin Desktop uses the same MCP config format under .devin/mcp.json.
Add to your Windsurf MCP config (.windsurf/mcp.json):
{
"mcpServers": {
"contextq": {
"command": "npx",
"args": ["-y", "@contextq/mcp"],
"env": {
"CONTEXT_API_KEY": "sk_live_YOUR_API_KEY",
"CONTEXT_API_URL": "https://ctx.example.com"
}
}
}
}Related MCP server: Agent Construct
What data is sent and tenant isolation
Only the requests your agent makes are sent. The MCP server is a stateless proxy -- it forwards each tool call to the ContextQ API via
CONTEXT_API_URLand returns the response. No telemetry, no background sync, no usage tracking beyond what your ContextQ server logs.Tenant-scoped API keys. Every ContextQ API key is bound to a single tenant. All
/api/*endpoints enforce tenant isolation -- an API key can only access the tenant it was issued for. Cross-tenant data leaks are impossible at the API layer.Per-request auth. Your
CONTEXT_API_KEYis sent as anAuthorization: Bearerheader on every call. It never appears in tool names, argument schemas, or responses returned to the LLM.
Client timeout configuration
A handful of ContextQ tools run LLM calls, kNN scans, or bulk DB operations server-side and can legitimately take longer than a typical MCP client's default request timeout. If your client aborts before the server responds, you will see a timeout error that looks like a broken tool — it usually isn't. Configure a longer per-server timeout for this MCP server rather than assuming the tool is hung.
Slow-class tools (recommend a longer timeout, e.g. 120000-180000 ms depending on workspace size):
Tool | Why it's slow |
| Clusters a workspace's contexts via vector similarity, then runs one LLM synthesis call per cluster. |
| Runs LLM judging over up to 20 nearest-neighbor contexts to decide links/archival. |
| Fetches/parses a source and runs LLM claim extraction + kNN diffing. Large or URL-sourced ingests already return |
| Re-clusters all of a tenant's contexts and runs one LLM synthesis call per cluster (admin scope). |
| Samples older contexts and asks the LLM to verdict each one (superadmin scope). |
| Applies a lifecycle/archive patch to up to 200 context ids in one call — bounded, but still slower than a single-row update. |
| Deletes up to 5000 |
| Clone or diff a workspace's full memory state (contexts, links, goal graph) — cost scales with workspace size. |
Everything else (ctx_search, ctx_get, ctx_save, ctx_list, agent_*, goal_*, relay_*, etc.) is ordinary CRUD/search and should complete well within a default client timeout.
These numbers are starting points, not guarantees — actual latency depends on your ContextQ server's hardware, workspace size, and configured LLM/embedding provider. Measure against your own deployment before tuning tighter.
.mcp.json per-server request_timeout_ms
Most MCP clients that support .mcp.json (including Claude Code) accept a per-server request_timeout_ms to override the client's default request timeout for every tool call on that server:
{
"mcpServers": {
"contextq": {
"command": "npx",
"args": ["-y", "@contextq/mcp"],
"env": {
"CONTEXT_API_KEY": "sk_live_YOUR_API_KEY",
"CONTEXT_API_URL": "https://ctx.example.com"
},
"request_timeout_ms": 120000
}
}
}request_timeout_ms applies per server, not per tool — if you regularly call slow-class tools, size it for the slowest one you expect to hit, not the average. Claude Code 2.1.206 fixed a bug where this field was silently ignored (a 60s default was applied regardless); confirm your Claude Code version is at least 2.1.206 if the setting doesn't seem to take effect.
Claude Code idle timeout
Independently of request_timeout_ms, Claude Code (2.1.187+) also enforces CLAUDE_CODE_MCP_TOOL_IDLE_TIMEOUT — an idle-abort timeout (default around 5 minutes) that fires if an MCP tool call produces no activity for that long. Set it in your shell environment (not .mcp.json) when calling slow-class tools against a large workspace:
export CLAUDE_CODE_MCP_TOOL_IDLE_TIMEOUT=300000 # milliseconds; raise if ctx_dream/ctx_ingest still time outTreat both settings as recommendations, not guarantees, of how long any given call will take.
API version compatibility
The 99-tool surface exposed by this MCP server is a direct projection of the ContextQ API (24 loaded by default, the rest via CONTEXT_MCP_TOOL_PROFILE=full or on-demand -- see "Client configuration" above). The tool count and signatures drift with the server. Pin compatible versions:
MCP package | ContextQ server API |
| ContextQ v2.x (99 tools) |
When upgrading your ContextQ server, check the changelog and bump the MCP package to the matching major version. A version mismatch may surface unknown tools or break call signatures.
License
MIT
Available Tools
24 toolsagent_bootAInspect
Boot an autonomous agent: ONE token-budgeted call returning everything needed to start or resume work. Call this FIRST in any agent run. Returns {agent, session:{...,role}, resume:{checkpoint_summary, open_tasks}, handoff:{tldr, source}, lessons:[], facts:[], brief, skills:[], skills_full, repo_map, siblings:[], budget:{limit, used, dropped}, client}. If a non-terminal session exists for this agent (or session_id is given), resume tells you exactly where you left off; handoff is the best-ranked latest handoff (a hand-written wrap SEED first). goal drives the facts, lessons AND skills retrieval. skills are parametrized procedures distilled from verified past runs matching the goal ({context_id, name, description, success_count, similarity}) -- check them BEFORE re-deriving a solution. On a session's later boots skills holds only new or changed entries and skills_full is false (empty then means nothing new, not no skills); pass full:true for the complete set. With project_id you also get a goal-graph brief {north_star, role, lane, next, blocked_on, blocking, done} and the session role is inferred from the matched goal node. With include_repo_map:true on a code-indexed workspace, repo_map carries top-ranked file signatures to answer "where is X handled" without grepping. siblings lists other active sessions in this workspace (last ~60 min, max 5) so you can coordinate via relay_* before touching shared resources. Slots fill resume > handoff > lessons > facts > brief; overflow is reported in budget.dropped. Trigger: resume or catch up on tracked work ("hôm trước tới đâu", "tiếp gì", "tóm lại đang làm gì", "what's next", "resume", "where did we leave off", "catch me up"). Skip for unrelated casual questions.
| Name | Required | Description | Default |
|---|---|---|---|
| full | No | T506: bypass the skills diff-since-last-boot behavior and always return the full current skills match set. Default false (repeat boots of the same session return only new/changed skills). | |
| goal | No | The objective for this run — drives relevant-facts + lessons retrieval AND goal-node role inference | |
| agent | Yes | Stable agent slug (handle the agent boots with every run, e.g. 'claude-code') | |
| workspace | No | Workspace slug to scope handoff + facts to | |
| agent_name | No | Human-readable name; used only when the agent is first created | |
| project_id | No | Project id to scope the goal-graph situation brief + role inference to (omit = no brief, classic pack) | |
| session_id | No | Resume a specific session by id (otherwise the latest active/paused session for this agent) | |
| token_budget | No | Max tokens for the assembled pack (default 4000) | |
| epistemic_min | No | Epistemic floor (T358) for the FACTS slot: only surface facts at or above this confidence tier (weakest->strongest: assumed < inferred < told < observed). Omit for no floor. | |
| include_repo_map | No | When true, include a token-budgeted repo map (entries with path+signatures) from code-indexed contexts. Only useful for workspaces indexed with contextq index. Default false. | |
| repo_map_token_budget | No | Token cap for the repo map slot (default ~2000, range 100-16000). Ignored when include_repo_map is false. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses state-dependent behavior (skills diff-since-last-boot on repeat boots, where empty means 'nothing new'), slot precedence ('resume > handoff > lessons > facts > brief'), and overflow reporting via budget.dropped. It also explains role inference and sibling-session scope (~60 min, max 5).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and gets denser but rarely wastes a sentence — the trigger examples and slot/resume rules all earn their place. It is a long wall of text, however, and could be broken into scannable lines for faster parsing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description enumerates the returned pack structure in detail (agent, session.role, resume, handoff, lessons, facts, brief, skills, repo_map, siblings, budget, client) and explains how each conditional branch is populated. An agent has everything needed to call it and interpret the response.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is already 100%, so the baseline is 3, but the description adds real meaning beyond the schema: 'goal' drives facts/lessons/skills retrieval AND role inference, 'full:true' overrides the diff behavior, and 'project_id' gates the goal-graph brief. It adds useful semantics for the highest-impact parameters rather than just restating names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Boot an autonomous agent') plus the distinguishing value proposition: a single token-budgeted call returning everything to start or resume work. It clearly positions itself apart from siblings like agent_session_start and agent_resume, and even names relay_* for coordination.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit routing: 'Call this FIRST in any agent run', a concrete trigger list for resume/catch-up work (with multilingual examples), and an explicit exclusion ('Skip for unrelated casual questions'). Also states when the resume/handoff/skills branches activate, so an agent knows exactly when this beats alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agent_checkpointAInspect
Snapshot the agent's working state so a restart/crash can resume from exactly here. state is an arbitrary JSON scratchpad (cursor, partial results, plan, open files). summary is a 1-line 'where I am'. Returns the checkpoint with its monotonic seq. Call periodically after each chunk of progress.
| Name | Required | Description | Default |
|---|---|---|---|
| state | No | Working-state scratchpad (arbitrary JSON) | |
| summary | No | One-line human-readable 'where I am' | |
| session_id | Yes | Session id to checkpoint | |
| token_estimate | No | Optional explicit token size of the state (auto-estimated if omitted) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses what gets stored (arbitrary JSON: cursor, partial results, plan, open files) and the return value (checkpoint plus monotonic seq), but says nothing about whether it overwrites prior checkpoints, lifecycle/consumption by resume, or permission needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, front-loaded with the purpose, then payload semantics, then return value, then cadence. No filler and every sentence carries a distinct fact.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema, the description supplies the essential missing context: payload shape, return value, and call cadence. It is largely sufficient for correct invocation, with the only gaps being overwrite/lifecycle behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description earns extra by expanding `state` into concrete examples (cursor, partial results, plan, open files) and clarifying `summary` as a one-line 'where I am', adding meaning beyond the terse schema strings. It does not address `token_estimate` or `session_id` semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb ('Snapshot') plus resource ('the agent's working state') with an explicit goal ('so a restart/crash can resume from exactly here'). It distinguishes itself from the surrounding ctx_* and agent_* siblings functionally via the resume framing, though it never names an alternative tool directly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Call periodically after each chunk of progress' gives a clear triggering condition for use. There is no when-not guidance and no explicit routing to siblings like agent_handoff or ctx_save, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agent_handoffAInspect
Generate a handoff document for the session's workspace at run end (wraps the dream handoff generator — LLM-synthesized TL;DR + in-progress + next-steps + open-questions). Links the handoff context back to the session. Pass complete=true to also mark the session completed. Requires an LLM provider configured on the server.
| Name | Required | Description | Default |
|---|---|---|---|
| dry_run | No | Generate without persisting the handoff context | |
| project | No | Optional project slug to narrow the handoff | |
| complete | No | Also set the session status to 'completed' | |
| workspace | No | Workspace slug (falls back to the session's workspace) | |
| session_id | Yes | Session to generate a handoff for | |
| since_days | No | Look-back window in days (default 7) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does meaningful work: it discloses the LLM-synthesized nature, that handoff context is linked back to the session, the side effect of complete=true marking the session completed, and the hard prerequisite of a configured LLM provider. It omits reversibility of the completed status and any auth/permission needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with what is produced and finished with the prerequisite; the parenthetical content list is information-dense rather than filler. Slightly packed but every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-param, no-annotation, no-output-schema mutation tool, the description covers the artifact produced, the main side effect, and the LLM prerequisite. It does not describe what happens to the returned handoff or how errors surface when no provider is configured, leaving a small gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (dry_run, project, complete, workspace, since_days, session_id) is already documented in the schema. The description only reiterates complete=true and adds that workspace falls back to the session's workspace implicitly; baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Generate a handoff document for the session's workspace') and enumerates contents (TL;DR, in-progress, next-steps, open-questions), so an agent knows exactly what artifact is produced. It does not explicitly differentiate itself from nearby siblings like agent_checkpoint or agent_session_end, which is the only thing keeping it from a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clear context: call 'at run end', and pass complete=true to additionally close the session. It also states the prerequisite that an LLM provider must be configured. No explicit when-not or named alternative to sibling tools, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agent_lesson_addAInspect
Record a lesson learned during a run so the agent doesn't repeat the failure. Embedded for goal-relevant recall at the next agent_boot. Mirrors the vibe-loop '## Lessons' log. scope controls breadth: 'session' (this run), 'agent' (this agent always), or 'workspace'.
| Name | Required | Description | Default |
|---|---|---|---|
| why | No | Why it failed | |
| scope | No | How broadly the lesson applies (default 'session') | |
| session_id | Yes | Session this lesson came from | |
| try_instead | No | What to do differently next time | |
| what_failed | Yes | What was attempted that failed |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden and does a decent job: it discloses the post-write behavior (embedded for goal-relevant recall at the next agent_boot) and unpacks the scope enum's effect ('session'/'agent'/'workspace'). It omits any dedup, overwrite, or permission behavior, but the core trait an agent needs is disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three front-loaded sentences that lead with the action and its payoff. The 'Mirrors the vibe-loop Lessons log' clause is mildly incidental jargon, but overall it is tight and well-ordered.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a write tool with no annotations and no output schema, the description covers purpose, timing, scope semantics, and downstream recall behavior. It leaves lifecycle details (overwrite, dedup, author attribution) unspecified, which is a minor gap rather than a blocking one.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all five parameters. The description adds genuine meaning for scope's three breadth tiers, but says nothing about what_failed, why, try_instead, or session_id beyond the self-evident schema text. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Record a lesson learned during a run') plus the outcome it serves ('so the agent doesn't repeat the failure'). An agent understands the intent immediately, though the description never explicitly separates this from the similarly-named ctx_remember sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Clear context: record a lesson after a run so it is recalled at the next agent_boot, with scope controlling breadth. It implies when to use it but names no exclusion or alternative, so it stops short of routing the agent away from ctx_remember or similar siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agent_resumeAInspect
Read the resume bundle for a session WITHOUT booting fresh: latest checkpoint, open tasks (pending/in_progress/blocked), and goal-relevant lessons. Use when you already know the session_id and just need to reload where you left off.
| Name | Required | Description | Default |
|---|---|---|---|
| session_id | Yes | Session id to resume |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the burden, and it does disclose the read-only nature plus the exact contents of the returned bundle (checkpoint, task statuses, lessons). It omits any note on permissions, side effects, or idempotency, which keeps it out of the top tier.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two compact sentences, front-loaded with what the tool returns, then the usage condition. No filler and every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by enumerating the bundle contents, and the single-parameter read tool needs nothing more. An agent has everything required to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter with 100% schema coverage, so the schema already documents session_id. The description adds marginal context by implying the id must be known in advance, but no format or source details beyond the schema. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Read) and resource (resume bundle), and enumerates the exact contents returned: latest checkpoint, open tasks, and goal-relevant lessons. The 'WITHOUT booting fresh' clause cleanly distinguishes it from the agent_boot sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear triggering condition ('when you already know the session_id and just need to reload where you left off') and contrasts behaviorally with booting fresh. It stops short of naming agent_boot explicitly as the alternative to use when the session_id is unknown.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agent_session_endAInspect
End or update an agent session's status. Use status='completed' when the goal is met, 'paused' to suspend (resume later from the checkpoint), 'stalled' when the vibe-loop stall detector trips, or 'abandoned' to drop the run. Setting completed/abandoned stamps ended_at.
| Name | Required | Description | Default |
|---|---|---|---|
| goal | No | Optionally revise the goal | |
| status | Yes | New session status | |
| metadata | No | Metadata to merge into the session | |
| session_id | Yes | Session id to update |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full disclosure burden and does add real behavior: completed/abandoned stamp ended_at, paused preserves a resumable checkpoint, stalled ties to the vibe-loop stall detector. It stops short of stating whether ending destroys session data, permission requirements, or what happens to associated tasks.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence that front-loads the core action and then enumerates statuses without filler. Every clause carries usable information; nothing is redundant with annotations or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 4 params including a nested metadata object and no output schema, the description covers status semantics and one side effect well. The metadata merge behavior and the return shape rely entirely on the schema, which is acceptable but leaves the nested-object handling underexplained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description goes further by giving operational meaning to the status enum values that the schema only labels 'New session status'. It also clarifies that goal optionally revises (rather than replaces) the goal, matching the schema's 'Optionally revise the goal'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('End or update an agent session's status') and clarifies the scope is status/goal/metadata rather than the broader lifecycle handled by agent_session_start, agent_resume, or agent_checkpoint. An agent can distinguish it from siblings without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description maps each status value to an explicit condition: 'completed' when the goal is met, 'paused' to suspend and resume later, 'stalled' when the stall detector trips, 'abandoned' to drop the run. That is high-quality when-to-use guidance, though it never names a sibling alternative or states when not to call this tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agent_session_startAInspect
Start a new agent session (a run with a goal). Returns the created session including its id. Use when beginning a fresh task that you want to track and resume. Pass parent_session_id to chain a resumed run to its predecessor.
| Name | Required | Description | Default |
|---|---|---|---|
| goal | No | The A-Z objective for this run | |
| role | No | Optional role this session plays (frontend, backend, design, ...). Usually inferred at boot from the matched goal node instead. | |
| agent | Yes | Stable agent slug | |
| project | No | Optional project slug within the workspace | |
| metadata | No | Arbitrary run metadata | |
| workspace | No | Workspace slug this run operates in | |
| agent_name | No | Human-readable name (used only on first creation) | |
| parent_session_id | No | Id of the session this one resumes/continues |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses that a session is created and that the response includes the id, but says nothing about side effects, whether re-calling with the same agent/goal is idempotent or duplicates, or error behavior for a mutation-style tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: definition, return value, and usage trigger come first, with the parent_session_id note last. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter creation tool with no annotations and no output schema, the description covers purpose, usage, and the return shape adequately. It omits behavioral details that matter for a creation call, such as what happens to existing sessions or how errors surface.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds genuine meaning only for parent_session_id ('chain a resumed run to its predecessor'), while the other seven parameters are left entirely to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Start a new agent session') and clarifies the domain term as 'a run with a goal', so an agent knows exactly what is being created. It does not explicitly name the sibling it contrasts with (agent_resume), leaving some differentiation to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear trigger ('Use when beginning a fresh task that you want to track and resume') plus a specific conditional for parent_session_id. It stops short of stating when NOT to use it versus agent_resume or where a resumed run belongs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agent_task_tickAInspect
Flip a task's status. Setting status='verified' REQUIRES non-empty evidence (real observed output: test result, HTTP status, exit code) — the no-self-certification rule. Returns 400 if you try to verify without evidence. Use this as each checklist item is proven. Trigger: user reports finishing a piece of tracked work ("xong rồi", "xong X", "done X", "done", "mark done", "finished X") — tick the matching session-scoped checklist item here (use goal_advance instead for a board-level task). Only call when it maps to a tracked item; ignore unrelated casual chatter.
| Name | Required | Description | Default |
|---|---|---|---|
| status | Yes | New status | |
| task_id | Yes | Task id to tick | |
| evidence | No | Real observed output proving the task (required to set 'verified') |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does so well: it discloses the no-self-certification rule, that status='verified' requires non-empty evidence, and the 400 failure mode. It stops short of stating reversibility, idempotency, or what a successful call returns, which are the remaining gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action and the hard constraint, then the trigger guidance. Dense but every clause carries information; the parenthetical phrase list is slightly verbose but serves a routing purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 3-param mutation tool with no annotations and no output schema, the definition supplies the conditional requirement, the error behavior, the routing trigger, and the sibling alternative. Only the success response shape and state-transition semantics are left unstated, which is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real meaning: it ties the `evidence` parameter to a specific semantic requirement (real observed output such as test result, HTTP status, exit code) and couples it conditionally to status='verified'. That is meaningfully beyond the bare schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Flip a task's status') and immediately distinguishes itself from the sibling goal_advance, which handles board-level tasks. An agent can tell exactly what this does and what it is not.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit trigger conditions are given ('user reports finishing a piece of tracked work' with concrete phrases), an exclusion is stated ('ignore unrelated casual chatter'), and the alternative tool is named for the board-level case. This is close to ideal routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
agent_task_upsertAInspect
Create or update one checklist item in a session's task tree. Omit task_id to create; pass task_id to update. verify_cmd names HOW the item is proven done (the agent must run it before ticking). Use parent_task_id for subtasks. This productizes the vibe goal-file checklist. Trigger: call at the START of a tracked piece of work to record a checklist item (session-scoped — for a task meant to persist across sessions use goal_add on the board instead). Only call when the work is actually being tracked; ignore unrelated casual chat.
| Name | Required | Description | Default |
|---|---|---|---|
| title | No | Task title (required when creating) | |
| status | No | Task status | |
| task_id | No | Existing task id to update (omit to create) | |
| context_id | No | Optional id of the durable context this task produced | |
| session_id | Yes | Session that owns this task | |
| verify_cmd | No | Command/observation that proves this task done | |
| order_index | No | Ordering within the session | |
| goal_node_id | No | Optional goal-graph node this task rolls up to (links session work to the project goal) | |
| parent_task_id | No | Parent task id for a subtask |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so the description carries the full behavioral burden. It does add real context the schema lacks — that verify_cmd must be run before ticking, and that the item is session-scoped — but it says nothing about update semantics (overwrite vs. merge of omitted fields), required permissions, or error behavior for a mutation tool with 9 params.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the create/update distinction and the trigger in a compact block. One sentence — 'This productizes the vibe goal-file checklist' — is internal jargon that earns no place and costs clarity, but the rest is tight.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter, no-annotation, no-output-schema mutation tool, the usage and routing story is complete but the mutation contract is not: what an update does to unset fields, whether status transitions are validated, and whether subtask creation requires an existing parent are all unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents task_id, parent_task_id, title, and the rest; the baseline of 3 applies. The description's gloss on verify_cmd ('names HOW the item is proven done') adds intent but not syntax, and its create-vs-update rule repeats the task_id schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb pair and resource ('Create or update one checklist item in a session's task tree') and immediately disambiguates the two modes via 'Omit task_id to create; pass task_id to update.' It also names the sibling goal_add and the condition that distinguishes them, so an agent can route without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit when-to-use ('call at the START of a tracked piece of work'), when-not ('ignore unrelated casual chat'), and the alternative for the cross-session case ('use goal_add on the board instead'). This is a full routing rule, not an implication.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ctx_deleteAInspect
Permanently delete a context entry by ID. This action cannot be undone. Use this only when you are sure the entry should be removed.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The context entry ID to delete |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden; it does disclose the most important trait — permanence and irreversibility ('This action cannot be undone'). It omits other relevant behavior: whether deleting a missing ID errors, whether removal cascades to related entries, and any permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the action and immediately followed by the destructive consequence. Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter, no-output-schema delete tool, the description covers the action, the input target, and the critical irreversibility warning. Only edge-case behavior (not-found handling, side effects on related entries) is unaddressed, which is a minor gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single 'id' parameter is already documented as 'The context entry ID to delete'. The description's 'by ID' adds no syntax, format, or lookup detail beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Permanently delete a context entry by ID'), which is unambiguous and easy to distinguish from read-oriented siblings like ctx_get or ctx_search. It stops short of naming a sibling (e.g., suggesting ctx_update as the non-destructive alternative), so it is clear but not sibling-differentiating.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Use this only when you are sure the entry should be removed' gives a weak precondition for use and implicitly warns against casual deletion. However, it names no explicit alternative (such as ctx_update for modification) and gives no guidance on when deletion is the wrong choice versus another tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ctx_getAInspect
Get a single context entry by its ID. Use this when you already know the exact context ID you want to read. Response includes lifecycle, atomic, quality_score, valid_from, and valid_to alongside the standard fields.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The context entry ID |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description carries the full behavioral burden. It discloses the extra return fields (lifecycle, atomic, quality_score, valid_from, valid_to), which is genuinely useful, but says nothing about what happens on a missing/invalid ID, permission requirements, or that this is a non-mutating read. Adequate but with real gaps for a no-annotation tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, front-loaded with the core action, then the usage condition, then the return payload. No redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully enumerates the notable return fields, and a one-param getter needs little more. The main omission is error/not-found behavior, which an agent would likely want.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with a single documented parameter, so the schema already carries the semantics and the baseline is 3. The description only restates "exact context ID" without adding format, type, or lookup constraints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ("Get a single context entry by its ID") and scopes it to a single record. It contrasts implicitly with retrieval-by-search via "single" and "exact context ID," but never names the sibling tools (ctx_search, ctx_list) it differs from.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Use this when you already know the exact context ID you want to read" gives a clear selecting condition and implies the when-not case (when you don't know the ID). It stops short of naming the alternative tool an agent should fall back to.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ctx_healthAInspect
Run the deep health probe and return the full report. Probes DB, Elasticsearch, embedding provider, LLM provider, and scheduler states. 30-second in-memory cache on the server. Returns {status, components, schedulers, queue} - status is ok|degraded|fail. Useful for ad-hoc prod health checks from MCP clients.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full behavioral burden, and it does disclose meaningful traits: the exact components probed, a 30-second server-side in-memory cache (i.e., results may be stale by up to 30s), and the returned shape with the ok|degraded|fail status domain. It omits any note about auth/permission requirements or whether it is strictly read-only, which would matter for a prod-facing probe.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the action and followed by scope, caching behavior, and return shape. The trailing 'from MCP clients' is slightly redundant, but no sentence is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter diagnostic with no output schema, the description supplies everything an agent needs: what is checked, the caching/staleness caveat, the return keys, and the status enumeration. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters and the schema is empty with 100% coverage, so there is nothing for the description to clarify. Baseline 4 applies; no parameter guidance is needed or missing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Run the deep health probe') and scopes it as deep vs. shallow by enumerating exactly what is probed (DB, Elasticsearch, embedding, LLM, schedulers). It does not explicitly distinguish itself from the sibling ctx_stats, which an agent could plausibly confuse with a health probe.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear context for use ('ad-hoc prod health checks from MCP clients'), which tells the agent this is an on-demand diagnostic rather than a routine data call. It offers no exclusions or named alternative (e.g., when to prefer ctx_stats instead), so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ctx_listAInspect
List context entries with optional filters. Use this to browse existing contexts by workspace, project, type, or tag without a search query.
| Name | Required | Description | Default |
|---|---|---|---|
| tag | No | Filter by a single tag | |
| type | No | Filter by context type | |
| limit | No | Max results to return (default 20) | |
| scope | No | Filter by visibility scope | |
| offset | No | Offset for pagination | |
| project | No | Filter by project identifier | |
| workspace | No | Filter by workspace identifier | |
| memoryKind | No | Filter to a single taxonomy kind: 'episodic', 'semantic', or 'procedural'. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. 'List' implies a safe read, but it does not explicitly state read-only semantics, default ordering, or pagination behavior — the latter is only inferable from the schema's limit/offset fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with the primary action front-loaded and the usage context second. No filler, though it is on the terse side.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For an 8-parameter list tool with no annotations and no output schema, the description covers purpose and filters but says nothing about the shape of returned entries, default result ordering, or pagination semantics. It is adequate but leaves gaps an agent would want filled.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description names four filter axes (workspace, project, type, tag) but omits scope, memoryKind, limit, and offset, adding little beyond what the schema already documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb+resource ('List context entries') and enumerates the filterable dimensions (workspace, project, type, tag). It implicitly distinguishes itself from ctx_search with 'without a search query', though it never names that sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly states the use case: browsing existing contexts by filter rather than searching. The contrast with a query-based tool is implied but there is no explicit when-not-to-use or naming of ctx_search as the alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ctx_load_tool_groupAInspect
Load one additional group of ContextQ tools into this session (group names come from ctx_tool_groups) so they become callable without reconnecting. Pass "all" to load every remaining ContextQ tool at once. Some MCP clients need to refresh their tool list to actually see newly loaded tools in the model's context -- if a loaded tool still doesn't show up, call it directly by name anyway (ContextQ accepts a tool call for any known tool name regardless of what tools/list currently returns), or restart this server with the environment variable CONTEXT_MCP_TOOL_PROFILE=full to get every tool from the start.
| Name | Required | Description | Default |
|---|---|---|---|
| group | Yes | Group name from ctx_tool_groups (e.g. "pkm", "admin", "relay", "knowledge-graph"), or "all" to load every remaining ContextQ tool. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. It discloses that loaded tools become callable without reconnecting, that some MCP clients need a tool-list refresh, that direct calls work even if tools/list does not show them, and that CONTEXT_MCP_TOOL_PROFILE=full loads everything upfront. It does not cover auth, idempotency, or return behavior, but it is substantially informative.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description front-loads the core action before moving to the "all" shortcut and client-refresh troubleshooting. The troubleshooting sentence is long but earns its place by explaining a likely failure mode and concrete alternatives.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool that mutates the session's available toolset, the description covers the operation, the special "all" value, client visibility caveats, and two fallback paths. It does not describe return values or invalid-group errors, but no output schema exists and the critical operational context is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the single required parameter already documents group names, examples, and the "all" special value. The description repeats that group names come from ctx_tool_groups and that "all" loads every remaining tool, but adds no syntax or format detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: "Load one additional group of ContextQ tools into this session." It distinguishes the action from the sibling ctx_tool_groups by explaining that group names come from that listing tool, so an agent can tell which tool loads versus lists groups.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It tells the agent to get group names from ctx_tool_groups and explains that passing "all" loads every remaining tool. It also gives clear fallback behavior for clients that do not refresh their tool list. It stops short of an explicit when-not-to-use condition, but the context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ctx_rememberAInspect
Extract durable memories from a raw multi-turn conversation and save them as deduped atomic contexts. Turn-aware sibling of ctx_ingest: the server builds a speaker-attributed transcript, extracts only durable facts/preferences via LLM (skipping chit-chat), and runs the claims through the SAME kNN-dedup + diff + create/update/archive pipeline ctx_ingest uses. Pass subjectId to scope memories to a single end-user of your application (Mem0-parity user_id) — dedup then only considers that subject's own prior memories, and every created context is tagged with that subjectId so ctx_search (subjectId param) and GET /api/memory can retrieve it later. Set dryRun=true to preview without persisting.
Long conversations run async — the response is { jobId, statusUrl } and you must poll ctx_ingest_status (or GET /api/ingest-jobs/:id) until status='succeeded' or 'failed'. Short conversations return the full result inline. Pass async=true/false to force a path explicitly.
| Name | Required | Description | Default |
|---|---|---|---|
| async | No | Force the async path (true) or sync path (false). Omit to let the server auto-pick — conversations longer than MEMORY_ASYNC_THRESHOLD messages (default 8) run async. | |
| dryRun | No | When true, run the full extract + diff pipeline but skip every DB write. Default false. | |
| project | No | Optional project identifier within the workspace. | |
| messages | Yes | Conversation turns in chronological order. | |
| agentSlug | No | Optional identifier of the agent that produced/consumed this conversation. Recorded as metadata only. | |
| maxClaims | No | Cap on claims extracted from the conversation. Default 10, hard max 25. | |
| sessionId | No | Optional conversation/session identifier. Recorded as metadata and on the audit row only. | |
| subjectId | No | End-user identity this conversation belongs to (Mem0-parity user_id). Scopes dedup and tags every created context so it can be retrieved later via ctx_search subjectId or GET /api/memory. | |
| workspace | No | Workspace identifier the extracted memories belong to. Falls back to the request's active scope when omitted. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so well: it discloses the LLM extraction step (durable facts only, skipping chit-chat), the shared kNN-dedup + diff + create/update/archive pipeline, the async threshold behavior, the { jobId, statusUrl } response and the need to poll ctx_ingest_status, plus dryRun's non-persisting semantics. This is rich operational context well beyond the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose and sibling relationship are front-loaded, and the three short paragraphs are dense with useful detail rather than filler. There is mild redundancy in the async explanation, which repeats what the 'async' parameter description already states.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by describing both response shapes (inline full result for short conversations, { jobId, statusUrl } for async) and the polling requirement. Nothing an agent needs to call it correctly appears to be missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all nine parameters; the description's explanations of subjectId scoping, dryRun, and async largely restate the schema descriptions rather than adding new syntax or constraints. Baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific compound verb (extract and save durable memories) and its input resource (a raw multi-turn conversation), and explicitly frames itself as the 'turn-aware sibling of ctx_ingest', letting an agent separate it from the ingestion tool without opening either schema. Scope and output (deduped atomic contexts) are both named.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete usage conditions: use subjectId to scope memories to a single end-user, use dryRun to preview, and how async selection works. It clearly explains the relationship to ctx_ingest but does not state when NOT to use this tool versus that sibling, so it falls short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ctx_saveAInspect
Save a new context entry (reference doc, feedback, project note, incident report, lesson learned, or user profile). Use this when you want to persist knowledge for future retrieval. Optional lifecycle/valid_from/valid_to flag the note's maturity and bi-temporal validity. Response includes atomic, quality_score, and lifecycle once the backend judge has run. Trigger: user asks you to remember/save something ("nhớ cái này", "lưu lại", "ghi nhớ giúp", "remember this", "save this", "note this down") — call whenever work-relevant info should persist across sessions. Only call when the request is actually about tracked work/memory; ignore unrelated casual chat.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes | Short, descriptive title | |
| tags | No | Tags for categorization and filtering | |
| type | Yes | Category of the context entry | |
| scope | No | Visibility scope (defaults to personal) | |
| content | Yes | Full content / body of the context entry | |
| project | No | Optional project identifier within the workspace | |
| metadata | No | Arbitrary key-value metadata | |
| valid_to | No | Bi-temporal: ISO 8601 timestamp when the fact stopped being true. Optional; null means still valid. | |
| lifecycle | No | Lifecycle state of the note. Omit to let the backend default to 'working'. Use 'fleeting' for transient captures, 'evergreen' for durable knowledge, 'archived' to retire from active surfacing. | |
| workspace | Yes | Workspace identifier that owns this context | |
| memoryKind | No | Taxonomy override: 'episodic' (an event/interaction happened), 'semantic' (durable factual/reference knowledge), or 'procedural' (how-to / lesson that changes future behavior). Omit to let the backend classify it from the context type (a cheap heuristic, optionally refined by the atomicity judge). | |
| valid_from | No | Bi-temporal: ISO 8601 timestamp when the fact this context describes started being true. Optional. | |
| description | Yes | One-line summary used for search ranking |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden well: it discloses that a backend judge runs asynchronously and returns atomic/quality_score/lifecycle, and that scope/lifecycle/memoryKind have backend defaults when omitted. It omits permissions/auth requirements and duplicate-handling behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded and each sentence carries information (types, triggers, defaults, response fields). The bilingual trigger list is slightly verbose but plausibly earns its place for matching.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully explains return fields (atomic, quality_score, lifecycle) and default behavior for optional params on a 13-param mutation tool. It still lacks error/validation and permission context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 13 parameters, including lifecycle, valid_from/valid_to and memoryKind. The description only restates the bi-temporal/maturity framing, adding little beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Save a new context entry') and enumerates the entry types it accepts. It implicitly distinguishes itself from ctx_update via 'new', but never names a sibling, so it falls short of full differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-call ('whenever work-relevant info should persist across sessions') plus when-not ('ignore unrelated casual chat'), reinforced with concrete trigger phrases in two languages. This leaves almost nothing to inference, though it does not name a sibling alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ctx_searchAInspect
Search / recall saved knowledge and memories using hybrid full-text + semantic search ranked by relevance — the default tool for 'what do I know about X' or 'did I already save this'. Use this when you need to find, remember, or look up existing knowledge by keyword or phrase. Set chunk_search=false to disable per-chunk passage matching, lifecycle_boost=false for legacy ranking, or include_archived=true to surface retired notes. Results may include lifecycle, quality_score, and matched_chunk per hit.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | No | Filter by tags (all must match) | |
| type | No | Filter by context type | |
| level | No | Layered representation level. 'full' (default) keeps the stored description; 'paragraph' replaces it with a ~120-word distill; 'sentence' replaces it with a ~25-word claim. Use 'sentence' for cheap high-density agent prompts where every token matters. Falls back to the stored description when the distill is not yet populated for a row. | |
| limit | No | Max results to return (default 20) | |
| query | Yes | Search query (full-text + semantic) | |
| scope | No | Filter by visibility scope | |
| offset | No | Offset for pagination | |
| project | No | Filter by project identifier | |
| subjectId | No | Filter to memories scoped to a single end-user (Mem0-parity user_id axis). Matches the subjectId used when the memory was created via ctx_remember / POST /api/memory. Omit to search across all subjects. | |
| workspace | No | Filter by workspace identifier | |
| memoryKind | No | Filter to a single taxonomy kind: 'episodic' (events/interactions), 'semantic' (durable reference knowledge), or 'procedural' (how-to / lessons). Omit to search across all kinds. | |
| chunk_search | No | When true (default), search at the chunk level so individual passages can match. When false, only whole-context fields are scored. | |
| epistemicMin | No | Epistemic floor (T358): only return contexts at or above this confidence tier (weakest->strongest: assumed < inferred < told < observed). E.g. 'told' excludes 'assumed'/'inferred' rows. Omit to search across all tiers. | |
| lifecycle_boost | No | When true (default), apply the evergreen/fleeting lifecycle multipliers to the ranking. Set false for legacy ts_rank * tagBoost * recencyBoost only. | |
| include_archived | No | When true, include lifecycle='archived' rows. Default false — archived notes are excluded from regular searches. | |
| trace_session_id | No | Scope trace-event fusion to one agent session (omit to search across the tenant's trace events). Ignored unless include_trace_events is true. | |
| include_trace_events | No | T375: when true, ALSO search episodic tool-call trace summaries ('what did I try before this worked?') and return them in a separate `traceEvents` field. Default false — this is fully additive and never changes `results` or its ranking. Combine with trace_session_id to scope to one session. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and largely meets it: it discloses the hybrid ranking approach, the effect of include_archived (surfaces retired notes), and the return fields (lifecycle, quality_score, matched_chunk). It omits auth/permission requirements and pagination semantics, so it isn't fully transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The core purpose is front-loaded and the whole thing is a tight four sentences. The second sentence ("find, remember, or look up") partly restates the first, a minor redundancy, but nothing else is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 17-parameter tool with 100% schema coverage and no output schema, the description covers the essence plus the key non-obvious flags and the return field shape. It leaves newer features (include_trace_events, epistemicMin tiers, memoryKind filtering) entirely to the schema, which is acceptable given full coverage but not fully self-contained.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema documents all 17 parameters and the baseline is 3. The description adds light rationale beyond the schema (legacy ranking for lifecycle_boost=false, surfacing retired notes for include_archived), so it earns slightly above baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (search/recall) plus resource (saved knowledge and memories), names the mechanism (hybrid full-text + semantic), and positions itself as "the default tool for 'what do I know about X'" — implicitly distinguishing it from save/get/list siblings. An agent can tell what this does and roughly where it fits without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear triggering contexts ("find, remember, or look up existing knowledge by keyword or phrase") and even per-parameter usage advice (chunk_search=false, lifecycle_boost=false, include_archived=true). It does not name explicit alternatives among the many ctx_* siblings or state when NOT to use it, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ctx_statsAInspect
Get aggregate statistics: total context count, breakdown by workspace, type, tag, recently updated entries, and orphan_rate (notes with no tags and no inbound references). Use this for an overview of what is stored and to spot disconnected knowledge.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It helpfully defines orphan_rate ('notes with no tags and no inbound references') and enumerates what the aggregate contains, but it never states that the operation is read-only, nor describes auth/rate-limit behavior. Decent but incomplete for an unannotated tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the concrete output enumeration and followed by the usage statement. Every listed metric earns its place by describing the return payload. Slightly list-dense but no wasted prose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description usefully compensates by enumerating the returned fields (totals, breakdowns, recents, orphan_rate). For a zero-parameter aggregation tool this is close to complete, though it omits return format and read-only confirmation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. There is nothing for the description to clarify beyond the schema, and it correctly adds no spurious parameter detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Get aggregate statistics') and resource, then enumerates the exact metrics returned (totals, breakdowns by workspace/type/tag, recently updated, orphan_rate). This is far more specific than a generic 'get stats'. It stops short of explicitly distinguishing itself from the similarly overview-oriented sibling ctx_health, so it earns a 4 rather than a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear when-to-use guidance: 'Use this for an overview of what is stored and to spot disconnected knowledge.' That maps the tool to two concrete scenarios. It offers no exclusions or explicit alternatives (e.g., how it differs from ctx_health or ctx_list), so it falls short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ctx_tool_groupsAInspect
List additional groups of ContextQ tools not loaded in this session by default -- code-graph lookup, admin/audit, relay handoff, world-model snapshots, saved searches, knowledge-graph traversal, bulk import/ingest, and more. Search here first if a ContextQ tool you expect (a saved search, a relay, a snapshot, a code reference) is missing from your current tool list. Returns each group's name, one-line purpose, member tool names, and how many of them are already loaded, plus how to load a group with ctx_load_tool_group.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it does disclose the return payload in detail (group name, one-line purpose, member tool names, how many are already loaded, and the load mechanism). It implicitly signals a read-only listing operation, though it does not explicitly state safety, permissions, or side effects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in sentence one and the key action ('search here first') in sentence two. The example enumeration (code-graph lookup, relay handoff, etc.) is somewhat list-heavy but genuinely aids discovery of unloaded tools.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and no annotations, the description fully compensates by describing exactly what is returned and how to act on it (load a group via ctx_load_tool_group). Nothing an agent needs in order to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. There is no parameter surface for the description to clarify, and schema coverage is effectively 100%.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('List') and resource ('additional groups of ContextQ tools not loaded in this session by default'), and immediately differentiates itself from all siblings as a meta-discovery/index tool. An agent can tell it apart from ctx_search or ctx_list without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit trigger condition: 'Search here first if a ContextQ tool you expect ... is missing from your current tool list.' It also names the follow-up sibling tool (ctx_load_tool_group) and the action to take, leaving nothing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ctx_updateAInspect
Update an existing context entry. Only the provided fields are changed; omitted fields remain unchanged. Use this to correct, append to, reclassify, or retire (archive) an existing entry. Optional lifecycle/valid_from/valid_to update note maturity and bi-temporal validity.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | The context entry ID to update | |
| name | No | New title | |
| tags | No | Replacement set of tags | |
| type | No | New context type | |
| scope | No | New visibility scope | |
| content | No | New content body | |
| metadata | No | Replacement metadata object | |
| valid_to | No | Bi-temporal: ISO 8601 timestamp when the fact stopped being true; null to keep open. | |
| lifecycle | No | New lifecycle state (fleeting/working/evergreen/archived). | |
| archivedAt | No | Set to null to unarchive, or ISO date string to archive | |
| memoryKind | No | Taxonomy override: 'episodic', 'semantic', or 'procedural'. Setting this locks the value against the fire-and-forget atomicity-judge LLM refinement on this write. | |
| valid_from | No | Bi-temporal: ISO 8601 timestamp when the fact started being true. | |
| description | No | New description |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses partial-update semantics ('only the provided fields are changed; omitted fields remain unchanged') and that archiving is possible, but says nothing about permissions, error behavior on a bad id, idempotency, or side effects of the memoryKind judge lock.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences with no redundancy. The core mutation semantics are front-loaded, and each sentence adds a distinct piece of information (partial update, use cases, temporal/lifecycle fields).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 13-parameter mutation tool with no annotations and no output schema, the description covers the essentials an agent needs: what is updated, partial-update behavior, and how to archive. It falls short only on failure modes and permission requirements.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents all 13 parameters, setting a baseline of 3. The description exceeds that by adding conceptual meaning: lifecycle maps to 'maturity' and valid_from/valid_to map to 'bi-temporal validity', helping an agent reason about those fields.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Update an existing context entry'), and the word 'existing' implicitly separates it from the create-oriented sibling ctx_save. However, it never names an alternative tool, so the differentiation is inferred rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete usage cases: 'correct, append to, reclassify, or retire (archive) an existing entry.' That is clear when-to-use guidance, but no when-not-to-use conditions or explicit routing to siblings (e.g. ctx_save for new entries) are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
goal_addAInspect
Add a node to the goal graph. Progressive elaboration: only title is required -- omit parent_id to create a root node (a vague node is created status='draft'); fill the rest as reality reveals it. kind: objective|milestone|goal|work_item|relay. owner_role/status/origin are free strings. Trigger: when a user (including a non-technical one) asks you to remember or hand off a piece of work for later ("thêm việc", "thêm task", "todo", "add a task", "add to the board"), create a node here with a valid status (draft is fine if details are vague) — this is how a casual request becomes a durable tracked task. Only call when the request is genuinely about work to track; ignore unrelated casual questions.
| Name | Required | Description | Default |
|---|---|---|---|
| do | No | One-sentence work order (stored as payload.do) | |
| kind | No | objective|milestone|goal|work_item|relay (default work_item) | |
| size | No | S|M|L|XL | |
| brief | No | Cold-executor brief: all six keys or omit (a partial brief is rejected). A work_item added without one gets a `hint` with the template. | |
| title | Yes | Node title (the only hard requirement) | |
| origin | No | greenfield|leverage|migrate|unknown | |
| status | No | draft|not_started|ready|in_progress|blocked|done|superseded | |
| content | No | Long description (creates a searchable contexts row) | |
| parent_id | No | Parent node id (containment tree; null/omit for a root) | |
| owner_role | No | Role that owns this node (e.g. frontend, backend, design) | |
| project_id | No | Project id (optional) | |
| session_id | No | Your agent session id, recorded on the node's history row | |
| verify_cmd | No | How a leaf is proven done | |
| effort_weeks | No | Estimated effort in weeks | |
| external_ref | No | External tracker ref, e.g. {jira: 'FIP-123'} | |
| target_weeks | No | Milestone target (weeks) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does well: it explains progressive elaboration (title-only, draft status for vague nodes), that a partial brief is rejected at the schema level, and that a work_item without a brief receives a `hint` containing the template. It does not disclose permissions, idempotency, or the shape of the response, which keeps it below a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The critical action and the 'only title required' rule are front-loaded, and every later sentence adds operational detail. The long trigger sentence with multi-lingual example utterances is heavier than typical but the examples are functional for matching, so the cost is justified rather than wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 16-parameter create tool with no annotations and no output schema, the description covers the core workflow well and flags the side behavior (hint generation) an agent would otherwise miss. It stops short of describing what is returned (e.g. the new node id) or permission requirements, which are the remaining gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3, but the description adds genuine meaning: it explains the mental model behind only-title-required, the draft default for vague nodes, and the consequence of omitting brief. The restating of kind/status/origin values adds little beyond the schema, but the progressive-elaboration framing is real added value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The opening sentence gives a specific verb+resource ('Add a node to the goal graph') and immediately frames it as a progressive-elaboration create operation, which cleanly separates it from read/advance siblings like goal_list, goal_frontier, and goal_advance. An agent knows exactly what this tool produces without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit trigger ('when a user asks you to remember or hand off a piece of work for later'), concrete example utterances, and an explicit exclusion ('Only call when the request is genuinely about work to track; ignore unrelated casual questions'). This is a rare case where when-to-use and when-not-to-use are both stated, so the only missing element is a named alternative tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
goal_advanceAInspect
Advance a node's status. A leaf moving to 'done' REQUIRES non-empty evidence (real observed output) — the no-self-certification rule. Parent status rolls up automatically from children. Trigger: when the user reports finishing a tracked piece of work ("xong rồi", "xong X", "done X", "done", "mark done", "finished X"), advance the matching board node's status here. Only call when the report maps to a tracked board item; ignore unrelated casual chatter.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | Named annotation kept under payload.notes, e.g. {name: "rechecked-2026-09-27", text: "..."} | |
| title | No | Rename the node (manifest task title) | |
| status | Yes | New status (draft|not_started|ready|in_progress|blocked|done|superseded). Always required: to edit a field below WITHOUT a status change, pass the node's CURRENT status. | |
| node_id | Yes | Node id to advance | |
| payload | No | Shallow-merged into the node's payload; `{"do": "<text>"}` is the manifest `do:` line. Keys not named here survive. | |
| evidence | No | Real observed output proving the node (required for a leaf -> done) | |
| session_id | No | Your agent session id, recorded on the node's history row | |
| verify_cmd | No | Replace the node's done-when text (manifest `done-when:`). Whole-value replace, max 2000 chars. | |
| order_index | No | Manifest position | |
| external_ref | No | Shallow-merged into the node's externalRef. Never trusted as a security predicate. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does disclose two meaningful invariants: leaf->done REQUIRES real evidence (no-self-certification) and parent status rolls up automatically. It does not cover permission/auth needs, idempotency, or what the mutation returns, so it falls short of a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the action and the critical evidence rule before the trigger list, with no filler. The multilingual trigger list is long but earns its place for routing; minor slack keeps it below a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter mutation tool with no output schema and no annotations, the description supplies the essential operational rules (evidence gate, rollup, trigger scope). Remaining gaps like partial-payload and external_ref merge behavior are left to the schema, which documents them adequately.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 10 parameters (including the note, payload merge, and status enum semantics). The description reinforces the evidence requirement for leaf->done but adds little syntax or format detail beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Advance a node's status') and immediately scopes it with the rollup and evidence rules. An agent can distinguish it from goal_add/goal_list/goal_frontier without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit trigger condition with concrete utterance examples ('xong rồi', 'done X', 'mark done'), plus a clear when-not rule ('Only call when the report maps to a tracked board item; ignore unrelated casual chatter'). This is exactly the routing guidance an agent needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
goal_frontierAInspect
The ready-frontier: nodes (optionally for a role, or scoped to one objective_run_id) whose ALL blocking dependencies are done and that aren't done yet — i.e. 'what can I start NOW'. Returns GoalNode[] with depsIn/depsOut ([{nodeId, kind}]). A node whose externalRef.blocker is set is waiting on a person or an external event, not on a dependency: treat it as not runnable.
| Name | Required | Description | Default |
|---|---|---|---|
| owner_role | No | Filter to a role's ready work | |
| project_id | No | Filter by project | |
| objective_run_id | No | Scope to one objective's run id |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses the return shape (GoalNode[] with depsIn/depsOut as [{nodeId, kind}]) and a meaningful behavioral nuance that nodes with externalRef.blocker are waiting on external events. It doesn't cover pagination or limits, but for a read tool the disclosure is substantial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core definition before qualifiers, and every clause (role/run_id scoping, return shape, blocker caveat) earns its place. The single long sentence is dense but not padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description correctly compensates by naming the return type and its fields. The blocker caveat adds operational detail an agent needs. Only minor gaps (pagination/limits) remain for this zero-annotation read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all three parameters. The description reinforces owner_role ('optionally for a role') and objective_run_id ('scoped to one objective_run_id') but ignores project_id entirely, adding nothing beyond what the schema provides. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description defines a specific, non-obvious concept ('the ready-frontier') and states the exact semantics: nodes whose blocking dependencies are all done. It effectively distinguishes the tool from siblings like goal_list and goal_advance by answering 'what can I start NOW'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrasing 'what can I start NOW' makes the use case clear, and the externalRef.blocker note tells the agent when a frontier node should NOT be treated as runnable. It does not, however, explicitly contrast this tool with goal_list or other goal siblings, leaving the routing decision partly to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
goal_listAInspect
List goal nodes, filtered. Use to read the graph (your lane, a status column, all milestones, or one objective's whole run via objective_run_id). Returns a COMPACT view by default (id, title, status, kind, externalRef.local_id/priority, blocked, depsIn/depsOut) and omits done nodes unless include_done is true or a status filter is given; call goal_get for one node's full payload, or pass view="full".
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | Filter by kind | |
| view | No | compact (default) drops payload and brief; full returns every field | |
| status | No | Filter by status | |
| owner_role | No | Filter by owning role | |
| project_id | No | Filter by project | |
| include_done | No | Include done nodes (default false) | |
| objective_run_id | No | Filter to nodes stamped with one objective's run id (from goal_set_objective/goal_get/goal_decompose) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does disclose real behavior: a COMPACT default view, done nodes omitted unless include_done is true or a status filter is given, and the option to pass view="full". It omits pagination, result caps, and any auth/permission notes, which are relevant for a graph-listing tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loads the core action, then adds one dense sentence covering return shape, default filtering, and the escalation path to goal_get/full view. The parenthetical field list is long but earns its place by preempting a schema lookup.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 7-param read tool with no annotations and no output schema, the description compensates well by describing both the default compact payload contents and when nodes are excluded. Remaining gaps (pagination/limits, ordering) are minor but real.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description adds meaning beyond the schema: it lists exactly which fields the compact view returns, explains the view default's effect on payload/brief, and clarifies the provenance of objective_run_id (from goal_set_objective/goal_get/goal_decompose).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('List goal nodes, filtered') and immediately enumerates the read modes (your lane, a status column, all milestones, one objective's run). It also explicitly names the sibling it is not, telling the agent to call goal_get for a single node's full payload.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives concrete when-to-use contexts (reading the graph by lane, status, milestones, or objective_run_id) and routes single-node reads to goal_get. It does not mention goal_frontier, which is the closest sibling, so the alternative coverage is clear but not exhaustive.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
24 tool updates
v2.1.1- First observed
agent_boot - First observed
agent_checkpoint - First observed
agent_handoff - First observed
agent_lesson_add - First observed
agent_resume - First observed
agent_session_end - First observed
agent_session_start - First observed
agent_task_tick - First observed
agent_task_upsert - First observed
ctx_delete - First observed
ctx_get - First observed
ctx_health - First observed
ctx_list - First observed
ctx_load_tool_group - First observed
ctx_remember - First observed
ctx_save - First observed
ctx_search - First observed
ctx_stats - First observed
ctx_tool_groups - First observed
ctx_update - First observed
goal_add - First observed
goal_advance - First observed
goal_frontier - First observed
goal_list
TDQS
Scored across 24 tools
Most tools have clearly distinct purposes (ctx_save vs ctx_search vs ctx_get vs ctx_list), and the context CRUD tools are well separated. However, there is a notable cluster of session/keyboard-overlapping tools: agent_checkpoint and agent_resume both expose session state, agent_task_upsert and goal_add both handle tasks, and agent_task_tick and goal_advance both flip task status. The descriptions do clarify the session-scoped vs board-level distinction, but a less attentive agent could still misselect.
The set uses a consistent snake_case convention throughout (ctx_save, agent_boot, goal_add). There is a minor inconsistency in grouping prefixes: context tools use ctx_, agent tools use agent_, and goal tools use goal_ without the agent_ prefix, but all follow a verb_noun or verb_noun_qualifier pattern. The only deviation is a few compound names like agent_task_upsert and ctx_load_tool_group that are slightly more verbose, which is acceptable.
24 tools is at the upper end of the reasonable range and borderline heavy for the core purpose of context/memory management. The set is further expanded by meta-tools (ctx_tool_groups, ctx_load_tool_group) that hint at even more tools behind a loading mechanism, making the actual surface feel bloated. A leaner consolidation of session and goal tools could improve usability.
The surface provides comprehensive coverage for its domain: full CRUD on context entries (save/search/list/get/update/delete), statistics, health checks, memory extraction, agent session lifecycle (start/end/boot/resume/checkpoint/handoff), task management, and goal graph operations. There are no obvious gaps for the stated purpose, and the tool-group mechanism offers a path to additional features without cluttering the default set.
Related MCP Connectors
AI Reasoning Cache & Consensus Layer with 11 MCP tools via Streamable HTTP.
Private-by-default, local-first memory/context/task orchestrator for MCP apps and agents.
Remote MCP server exposing SMI Aware tools, resources, and skills over Streamable HTTP.
Generate contextual prompts and reusable agent skills, evaluate prompts with the 16-dimension Prompt Score, and manage saved work in PromptDrive. Twelve MCP tools also provide authorized access to private Memory for source-grounded answers. Connect over Streamable HTTP using OAuth 2.1 and PKCE. Generation consumes account quota and automatically saves successful results; Memory access follows account permissions and plan limits.
Related MCP Servers
- AlicenseAqualityCmaintenanceA Model Context Protocol server that exposes a MeshCore node and the mesh reachable through it as a clean, high-signal interface for AI agents or tools.3016 npm2MIT
- AlicenseNot gradedqualityDmaintenanceModel Context Protocol server that standardizes tool discovery, execution, and context management for AI applications.MIT

MemoryOS MCP Serverofficial
AlicenseCqualityBmaintenanceEnables agent runtimes, IDEs, and local AI tools to access and manage MemoryOS tenant memory, domain schemas, and cross-agent universal memory through Model Context Protocol tools.26MIT- AlicenseAqualityCmaintenanceProvides a Model Context Protocol server with tools for knowledge base search, retrieval, and addition, plus system time, resources, and prompts. Supports stdio and Streamable HTTP transports for local and shared deployments.4MIT