Skip to main content
Glama

Server Details

Shared memory for coding agents. Stop re-explaining your codebase every session.

If you are the author of this connector, you can claim ownership by verifying the domain or GitHub account it belongs to. Claimed connector authors can inspect health checks, view analytics, and manage their listing.
Status
Healthy
Uptime
91.1% over 42 days
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL
Repository
teamshared-ai/teamshared-plugin
GitHub Stars
0

TDQS

C2.8/5.0

Scored across 98 tools

Disambiguation2/5

Many tools overlap in memory retrieval and context assembly (memory_recall, memory_think, memory_assemble_context, memory_context_pack, memory_profile_get), file/storage upload and publish, and process/event streams. Descriptions are detailed but an agent can easily misselect among similar-sounding tools or aliases.

Naming Consistency4/5

Nearly all names use snake_case with consistent domain prefixes (memory_, work_, file_, storage_, process_run_, agent_run_, org_). A few generic names (call_tool, search_tools, health, version) and alias mismatches (playbook vs procedure) are minor deviations.

Tool Count1/5

98 tools is far beyond typical scoped MCP servers and creates an unwieldy surface for agents. Even for a broad platform, heavy consolidation is needed to make the set navigable.

Completeness3/5

Core CRUD/lifecycle is covered for work, projects, files, storage, memory, and process runs, but notable gaps exist: no memory_forget despite being referenced, no delete/archive for files/storage/projects, and no project update. Some write operations are only reachable through gated approval.

Available Tools

98 tools
account_briefAInspect

Weekly account file for one strategic Person or Organization.

Writes one shared file (stakeholders, initiative, objections, last evidence, next review, five-line brief) attached as an artifact on a create-once account work item. Returns {changed: false} and does not comment when the picture is unchanged. Weekly schedule is bot-side.

ParametersJSON Schema
NameRequiredDescriptionDefault
slugYesPerson or Organization slug (not Contact or Deal)
agentNoOverride agent identity

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden, and it delivers real behavioral detail: it returns '{changed: false}' and does not comment when the picture is unchanged, and it writes to a 'create-once' work item rather than creating a new one each run. This exceeds what the schema alone reveals, though it does not cover overwrite vs. append or auth requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is tightly written and front-loaded: purpose in the first sentence, artifact details, no-op return behavior, and bot-side scheduling each in their own clause. Every sentence earns its place with zero filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values are covered. The description explains what the file contains, where it attaches, the no-op behavior, and the scheduling model. The only notable gap is what happens on a changed picture (presumably it does comment), which is implied but not stated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are already documented. The description adds minimal parameter-specific meaning beyond the schema — it reinforces that slug targets a Person/Organization rather than a Contact or Deal, but introduces no new syntax or formatting guidance. The baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Writes one shared file (stakeholders, initiative, objections, last evidence, next review, five-line brief) attached as an artifact on a create-once account work item.' This is concrete enough to distinguish it from generic siblings like file_create, work_create, and storage_attach, which have no weekly-brief semantics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides scoping context ('for one strategic Person or Organization', 'Weekly schedule is bot-side') and the schema notes slug is 'not Contact or Deal', which narrows usage. However, it never names an alternative or states an explicit when-not-to-use condition, leaving the agent to infer that this is a scheduled bot task rather than a manually invoked tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

agent_run_cancelCInspect

Cancel the active TeamShared Cloud Agent run.

ParametersJSON Schema
NameRequiredDescriptionDefault
agentNoOverride agent attribution label
run_idNoTeamShared run UUID or Cursor run-… id
cursor_agent_idNoTeamShared agent id (ts-…, latest run)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It indicates a mutating action ('Cancel') but does not explain consequences such as whether the run is irreversibly stopped, whether resources are cleaned up, whether confirmation is needed, or what happens to in-flight work. The description adds no behavioral context beyond the verb itself.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single front-loaded sentence with no wasted words. It is concise and immediately identifiable as an action statement. It loses one point because the imprecise 'active' qualifier is not clarified, but overall it is well-sized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has three optional parameters and an output schema, but no annotations and no usage context. The description does not explain how to choose between run_id and cursor_agent_id, what 'active' means in this context, or what side effects cancellation has. For a mutating tool, this is incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds no parameter-level meaning beyond what the schema already provides, and the schema adequately describes run_id, cursor_agent_id, and agent. No additional insight is needed, but none is offered.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Cancel') and resource ('TeamShared Cloud Agent run'), making the basic action clear. However, the qualifier 'active' is slightly misleading because the schema allows specifying a run_id or cursor_agent_id for a particular run, not necessarily the active one. Still, the core purpose is unambiguous and distinct from sibling agent_run_* tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives like agent_run_status, agent_run_followup, or agent_run_hitl_decide. It does not describe prerequisites, side effects to check before cancelling, or situations where cancellation would be appropriate. The agent is left to infer usage from the tool name alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

agent_run_followupAInspect

Enqueue a follow-up prompt on an existing TeamShared Cloud Agent.

Runs on the seat's recorded model, re-checked like agent_run_start; a swap is reported as model_substitution. Keeps the run's tools narrowing and agent record.

ParametersJSON Schema
NameRequiredDescriptionDefault
agentNoOverride agent attribution label
promptYesFollow-up prompt for an existing cloud agent
run_idNoTeamShared run UUID or Cursor run-… id
cursor_agent_idNoTeamShared agent id (ts-…) when run_id is omitted

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does meaningful work: it discloses that the run uses the seat's recorded model, that a model swap is surfaced as model_substitution, and that the run's tool narrowing and agent record are preserved. It still omits permission/auth requirements and whether execution is asynchronous, but the behavioral disclosure is well above baseline.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short lines with the action and target front-loaded and no filler. The middle sentence packs some jargon ('model_substitution') but each clause adds substantive behavior, so only minor tightening is possible.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers the two non-obvious behaviors (model substitution reporting, tool-narrowing persistence) for a no-annotation mutation-style tool. Missing only auth/async semantics, which keeps it short of a 5.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so agent, prompt, run_id, and cursor_agent_id are already documented in the schema. The description adds no parameter-level syntax or format detail beyond noting the run must already exist, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Enqueue a follow-up prompt on an existing TeamShared Cloud Agent.' The word 'follow-up' plus 'existing' cleanly separates it from agent_run_start, and it names that sibling directly, so an agent can route between the two without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The prerequisite (an already-existing cloud agent run) is implied by 'existing' and the analogy to agent_run_start, but there is no explicit when-to-use/when-not or clear alternative selection between this and agent_run_start. Usage is inferable but not spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

agent_run_hitl_decideAInspect

Resume a paused agent run after a human decides on a gated tool call (#872).

ParametersJSON Schema
NameRequiredDescriptionDefault
agentNoOverride agent identity
reasonNoRequired for reject; optional context for approve/edit
run_idYesTeamShared run UUID awaiting approval
decisionYesapprove, reject, or edit the pending destructive tool call
argumentsNoEdited tool arguments when decision=edit

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description carries the full behavioral disclosure burden. It states that a run is resumed and that a human has decided, but it does not disclose the consequences of that decision: approving may execute the previously gated call, rejecting may deny it, and editing changes the arguments before execution. For a mutation tool, this is a meaningful gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One short, front-loaded sentence that states the action first and the trigger second. There is no filler, and it is easy for an agent to scan and understand in one pass.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a full schema, an output schema, and no nested objects, the description does not need to restate parameters or return values. However, the absence of any behavioral/safety context, no annotations, and no statement about what approve/reject/edit causes, leaves the context only partially complete for a tool that gates a destructive call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all five parameters, including the required reason for reject and the arguments field for edit. The prose adds no parameter-level detail beyond the phrase 'gated tool call,' which gives context but does not add semantics beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action ('Resume'), a specific resource ('a paused agent run'), and the exact trigger ('after a human decides on a gated tool call'). This is not a tautology and clearly differentiates the tool from siblings such as agent_run_cancel, agent_run_status, and agent_run_followup.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear condition for use: a paused agent run that is waiting on a human decision about a gated tool call. It does not explicitly spell out when not to use it or name alternatives like request_gated_approval or agent_run_followup, so it stops short of exhaustive routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

agent_run_listAInspect

List TeamShared Cloud Agent runs in this org.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
offsetNo
work_idNoFilter to one work item UUID
cursor_agent_idNoFilter to one TeamShared agent id (ts-…)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It states a read-only action ('List') but does not mention pagination, ordering, auth requirements, or any side effects. For a list operation, it is reasonable to assume safety, but the description fails to explicitly confirm that, leaving the agent to infer behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence that states the purpose without any fluff. It efficiently communicates the essential action and scope, making it easy for an agent to parse quickly. There is no redundant information or unnecessary detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simplicity of a list operation and the presence of an output schema, the description is largely sufficient. It mentions the key scope ('in this org') and the entity type, which is enough for basic use. However, it could briefly note that results are paginated via limit/offset to fully prepare the agent, but this is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds no parameter-specific meaning beyond what the schema already provides. Two of the four parameters (work_id and cursor_agent_id) have schema descriptions, while limit and offset are self-explanatory with defaults and bounds. With 50% schema coverage, the description could have clarified usage, but it does not, leaving the agent to rely solely on schema hints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb ('List') and a well-defined resource ('TeamShared Cloud Agent runs in this org'). It differentiates from sibling list tools like process_run_list and work_list by naming the exact entity type. The scope 'in this org' adds precision, leaving no ambiguity about what is being listed.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no explicit guidance on when to use this tool versus alternatives. While the name and description make it obvious this is for listing agent runs, it does not mention exclusions or point to other tools for different list scenarios. Usage is implied rather than stated, which is acceptable but not proactive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

agent_run_startAInspect

Start a TeamShared Cloud Agent (router + FastMCP). Not a Cursor VM.

The worker mints a short-lived org tsk_ and runs the tool loop against in-process FastMCP. github= is an optional checkout. One active run per work item — a busy run returns status=busy. When the compact bootstrap profile is non-empty it is prepended (capped). When work_id has playbook_slug, that named playbook is prepended after the profile, its skills listed by name for the run to load with memory_skill_get. Unset slug or omitted work_id: no playbook inject. run.model is the model that will run: a model= the router would refuse is substituted (see model_substitution). agent_slug= binds the run to an agent record (#733); the result's agent.inherited says which settings came from it.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoNoOptional workspace slug stored on the run
agentNoOverride agent attribution label
modelNoRouter model id (prefixed, e.g. openrouter/...); omit to use TEAMSHARED_AGENT_RUN_MODEL. An id the router would refuse (no routable prefix) is replaced by that default and the result carries model_substitution={requested, model, reason}.
toolsNoTool name/prefix patterns this run may bind (e.g. ['work_*','memory_recall']). Narrows the box allowlist; never widens it. Omit for the deployment default.
githubNoowner/repo — optional checkout for the worker. Omit for a no-repo run
promptYesTask prompt for the TeamShared Cloud Agent
work_idNoWork item UUID; reuses that item's TeamShared agent
agent_slugNoAgent record to run as (agent-<label> or <label>; see agent_list). The run inherits the record's model and tools: an explicit model= wins, an explicit tools= can only narrow the agent's. Unknown, paused or archived agents are refused.
spawned_byNoParent Agent slug or label (dispatcher that spawned this worker)
parent_run_idNoParent agent run UUID when this is a nested child run; defaults to the run that invoked this tool
auto_create_prNoHint to open a PR when a coding worker finishes (default true)
durable_memoryNoNested runs are ephemeral: no capture session or run memory of their own, the parent's work thread is the record. Set true for a specialist that must keep its own memory.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the full burden and does so richly: single-active-run concurrency, busy status, capped profile prepending, playbook injection order, model substitution with model_substitution={requested, model, reason}, and agent_slug inheritance with agent.inherited reporting. This is well beyond what a bare tool name conveys.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded and appropriately sized for a 12-param tool, with the key disambiguation up front. Some sentences are awkwardly terse ('Unset slug or omitted work_id: no playbook inject.') and layers of parentheticals reduce scanability, but little is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and the description covers the concurrency model, injection behavior, substitution, and inheritance an agent needs to invoke this correctly. Nothing material is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds genuinely non-redundant semantics: how model= is substituted when the router would refuse it, how agent_slug= inherits model/tools (explicit model wins, explicit tools can only narrow), and the ephemeral-vs-durable memory of nested runs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource ('Start a TeamShared Cloud Agent (router + FastMCP)') and immediately disambiguates with 'Not a Cursor VM.' An agent can distinguish this from agent_run_cancel/followup/status siblings without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives situational constraints (one active run per work item; a busy run returns status=busy; playbook/profile injection rules) but never names an alternative tool or states when to prefer agent_run_followup or agent_run_status instead. Usage is implied rather than routed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

agent_run_statusBInspect

Read the latest status of a TeamShared Cloud Agent run.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idNoTeamShared run UUID or Cursor run-… id
cursor_agent_idNoTeamShared agent id (ts-…, latest run)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations available, the description carries full behavioral disclosure. 'Read' indicates a non-mutating operation, which is helpful, but it does not describe behavior when both ids are omitted or invalid, whether it can return partial/in-progress state, or any latency/auth considerations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with no filler: it states the action, resource, and scope efficiently. Every word contributes to understanding the tool's purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema covers return values, and the schema covers parameter semantics, but the description itself lacks guidance on when to use this tool and how the optional ids interact. For a tool with two optional parameters, the absence of usage context leaves a noticeable gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and the schema already documents both parameters, including default behavior for cursor_agent_id. The description adds no parameter-level meaning, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Read'), a resource ('TeamShared Cloud Agent run'), and qualifies it as 'latest status', making the operation clear. It is distinguishable from siblings like agent_run_list or agent_run_cancel, though it does not explicitly name the alternative it is not.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus agent_run_list or agent_run_start, and no exclusions or conditions. It also does not explain how to choose between run_id and cursor_agent_id, leaving usage context entirely to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

attention_snapshotAInspect

Org attention briefing from open work, project status, and dream.

Work is the signal — not Slack or email. Writes one shared file (attached as an artifact on a create-once org attention work item). Returns {changed: false} and does not comment when the picture is unchanged. Hourly schedule is bot-side.

ParametersJSON Schema
NameRequiredDescriptionDefault
agentNoOverride agent identity

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and excels: it discloses that it writes a shared file attached as an artifact on a create-once work item, returns {changed: false} and does not comment when unchanged, and notes the schedule is bot-side. These are concrete behavioral traits beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is exceptionally concise and well-structured: two sentences front-load the purpose, then add key behavioral details. Every sentence earns its place, with no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the core behavior, side effects, and scheduling, which is strong for a tool with a single optional parameter and an output schema. It lacks explicit prerequisites (e.g., org binding) and return behavior when changed, but those are minor given the output schema exists and the tool is simple.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter 'agent' is fully described in the schema ('Override agent identity'), and the description adds no further meaning. Since schema coverage is 100%, the baseline of 3 is appropriate; the description does not enhance parameter understanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it produces an 'Org attention briefing' from 'open work, project status, and dream', giving a specific verb and resource. It also clarifies the data source ('Work is the signal — not Slack or email'). However, it does not name or contrast any sibling tools, so an agent might not easily distinguish it from similar briefing tools like account_brief.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus alternatives. The only usage hint is that the hourly schedule is bot-side, implying it is not intended for manual invocation, but no when-to-use or when-not-to-use instructions are provided, and no alternative tools are mentioned.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

call_toolAInspect

Call a tool by name with the given arguments.

Use this to execute tools discovered via search_tools.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesThe name of the tool to call
argumentsNoArguments to pass to the tool

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden for behavioral disclosure. It merely restates the schema without describing side effects, return values, error conditions, or the fact that calling arbitrary tools may have destructive consequences. This is a significant transparency gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is exactly two sentences, front-loaded with the primary action and followed by a usage hint. Every word earns its place, with no redundancy or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple, but with no output schema and no annotations, the description could have mentioned that the tool returns the called tool's output or that invalid names will error. The core functionality is clear, but some behavioral expectations are missing, so it is minimally viable rather than complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, with both parameters described. The description adds little beyond the schema, repeating 'by name with the given arguments' without explaining how arguments map to the called tool's parameters. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action: 'Call a tool by name with the given arguments,' which specifies the verb and resource. It also differentiates from search_tools by noting the tool is for executing discovered tools, but it does not explicitly distinguish from the sibling executar_lote, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear usage context: 'Use this to execute tools discovered via search_tools.' This tells the agent when the tool is appropriate, but it does not mention exclusions or alternatives, so it is not a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

context_commitAInspect

Turn-end batch: assistant summary + durable writes + optional close.

One call replaces the end-of-turn memory_session_append + memory_remember (+ memory_session_close + memory_state_set) sequence. The append self-heals expired sessions; the response's session_id is authoritative. Returns {session_id, turn_count, reopened, memories, closed, org, bound}.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoNoWorkspace slug; scopes fact tags and the state pointer.
agentNoOverride agent identity
closeNoClose the session (queueing distillation) and clear the state pointer. Pass true when the task is done or the user says goodbye.
factsNoDurable memories to write in the same call: [{"content": "...", "kind": "fact|preference|event|note|outreach", "subject": "...", "tags": [...]}]. Only include things still true next week.
githubNoGitHub owner/repo tag for the facts.
summaryYesFaithful summary of your reply — appended as the assistant turn.
namespaceNoOrg-allowlisted container slug (repo:owner/name, github:owner/repo, or project:name). Tags every fact. Omitted: default from github= then repo=. Invalid/unknown slugs fail. See docs/memory-namespaces.md.
session_idNoWorking-memory session to commit to. Omit to resolve it from the conversation/active-session state pointer (requires repo).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral disclosure burden. It adds useful traits such as self-healing expired sessions, the response's session_id being authoritative, and the returned fields, but it omits side effects of close (queueing distillation, clearing the state pointer) and any permission or failure behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences front-load the one-line summary, then give the replacement relationship, behavioral caveat, and return shape. There is no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the 8-parameter schema with 100% coverage and an output schema, the description provides the essential selection and behavioral context: when to use it, what it replaces, self-healing behavior, and authoritative session_id. The schema fills the remaining parameter and return-value details, so nothing needed to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% with detailed parameter descriptions, so the baseline is 3. The description only maps summary/facts/close to 'assistant summary + durable writes + optional close' without adding new syntax or format details beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening line defines the tool as a 'Turn-end batch' combining assistant summary, durable writes, and optional close, and the second sentence names the exact memory_* sequence it replaces. This makes its purpose and distinction from sibling memory tools immediately clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly positions the tool as the end-of-turn replacement for memory_session_append + memory_remember (+ memory_session_close + memory_state_set), so an agent knows when to prefer it. It does not explicitly state when not to use it, such as for mid-turn appends, so it lacks full exclusion guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

context_compressAInspect

Compress a prompt payload before it reaches an LLM.

Shrinks JSON tool outputs, logs, and long text using SmartCrusher-lite sampling. Originals are stored in CCR (Redis) with ref= markers for context_retrieve. Always runs; tune thresholds via TEAMSHARED_COMPRESS_*.

ParametersJSON Schema
NameRequiredDescriptionDefault
messagesYesOpenAI-style chat messages to compress before sending to an LLM. User messages are preserved; long tool/assistant/system blocks shrink.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose side effects and behavior. It does mention that originals are stored in CCR (Redis) with ref= markers for retrieval, and that it 'Always runs' and thresholds are tunable. However, it doesn't describe the return value, potential data loss, or any permissions needed. It gives some behavioral transparency but leaves gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, with the key purpose in the first sentence. It provides necessary details in two more sentences without excess. The structure front-loads the purpose and then elaborates on behavior and configuration. No fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has a single parameter with good schema coverage and an output schema (not shown). The description explains the purpose and storage side effect but doesn't detail the return format (presumably in output schema), nor does it mention any prerequisites like Redis availability or error conditions. Given it's a straightforward compression tool with one input, it's fairly complete but could state expected output behavior more explicitly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description already covers the 'messages' parameter well, stating it's OpenAI-style chat messages and that user messages are preserved while long tool/assistant/system blocks shrink. Since schema coverage is 100%, the description adds only minor reinforcement (mentions JSON tool outputs and logs) but doesn't significantly extend meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it compresses a prompt payload before sending to an LLM, and specifies what it shrinks (JSON tool outputs, logs, long text). It distinguishes itself from other context tools by focusing on compression and storage of originals. It's not a tautology and gives a specific verb-resource pairing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description says 'Always runs' but does not provide explicit when-to-use or when-not-to-use guidance, nor does it name alternatives. It hints at a companion tool (context_retrieve) but doesn't tell the agent when to choose this over other context tools like context_normalize or context_prepare. The 'always runs' suggests it might be automatic, which is a usage signal but not a clear directive.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

context_normalizeAInspect

Strip, clean, and compress a non-teamshared tool output for agent context.

Trims recall-style payloads, shrinks large JSON/logs, and stores originals in CCR when compressed. Prefer letting MCP middleware handle teamshared tools automatically; call this for Shell, Grep, or other harness tools.

ParametersJSON Schema
NameRequiredDescriptionDefault
outputYesRaw tool output string (usually JSON).
tool_nameYesName of the tool whose output you are trimming.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it discloses real behavioral traits: it trims recall-style payloads, shrinks large JSON/logs, and stores originals in CCR when compressed. This goes beyond the schema and gives agents a clear sense of side effects and preservation behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the core purpose; the second sentence adds valuable usage and side-effect context. 'Strip, clean, and compress' is slightly redundant, and 'CCR' is not expanded, but the overall structure is efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter tool with an output schema, the description covers purpose, usage boundaries, processing behavior, and a notable side effect. Nothing essential is missing for an agent to decide when and how to invoke it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both parameters thoroughly. The description adds context about target tools (Shell, Grep) but no additional parameter-level syntax or formatting details, matching the baseline for full schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action ('Strip, clean, and compress') on a specific resource ('non-teamshared tool output') and states the intended use ('for agent context'). It also distinguishes this tool from the middleware path for teamshared tools, so an agent can tell it apart from siblings like context_compress.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance: call this for Shell, Grep, or other harness tools, and prefer MCP middleware for teamshared tools. It names exclusions and alternatives, leaving little to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

context_prepareAInspect

Pre-LLM pipeline: session append → compress incoming history → enrich.

Returns compressed messages, optional additional_context (org memory), session_id, and stats. Use before sending a turn to your LLM when you want teamshared to shrink tool bloat and inject recall. Server-side MCP middleware already normalizes teamshared tool responses; this covers the rest of the prompt.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoNoWorkspace slug for scoped recall enrichment.
enrichNoAssemble org memory and append as `additional_context`.
githubNoGitHub `owner/repo` for scoped recall enrichment.
promptNoLatest user prompt when you do not have full message history.
messagesNoOpenAI-style chat messages to run through the pre-LLM pipeline. Provide this or `prompt`.
session_idNoWorking-memory session to append the user turn to.
token_budgetNoSoft token cap for assembled context.
append_sessionNoAppend the latest user message to the working session.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the burden. It discloses the pipeline stages and return payload, and it clarifies that teamshared normalization is already handled elsewhere. But it does not state whether session append persists state, whether compression discards messages, or any side effects or permissions, which matters for a tool that mutates a working session.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences front-load the purpose and returns, then give the usage context and scope. No filler; the pipeline arrow notation is efficient and every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers purpose, return values, and usage timing, and the output schema handles return structure. But with eight parameters and siblings like context_compress, context_normalize, and memory_assemble_context, it does not fully disambiguate when to choose this tool over those, nor does it mention the messages-vs-prompt requirement or side effects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all 8 parameters. The description adds no parameter-level detail beyond the schema; it only restates the concepts of session append, compression, and enrichment. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a concrete pipeline ('session append → compress incoming history → enrich') and lists the exact return fields, so an agent can tell this is a pre-LLM context assembly tool. It gestures at differentiation by saying it 'covers the rest of the prompt' after middleware normalizes teamshared responses, but it never names a sibling such as context_compress or memory_assemble_context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit trigger: 'Use before sending a turn to your LLM when you want teamshared to shrink tool bloat and inject recall.' It also clarifies scope relative to server-side middleware. However, it provides no exclusions or named alternatives for cases like standalone compression (context_compress) or memory-only assembly (memory_assemble_context).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

credits_balanceAInspect

Return the current org's prepaid router credit balance.

Balance is in micro-USD (1 USD = 1,000,000 micro-USD). Read-only — no write tool yet (credit purchases land in Phase 2 via Stripe).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description correctly carries the full burden. It explicitly discloses the read-only nature ('Read-only'), the return unit (micro-USD with the conversion factor spelled out), and the roadmap context that writes are not yet available. This provides meaningful behavioral context beyond what structured metadata would cover.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with zero waste. The first sentence front-loads the core purpose, and the second adds the essential unit detail and read-only disclosure. Every word earns its place; there is no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only tool with an output schema covering the return value, the description covers all that's needed: what it does, the unit of measurement, the scope, and the read-only constraint. The only minor gap is not describing edge cases (e.g., behavior with zero balance), which is a negligible concern for such a simple tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, and the rubric grants a baseline of 4 for this case since there is nothing to document. Schema coverage is trivially complete at 100% with an empty object. The description correctly needs no parameter explanation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Return the current org's prepaid router credit balance.' This is unambiguous and clearly distinguishes itself from all sibling tools, none of which deal with credits or balances. The scope ('current org') and object ('prepaid router credit') are precisely defined.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The purpose sentence inherently signals when to use it — when an agent needs the org's credit balance. The 'Read-only — no write tool yet (credit purchases land in Phase 2 via Stripe)' line clarifies what it is NOT (a purchase/credit tool), providing useful exclusionary context even though no direct alternative tool is named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

diagram_generateAInspect

Sketch a teamshared.diagram/v1 from GitHub manifests and store it.

Reads the repo tree plus README, compose, package manifests, infra/, and SQL migrations. Produces mode=generated Mermaid (C4-container or ER sketch — not architecture recovery). Source metadata (github, ref, commit, paths, content_hash) lives on the document; TeamShared is not a git host. Pass work_id / project_id to attach. Refresh later with diagram_sync.

This generator emits teamshared.diagram/v1 only. For the layered 3D scene (teamshared.diagram/v2), author the scene yourself with file_create(content_format='diagram').

ParametersJSON Schema
NameRequiredDescriptionDefault
refNoGit ref (branch, tag, or commit)main
kindNo'software' (C4-container sketch), 'data' (ER), or 'infra'software
viewNoDiagram view label stored on teamshared.diagram/v1 (defaults: C4-container / ER / containers)
agentNoOverride agent identity
titleNoShared-file title; defaults to '{github} {kind} {view}'
githubNoowner/repo to read via the GitHub API. Optional when work_id= already has github= set — that work field is the coding repo pointer, not a second git host.
work_idNoAttach the new shared file to this work item (kind=artifact)
project_idNoAttach the new shared file to this project (kind=artifact)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses the output format (Mermaid), the mode ('generated'), the limitation (not architecture recovery), the fact that TeamShared is not a git host, and the version restriction. It does not explicitly mention side effects like permissions or failure behavior, but the read-and-store nature is conveyed. This is strong but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is moderately long but well-structured, with the core purpose front-loaded and supplementary details (source reading, metadata, alternatives) following logically. Each sentence adds information; no filler. It is slightly verbose but remains efficient for a tool with 8 parameters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (8 optional params, output schema present), the description covers the main workflow, version constraints, and usage context. It explains the source inputs and the attach mechanism. It does not detail the exact output structure, but that is delegated to the output schema. Overall, it provides enough for an agent to decide when and how to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds value beyond the schema by clarifying the github parameter's optionality when work_id already carries a github pointer, and by noting default behaviors for view and title. These clarifications help avoid misuse, making it a 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb ('Sketch') and resource ('a teamshared.diagram/v1 from GitHub manifests'), and immediately distinguishes it from the sibling diagram_sync (refresh) and file_create (v2 authoring). It clearly states what it does and what it is not (architecture recovery), leaving no ambiguity about its role.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states when to use this tool (to generate a v1 diagram from GitHub manifests) and when not to (for the layered v2 scene, use file_create with content_format='diagram'). It also references diagram_sync for later refresh, providing clear routing among alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

diagram_syncAInspect

Re-fetch mapped GitHub paths and append a version if they changed.

No-op when commit + content hash match. mode=authored is refused unless force=true (then the new version is mode=generated). No webhook or periodic worker — call this when you want a refresh. teamshared.diagram/v2 scenes are authored, so this returns changed=false, reason='scene' instead of overwriting them.

ParametersJSON Schema
NameRequiredDescriptionDefault
refNoOverride source.ref (defaults to the document's stored ref)
agentNoOverride agent identity
forceNoOverwrite an authored diagram. Generated diagrams sync without this; authored mode is never overwritten unless force=true
file_idYesExisting diagram shared-file UUID

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full disclosure burden and does so exceptionally: no-op on matching commit/content hash, refusal of authored mode unless force=true, forced versions becoming mode=generated, and scene inputs returning changed=false with reason='scene'. These are concrete behavioral traits beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact, front-loaded with the primary action, and every sentence carries distinct information: main behavior, no-op condition, forced-authored behavior, manual invocation, and scene edge case. There is no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity and the existence of an output schema, the description covers the essential operational context: when it does nothing, when it refuses, what happens under force, and how scene diagrams are handled. The lack of annotations is compensated by unusually thorough behavioral detail.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaning beyond the schema by explaining the authored/generated mode interaction and what force actually produces ('new version is mode=generated'). It does not elaborate on ref or agent, but those are already well documented in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Re-fetch mapped GitHub paths and append a version if they changed.' This clearly conveys the operation and outcome. It does not explicitly name or contrast a sibling like diagram_generate, so it stops short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives an explicit invocation trigger: 'No webhook or periodic worker — call this when you want a refresh.' It also clarifies when the tool refuses (authored mode without force) and when it no-ops. It does not name alternative tools, so it misses the explicit alternatives part of the top score.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

event_appendAInspect

Append one event to a stream. History is never mutated.

ParametersJSON Schema
NameRequiredDescriptionDefault
typeYesDotted event type, e.g. human.update, agent.run.started, or process.run.created
actorNoOptional actor override: human:… / agent:… / computer:…
agentNoOverride agent attribution label
payloadNoJSON object payload (corrections are new events)
stream_idYesStream id: work:{uuid} (task timeline), run:{uuid} (Process Run), entity:{uuid} (ontology node), org:{uuid} (schema), project:{uuid}, session:{id}, agent_run:{uuid}, memory:{uuid} (recall feedback)
on_behalf_ofNoOptional delegated actor (same dialect as actor)
correlation_idNoOptional correlation id to group related events

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the burden of behavioral disclosure. It does disclose a key behavioral trait: 'History is never mutated,' which is important for an append operation. However, it does not mention other behaviors such as idempotency, duplicate handling, or failure semantics, leaving some gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences with no wasted words. The core operation is front-loaded, and the behavioral guarantee is stated immediately after, making it easy to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple append operation with full schema coverage and an output schema, the description is largely complete. It states the core behavior and the most important guarantee (history immutability). It could be slightly stronger with usage context, but nothing critical is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters thoroughly. The tool description adds no new parameter-level meaning beyond what is already present, which matches the baseline for full schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the action and resource: 'Append one event to a stream.' It is specific enough to distinguish from read-only siblings like event_list, though it does not explicitly differentiate from other append-like tools such as memory_session_append.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is for recording a new event in a stream and notes that history is never mutated, which gives some context for when it is appropriate. However, it provides no explicit guidance on when to choose this tool over alternatives like event_list or memory_session_append.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

event_listBInspect

List append-only events for a stream (timeline order).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
sinceNoInclusive ISO timestamp lower bound
untilNoInclusive ISO timestamp upper bound
exportNoIf true, also return a JSONL body for the stream
offsetNo
stream_idNoFilter to one stream (work:{uuid} / run:{uuid} / …)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral burden. It does disclose two useful traits: events are append-only and results are in timeline order. However, it omits other behavioral details like pagination behavior, the meaning of the export flag, or what happens when stream_id is null.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no filler. 'Append-only' and 'timeline order' are both high-signal qualifiers, and the core purpose is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema and partial parameter descriptions reduce the burden on the tool description. Still, the description leaves important context unstated: stream_id is optional, export changes the response shape, and there is no mention of when to prefer this over related list tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 67% (4 of 6 parameters described), but the description itself adds no parameter semantics. The uncovered parameters (limit, offset) are simple, but the description does not compensate for the moderate coverage gap or explain the stream filter in terms of the schema's stream_id.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('List') and resource ('append-only events for a stream'), with an ordering qualifier ('timeline order'). It is clear enough to be distinguished from the write-oriented sibling event_append, though it does not explicitly name alternatives or clarify that stream_id is optional.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus alternatives such as process_run_events or event_append. There are no explicit when-to-use conditions, exclusions, or alternative tool references.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

evidence_attachAInspect

Attach evidence to a Process run via an append-only event.

Source of truth is evidence.attached on run:{id}. A second attach to the same slot supersedes the previous pointer. Agents should prefer this (or file_create(work_id=) auto-attach) over mutating a case-file row.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNofile, drive, url, pr, comment, or text
slotNoEvidence slot (scan-notes). Omit to auto-slot from the pointer
agentNoOverride agent attribution label
titleNoHuman label for this attachment
run_idYesProcess run UUID — evidence lands on run:{id}
pointerYesPointer: file:<uuid>, drive:<uuid>, pr:<url>, or https://…
work_idNoOptional work UUID recorded on the event
project_idNoOptional project UUID recorded on the event
supersedesNoOptional prior pointer this attachment replaces

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It discloses that the operation is append-only, that the source of truth is evidence.attached on run:{id}, and that a second attach to the same slot supersedes the previous pointer. It does not mention auth or error cases, but the core behavioral contract is transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two efficient sentences: the first states the action and core mechanism, the second explains the source-of-truth, supersession, and preference guidance. Every clause adds value, and the purpose is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter tool with an output schema, the description covers the essential usage context, alternatives, and behavioral nuances. It does not mention potential errors or validation, but the output schema likely handles that. The description is sufficiently complete for an agent to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents every parameter. The description adds context around slot/pointer supersession but does not elaborate on individual parameter formats beyond what the schema provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb (attach) and resource (evidence to a Process run) and explains the append-only mechanism. It also distinguishes from alternatives like file_create auto-attach and direct case-file mutation, making the tool's purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit guidance: agents should prefer this tool (or file_create with work_id) over mutating a case-file row, and it explains the supersession behavior for same-slot attaches. This tells the agent exactly when to use it and what to avoid, leaving no ambiguity about alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

evidence_getAInspect

Read the case-file spine (required / received / missing / superseded).

Projection over run:{id} events. Work/Project pass-throughs merge every linked run. Pass one of run_id, work_id, or project_id.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idNoProcess run UUID
work_idNoWork UUID — folds evidence from linked runs
project_idNoProject UUID — folds evidence from linked runs

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses the read-only nature ('Read'), explains the projection over events, and details the merging behavior for work/project pass-throughs. It does not mention side effects (expected none for a read) or access-control requirements, but the disclosed behavior is sufficient for safe invocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, each earning its place. The core purpose is front-loaded in the first sentence, and the parameter guidance is compact and direct. No fluff or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read tool with an output schema, the description covers the operation, the parameters, and the key merging behavior. It does not mention pagination or limits, but these are unlikely to be critical for a spine read and are partially covered by the output schema signal. Adequate for the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds meaningful value beyond the schema: it clarifies that exactly one ID should be passed ('Pass one of'), and it explains that work_id and project_id fold evidence from linked runs, which interprets their otherwise generic schema descriptions. This exceeds baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise verb and resource: 'Read the case-file spine' and lists the states it covers (required/received/missing/superseded). This clearly distinguishes it from write-type siblings like evidence_attach and event-listing tools like event_list, so an agent can tell it apart without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit parameter usage ('Pass one of run_id, work_id, or project_id') but does not state when to prefer this over alternatives, nor does it give negative guidance like 'use evidence_attach to add evidence' or 'use event_list for raw events'. The purpose implies when to use it, but explicit routeing is missing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

evidence_requireAInspect

Declare required evidence on a Process run (append-only event).

Folded with received attachments into missing / received / superseded.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNoExpected kind: file, drive, url, pr, comment, or text
slotYesRequired slot token (scan-notes)
stepNoOptional process step this slot belongs to
agentNoOverride agent attribution label
titleNoHuman label for the required slot
run_idYesProcess run UUID

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description partially carries the behavioral burden by disclosing that this is an 'append-only event' and that it folds into missing/received/superseded states. However, it does not explain failure modes, permissions, idempotency, or what happens when evidence is declared for a non-existent or already-completed run.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and the main purpose is front-loaded in the first sentence. The second sentence is dense but informative, though the phrase 'folded with received attachments' is somewhat compressed and may require domain knowledge to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the moderate complexity, full schema coverage, and presence of an output schema, the description covers the core operation but leaves usage guidance and the meaning of the folding behavior implicit. An agent could call it correctly, but not with full confidence about when it should be chosen over evidence_attach.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all six parameters and their meanings. The description adds no parameter-level detail beyond what the schema provides, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Uses a specific verb and resource: 'Declare required evidence on a Process run' immediately identifies the operation and target. The phrase 'append-only event' plus the contrast with sibling tools like evidence_attach and evidence_get makes its role in the evidence lifecycle clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intent is implied by the action word 'Declare' and by the sibling context, but the description never explicitly says when to use this tool instead of evidence_attach or evidence_get. No exclusions or alternative selection criteria are given, so an agent must infer the boundary between requiring evidence and attaching evidence.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

file_createAInspect

Create a new versioned shared file in the caller's org.

Shared files default to private. Call file_publish to generate the public share URL (/s/{share_token}). Each file_update creates a new immutable version row. Pass work_id to attach it to a task and/or project_id to attach it to a project (same shared_files row; project attachments are kind=artifact only).

content_format='diagram' takes either diagram schema:

  • teamshared.diagram/v1 — body: {format: mermaid, text: "..."}, rendered as Mermaid. Best for flowcharts, sequences, and ER sketches.

  • teamshared.diagram/v2 — body: {format: scene, layers: [...], nodes: [...], edges: [...]}, rendered as a layered 3D scene where each layer is a plane (infra / software / data by convention), each node sits on one layer, and edges may cross layers. kind (default system), view (layers), mode (authored) and body.camera are optional. Every node needs id and layer; shape is box | cylinder | sphere | cone; edges are {from, to, label?, kind?} with kind one of sync | async | data | depends. Omit x/z to get an automatic grid on the plane. Caps: 8 layers, 300 nodes, 600 edges, 256 KiB. Unknown keys are rejected.

ParametersJSON Schema
NameRequiredDescriptionDefault
agentNoOverride agent identity
titleYesFile title
contentYesFile body: markdown, raw HTML, a diagram document JSON/YAML when content_format='diagram', or teamshared.sheet/v1 JSON/YAML when content_format='sheet' (stored as JSON)
work_idNoAttach the new file to this work item UUID
project_idNoAttach the new file to this project UUID (kind=artifact)
content_formatNo'markdown' (allowlist sanitizer), 'html' (sanitized raw HTML), 'diagram' (teamshared.diagram/v1 Mermaid or teamshared.diagram/v2 layered 3D scene, JSON/YAML), or 'sheet' (teamshared.sheet/v1 JSON/YAML; typed grid, stored as JSON)markdown

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description fully carries the behavioral burden. It discloses default privacy, the need for explicit publishing, immutable version behavior, attachment semantics, diagram format details, caps, and strict schema validation ('Unknown keys are rejected').

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but densely informative, front-loads the core purpose, and uses structured paragraphs and bulleted variants. Every major sentence adds operational value, especially the diagram schema details that would otherwise require external documentation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex creation tool with 6 parameters and multiple content formats, the description covers defaults, publishing workflow, versioning, attachments, diagram schemas, limits, and validation behavior. An agent has enough context to invoke the tool correctly; the output schema covers return-value needs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Though schema coverage is 100%, the description substantially enriches parameter meaning. It explains how content_format=diagram maps to two distinct diagram schemas, defines node/edge requirements, caps, and clarifies that work_id/project_id attach to the same shared_files row with kind=artifact for projects.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence names a specific verb, resource, and scope: 'Create a new versioned shared file in the caller's org.' It clearly distinguishes the tool from siblings like file_update and file_publish by emphasizing creation, versioning, and the caller's org.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly routes to alternatives: call file_publish for public share URLs, and notes that file_update is the operation that creates new immutable version rows. It also gives attachment guidance for work_id and project_id, making the tool's role in the workflow unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

file_getAInspect

Fetch a shared file with its latest version content.

Includes public_url (the /s/{slug} link) when the file is published.

ParametersJSON Schema
NameRequiredDescriptionDefault
file_idYesFile UUID

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description must carry the behavioral burden. It does disclose that the response contains the latest version content and conditionally includes a public_url when published. Still, it does not address access/sharing prerequisites, error behavior, or explicitly confirm the absence of side effects beyond the word 'Fetch'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences with no filler. The core purpose is front-loaded, followed by a single specific note about the conditional public_url field.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple single-parameter tool with an output schema, the description covers the essentials: what to fetch, what version to expect, and a notable conditional return field. Minor omissions like authentication expectations and failure semantics are not critical for a straightforward read operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the only parameter, file_id, is described as 'File UUID'. The description adds no extra meaning about the parameter, such as ownership, sharing scope, or how to find the ID, so the schema already carries the full semantic weight.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states a specific operation ('Fetch') on a specific resource ('a shared file') and the key return characteristic ('latest version content'). This clearly differentiates it from sibling tools like file_create, file_update, file_publish, and file_list, which use different verbs or operate on collections.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The purpose sentence implies usage for retrieving a single file, and it is naturally the get-by-id counterpart to file_list. However, there is no explicit guidance about when to prefer this tool over storage_get or when it should not be used, leaving the agent to infer context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

file_listAInspect

List active shared files in the caller's org, newest update first.

Each file includes public_url (the /s/{slug} link, or /s/{share_token} if no slug) when published, plus slug and share_token. Use query to find a file by title without listing everything — e.g. file_list(query="yield vault"). Pass work_id to list documents + the agent transcript on a task (Storage artifact joins; transcripts stay on work_item_files), or project_id for a project's documents. storage_list is canonical for every Storage join including blobs. project_id wins if both filters are set. Archived files stay joined but are omitted.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryNoOptional case-insensitive title substring to filter by, e.g. 'yield vault'
work_idNoOnly files attached to this work item UUID
project_idNoOnly files attached to this project UUID

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full behavioral burden and does so well: it specifies ordering, included fields (public_url, slug, share_token), when fields are present ('when published'), the 'Archived files stay joined but are omitted' exclusion, filter precedence, and the special work_id/Storage artifact behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose is front-loaded in the first sentence, and the remaining paragraphs add dense, non-redundant detail. Every sentence contributes meaningful guidance—order, fields, filters, alternatives, precedence, and exclusions—without filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists, the description covers everything else an agent needs: scope, ordering, filter semantics, field inclusion conditions, archive behavior, precedence, and the relationship to storage_list. The only uncovered param, limit, is fully specified in the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 75%, so baseline is 3, but the description adds real value beyond the schema: a concrete query example, the work_id transcript-join meaning, the project_id document listing behavior, and explicit filter precedence. Only 'limit' receives no added semantic context, but the schema already documents its constraints and default.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states a specific verb, resource, and scope: 'List active shared files in the caller's org, newest update first.' This clearly identifies what the tool does and is distinguishable from siblings like storage_list, file_get, and file_create.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage guidance is explicit: 'Use query to find a file by title without listing everything', 'Pass work_id to list documents + the agent transcript', and 'storage_list is canonical for every Storage join including blobs'. It even clarifies precedence ('project_id wins if both filters are set'), giving the agent concrete routing rules.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

file_publishAInspect

Publish a shared file: generate the public share token + slug and URL.

Idempotent: returns the existing token/slug if already published. The latest rendered HTML is eagerly pushed to the Railway bucket. The public URL is /s/{slug} (human-readable, from the title) with /s/{share_token} as a fallback; both are returned in the response.

Agent seats (org tsk_ / agent_run) are refused — call request_gated_approval(gate="publish") instead. Console humans (ts_session) still publish.

ParametersJSON Schema
NameRequiredDescriptionDefault
file_idYesFile UUID to publish

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals idempotency, eager pushing of rendered HTML to the Railway bucket, URL structure with a fallback, and access restrictions—all beyond what a schema would convey. It does not cover error handling (e.g., what happens if the file does not exist) or rate limits, but the core behaviors are transparent and align with the operation's nature.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is longer than a single sentence but well-structured: the first line states the core purpose, the second details idempotency and output mechanics, and the third addresses access control. Every sentence adds substantive value; there is no fluff. The pacing is efficient, though it could be tightened by moving the Railway bucket detail to the output schema or a separate note.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (multiple output URLs, gating, idempotency) and that an output schema exists, the description covers the most critical runtime context: what the agent will receive (both URLs), that it is safe to call repeatedly, and the exact gate alternative. It does not describe preconditions (e.g., file must exist and be in a publishable state) or error scenarios, but the provided information is sufficient for correct invocation in the common case.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the file_id parameter is already fully documented in the schema ('File UUID to publish'). The tool description adds no additional semantic detail about the parameter (e.g., format, validation, or edge cases). Since the schema fully covers the parameter, the baseline score of 3 is appropriate—the description does not need to repeat schema information.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise action: 'Publish a shared file: generate the public share token + slug and URL.' It clearly identifies the resource (shared file) and the output (token, slug, URL). It also distinguishes itself from other publishing tools (e.g., storage_publish) by detailing the specific token/slug mechanism and the /s/ URL format. This leaves no ambiguity about what the tool accomplishes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use this tool and when not to, naming the alternative: 'Agent seats (org tsk_ / agent_run) are refused — call request_gated_approval(gate="publish") instead. Console humans (ts_session) still publish.' It also implies usage context through idempotency, informing agents that repeated calls are safe and return existing tokens. This is explicit, actionable guidance that completely removes guesswork.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

file_updateAInspect

Append a new version to an existing shared file (version = prior max + 1).

If the file is published, the new version is eagerly mirrored to the Railway bucket. Old versions are never mutated.

ParametersJSON Schema
NameRequiredDescriptionDefault
agentNoOverride agent identity
contentYesNew file body (markdown, raw HTML, teamshared.diagram/v1 or teamshared.diagram/v2, or teamshared.sheet/v1 JSON/YAML)
file_idYesFile UUID to update
content_formatNoOverride content format ('markdown', 'html', 'diagram', or 'sheet'); defaults to the file's current format

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral burden. It discloses key traits: new versions are appended, version numbers increment deterministically, published files are eagerly mirrored, and old versions are never mutated. This is strong, though it omits permission requirements and error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the core action front-loaded and no filler. Every clause adds useful information about behavior or constraints.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists and parameter coverage is complete, the description covers the essential behavioral context: append semantics, version calculation, mirroring, and immutability. It does not discuss authorization or failure modes, but these are not critical for basic invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters. The description adds versioning context but does not materially explain individual parameters beyond what the schema states, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Append') and resource ('existing shared file') and clarifies the versioning semantics ('version = prior max + 1'). This clearly distinguishes it from siblings like file_create, file_publish, and file_get.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'existing shared file' implies this is for updating an already-created file rather than creating a new one, and the append-only behavior sets it apart from file_create. However, it does not explicitly name alternatives or state when not to use this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

file_upload_requestAInspect

Get a one-time uploader script to push a local file into a shared file.

For large local HTML/Markdown files that don't fit inline in file_create/file_update. Returns upload_url, upload_token, an expires_in_seconds TTL, and a self-deleting Python script. Save the script to disk and run python3 upload.py /path/to/file; it reads the file, POSTs it to the server with the one-time token, prints the resulting file id (and public URL if publish=true), and deletes itself on success. The token is single-use and expires in ~10 min.

Update mode: pass file_id to append the uploaded body as a new version to an existing shared file (the title is ignored; the existing file's title/slug/share_token are preserved, and the bucket mirror is re-published to the new version when the file is already published). This is the supported way to push a new version of a large file.

ParametersJSON Schema
NameRequiredDescriptionDefault
agentNoOverride agent identity
titleYesFile title (used only when creating a new file; ignored in update mode)
file_idNoExisting file UUID to append a new version to (update mode). Omit to create a new file.
publishNoIf true, the file is published immediately (returns public URLs). In update mode this is idempotent if already published.
work_idNoAttach the uploaded file to this work item UUID
filenameNoOptional filename (used for format sniffing and as the script's default path)
project_idNoAttach the uploaded file to this project UUID (kind=artifact)
content_formatNo'html', 'markdown', 'diagram' (teamshared.diagram/v1 JSON/YAML), 'sheet' (teamshared.sheet/v1 JSON/YAML), or 'auto' (sniff from the file extension; *.diagram.yaml/json, *.sheet.json/yaml)auto
upload_base_urlNoOptional server origin (e.g. https://teamshared.com). Defaults to settings.public_url.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burdencents for behavioral disclosure. It thoroughly describes the script's lifecycle (reads file, POSTs, prints id/URL, deletes itself on success), token security (single-use, ~10 min expiry), and update-mode side effects (preserves title/slug/share_token, re-publishes mirror). This goes well beyond a simple API call description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is somewhat long but well-structured with a main paragraph, a code-execution block, and a dedicated 'Update mode' section. Information is front-loaded with the core purpose, and every sentence contributes to usage or behavior. Slight verbosity in the script walkthrough prevents a perfect score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex tool with 9 parameters and an output schema, the description covers all essential aspects: when to use, full script execution flow, return values, update-mode semantics, and token lifecycle. Output schema already documents return fields, so the description needn't repeat them. No significant missing information for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds contextual nuance like 'title ignored in update mode' and the script's behavior related to publish, but these are also present in the schema's field descriptions. The description doesn't add significant new parameter semantics beyond what the schema already documents.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Get... push'), a resource ('a local file into a shared file'), and explicitly distinguishes itself from file_create/file_update by specifying its use case for large files that don't fit inline. This clearly separates it from siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use this tool ('For large local HTML/Markdown files that don't fit inline in file_create/file_update') and provides a specific alternative path for updating ('pass file_id'). It also labels update mode as 'the supported way to push a new version of a large file,' giving clear routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

healthAInspect

Liveness + dependency probe.

Returns {"status", "version", "components": {server, redis, postgres, semantic, distiller, graph, ollama}}. semantic is the pgvector + embedder store. Optional deps report "disabled" when off and do not degrade overall status. Always cheap; safe to poll on a 10s interval. Used by Docker healthcheck and the /health HTTP route.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden and handles it well. It discloses that optional dependencies report 'disabled' without degrading overall status, that the call is cheap, and that it is safe to poll frequently. These are meaningful behavioral traits beyond the bare input/output contract.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and well-structured: purpose first, followed by return shape, component semantics, failure behavior, and operational guidance. Every sentence adds useful information, and none repeats the input schema or annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter health probe with an output schema present, the description is complete. It explains the return structure, defines the ambiguous 'semantic' component, clarifies optional dependency behavior, and gives explicit polling guidance. An agent has everything needed to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so there is no parameter documentation burden. The description focuses on the response shape and component meanings instead, which is the appropriate use of description space given the empty input schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Liveness + dependency probe', which clearly states the tool's function, and then specifies exactly what it returns: status, version, and component health. This is distinct from the many sibling tools because it is explicitly the healthcheck/status probe rather than an account, memory, or work operation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear usage context: 'Always cheap; safe to poll on a 10s interval. Used by Docker healthcheck and the /health HTTP route.' This tells an agent when polling is appropriate, though it does not explicitly compare against alternatives such as `version` or state when not to use this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mcp_authAInspect

Last-resort email + OTP bind for this MCP session (headless only).

Prefer account-level Cursor Connect (Cloud / Grok Bot inherit it) or a tsk_ header. Do not call this as the first hop. When the host has no token:

  1. mcp_auth(email="you@example.com") — we email a 6-digit code.

  2. Ask the human for the code, then mcp_auth(email="you@example.com", code="123456").

  3. If status=need_org, call again with org_id=.

After status=authenticated, later tools on this streamable-HTTP session run as that person. Do not store the code or any token.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeNoOne-time code from the login email. Pass together with email= after mcp_auth returned status=code_sent.
emailNoEmail to send a 6-digit login code to. Same OTP as the console and Cursor Connect. Omit (with no code) to see whether this session is already signed in.
org_idNoOrganization to attach when the email belongs to more than one org (status=need_org).

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden and does substantial work: it discloses side effects (emails a 6-digit code), session-wide statefulness (later tools run as that person), the headless-only constraint, and a security rule (do not store code/token). It stops short of covering failure/retry behavior or expiry, but the key stateful and side-effectful traits are clearly surfaced.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence earns its place: purpose and constraint in line one, routing guidance in line two, and a numbered invocation sequence. The most important qualifier (last-resort, headless-only) is front-loaded, and the step numbering makes the multi-call flow easy to follow without padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema exists (so return values need not be spelled out) and 100% parameter coverage, the description completes the picture with the flow states, the conditional org branch, and post-auth session behavior. The only minor gap is failure mode handling (e.g., wrong code), but nothing needed to invoke the tool correctly across its documented states is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the parameter descriptions already explain email, code, and org_id semantics including the code_sent and need_org statuses. The description adds orchestration context (which order to pass parameters in), but per-parameter meaning is fully carried by the schema, matching the baseline-3 case for high coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource statement: 'email + OTP bind for this MCP session (headless only).' The 'last-resort' qualifier and 'headless only' constraint make its role unmistakable, and among the sibling list (org_bind, health, version) there is no competing auth tool, so an agent can identify it immediately.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells the agent when NOT to call it ('Do not call this as the first hop'), names the preferred alternatives (Cursor Connect, 'tsk_' header), and gives a concrete condition for when it is appropriate ('When the host has no token'). The numbered 3-step flow with the conditional need_org branch leaves no ambiguity about invocation order.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_agent_ensureAInspect

Get-or-create this org's ontology Agent + shared memory profile.

Harness agents slug to agent-<label>. Cloud agents slug to agent-<cursor_agent_id>. Other agents read the profile with memory_agent_get or memory_entity_view.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoHarness label (cursor, grok, dispatcher). Omit for the caller.
agentNoOverride agent attribution label
modelNoCursor model id when known (composer-2.5, auto-smart, …)
harnessNoHarness type that owns this agent (cursor, grok-bot, …)
runtimeNoharness (MCP client) or cloud (Cursor cloud agent)
spawned_byNoParent Agent slug or label (dispatcher that spawned this cloud agent)
cursor_agent_idNoCursor cloud agent id (bc-…). Sets runtime=cloud when provided.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It openly states the side effect (get-or-create), the org scope, and the exact slug derivation for harness vs. cloud agents. It does not say whether existing profiles are silently returned or updated, but 'Get-or-create' communicates the core non-destructive intent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the most important information: the operation and its resource. The slugging rules and read-alternative pointer are each one clause with no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema handles return values and the schema covers all 7 optional parameters, the description is largely complete. It misses a small but useful pointer to memory_agent_set for explicit updates, but the get-or-create semantics and read alternatives cover the main decision paths.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds value beyond the schema by explaining how parameters map to the agent slug (name to agent-<label>, cursor_agent_id to agent-<cursor_agent_id>), which is not derivable from the individual parameter descriptions alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise verb phrase 'Get-or-create' and a specific resource ('this org's ontology Agent + shared memory profile'), so an agent knows exactly what the tool does. The slugging convention distinguishes it from read-only siblings like memory_agent_get and memory_entity_view, and the 'ensure' semantics are immediately clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives actionable context: use this tool to get-or-create an agent profile, and read the profile with memory_agent_get or memory_entity_view. It does not explicitly say when to use memory_agent_set or memory_agent_list instead, but the get-or-create framing and named read alternatives provide decent routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_agent_getBInspect

Read another agent's (or your own) org-shared ontology memory.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoHarness label. Omit together with slug/id to fetch the caller.
slugNoOntology slug, e.g. agent-cursor or agent-bc-…
cursor_agent_idNoCursor cloud agent id (bc-…)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It only says 'Read' and 'org-shared ontology memory,' which implies a safe read operation, but it does not disclose what happens when multiple identifiers are provided, whether it errors on ambiguous inputs, or what the output structure is despite an output schema existing. It also does not clarify auth or sharing semantics beyond 'org-shared.' For a read tool with zero annotation coverage, this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, clear sentence with no filler, front-loading the core action and resource. It is concise and to the point, though it could add a little more context without becoming verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read tool with an output schema and fully documented parameters, the description is mostly sufficient, but it lacks behavioral guidance on parameter combinations and tool-specific nuances. The output schema covers return values, so that's handled. However, a clearer usage distinction from memory_agent_list or memory_recall would complete the picture. Minimum viable but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters (name, slug, cursor_agent_id) with their roles. The description adds minimal extra meaning beyond saying 'read memory' and doesn't explain how the parameters interact (e.g., 'omit together with slug/id to fetch the caller' is already in the schema). Baseline 3 is appropriate since the description doesn't contradict or significantly extend the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description clearly states the action ('Read') and resource ('another agent's (or your own) org-shared ontology memory'), which is specific enough. It is distinct from siblings like memory_agent_list (which likely lists memories) and memory_agent_ensure/set (which write), though it doesn't explicitly name them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for reading ontology memory, and the parameter docs hint that omitting identifiers fetches the caller, giving a sense of when to use defaults. However, it does not explicitly say when to prefer this over siblings like memory_recall or memory_agent_list, nor does it provide exclusions or routing guidance. This is implied by the resource name but not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_agent_listAInspect

List Agent ontology memories in this org (other agents can reference them).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
offsetNo
runtimeNoOptional filter: harness or cloud

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the behavioral disclosure burden. 'List' conveys that this is a read-only operation and 'in this org' sets the scope, but nothing is mentioned about ordering, pagination behavior, or any side effects. These are minor for a simple list tool, but the description adds little beyond what the name and schema already imply.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler: the verb, resource, and scope come first, and the parenthetical about agent references earns its place by clarifying why the list matters. Nothing extraneous is present.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple, read-only list tool with no required parameters and an existing output schema, the description covers the essential context: what is being listed and in which org. It does not explain the runtime filter or result ordering, but the schema already documents runtime and the output schema covers return structure, so the gaps are minor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33% (runtime is described as 'harness or cloud'), and the description provides zero parameter guidance. Limit and offset are left to their self-evident names and defaults; since coverage is below 50%, the description should have compensated, but it does not. The agent must infer the pagination filter semantics from the schema alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description uses a specific verb ('List'), an explicit resource ('Agent ontology memories'), and a scope ('in this org'), while adding the functional note that other agents can reference them. This clearly distinguishes it from sibling tools like memory_agent_get (single fetch) and memory_agent_set (write).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The context is clear: use this to get a directory of Agent ontology memories for the current org, especially when you need identifiers other agents can reference. It does not explicitly enumerate when not to use it or name alternatives, but the listing semantics and org scope are sufficient for an agent to select it over the single-entity siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_agent_setAInspect

Replace an agent's org-shared memory profile (creates the entity if needed).

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoHarness label. Omit with slug/id to write the caller.
slugNoOntology slug, e.g. agent-cursor
agentNoOverride agent attribution label
body_mdYesCompressed markdown profile (role, current task, constraints). Keep short; server caps length. Org-visible — not a private soul.
cursor_agent_idNoCursor cloud agent id (bc-…)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry behavioral disclosure; it does communicate that this is a mutating, upsert-style operation (replace or create). It does not cover permission requirements, reversibility, or consequences for existing memory, so transparency is adequate but incomplete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One compact sentence states the action, scope, and key edge case (creates if needed) with no filler. The parenthetical carries meaningful behavioral information without bloating the description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a five-parameter mutation tool with an output schema, this is minimally viable: schema covers parameters and return values, and the description covers the essential operation. It still lacks guidance on alternatives and exclusions (e.g., private vs. org-shared scope boundary is only in the schema), so it is not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the structured parameter docs already explain every argument. The tool description adds no parameter-level detail beyond the core concept, which is fine under the baseline for full schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Replace') with a clear resource ('an agent's org-shared memory profile') and adds the upsert behavior ('creates the entity if needed'). This distinguishes it from read-only siblings like memory_agent_get and memory_agent_list, and from ensure-style semantics, making the tool's role unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use it: when an agent's shared memory profile should be replaced or created. However, it does not explicitly state when not to use it or name alternatives such as memory_agent_ensure, leaving the routing decision partly to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_assemble_contextAInspect

Assemble one token-budgeted, cited context pack for a task.

Fans recall across semantic, episodic, procedural, skill, strategic, work, working pillars and the optional graph in parallel through the secure retrieval path, then ranks and packs the result into a single sectioned markdown bundle. Use this once at the start of a task instead of issuing serial memory_recall / memory_procedure_get / memory_graph_related calls. Returns rendered (the pack), tokens_used, counts_by_pillar, and the kept records.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoNoWorkspace slug of your current repo. Boosts repo-scoped memories in the pack; pass it for code/repo-specific work.
taskYesWhat you are about to do (the task/question driving recall)
githubNoGitHub repository as owner/repo (boosts github-tagged memories)
open_filesNoPaths of files currently open/relevant; their names seed the graph-relationship lookup.
k_per_pillarNoMax records to recall per pillar
token_budgetNoApprox token budget for the rendered pack (default 1500)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses that the tool recalls across pillars in parallel, ranks and packs results, and returns specific fields. However, it does not explicitly state whether it is read-only or has side effects. The 'recall' language implies read-only, but without annotation support, a clear statement about non-mutation would be ideal. This minor gap prevents a perfect score.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with a clear first sentence stating the main action, followed by a detailed explanation of the recall mechanism and a usage note. It is somewhat verbose but every sentence contributes meaning. It is front-loaded with the core purpose, making it easy for an agent to grasp quickly. Slightly tighter wording could earn a 5, but it remains efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the tool's core function, usage timing, and return fields (rendered, tokens_used, counts_by_pillar, records). Given that an output schema exists, the return format is likely specified there. The description does not discuss error conditions or prerequisites, but for a context-assembly tool with 6 parameters and optional graph support, the provided information is sufficient for correct invocation. A mention of potential failure modes would make it more complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 100% parameter description coverage, so the baseline is 3. The tool description does not add semantic value beyond what the schema already provides. It mentions the 'task' indirectly and refers to the pack, but does not elaborate on parameters like 'repo', 'github', or 'open_files' in a way that would improve understanding. It meets the baseline but does not exceed it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Assemble one token-budgeted, cited context pack for a task.' It clearly defines the output as a sectioned markdown bundle and differentiates itself from sibling tools by naming the alternative serial calls (memory_recall, memory_procedure_get, memory_graph_related). The purpose is unambiguous and distinguishable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use the tool: 'Use this once at the start of a task instead of issuing serial memory_recall / memory_procedure_get / memory_graph_related calls.' This directly instructs the agent on the appropriate invocation context and names the alternatives to avoid, which is exemplary usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_changes_sinceAInspect

Resume-handoff delta of durable org memory since an opaque cursor.

Returns current-truth rows newer than cursor (oldest-first) plus a new opaque cursor. Use this to continue another session's work; use memory_recall for keyword search. Not a full checkpoint/context pack.

ParametersJSON Schema
NameRequiredDescriptionDefault
kNoMax durable changes to return (oldest-first)
kindNoOptional durable kind filter (fact|preference|event|note|outreach)
repoNoWorkspace slug filter. When set, only memories tagged for this repo are returned (unlike memory_recall, which boosts).
cursorNoOpaque resume watermark from a previous memory_changes_since call. Clients must persist and return it as-is — do not parse or mint one. Omit to take a bootstrap page of the newest current-truth durable rows and receive a cursor to watch from.
githubNoGitHub owner/repo filter. When set, only memories tagged github:<owner>/<repo> are returned.
pillarNoOptional pillar filter. This feed is memory_items only (semantic / episodic). Other pillars return an empty page.
include_supersededNoWhen true, also return superseded/merged history. Default is current truth only, same as memory_recall.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It discloses the output ordering (oldest-first), the fact that it returns a new opaque cursor, and the limitation that it is not a full checkpoint. It implies a read operation (returns delta), but does not explicitly state read-only or describe side effects. It adds useful context beyond the schema (e.g., resume-handoff, current-truth), so it is strong but not perfect.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, front-loaded with the core purpose. It packs a clear purpose, return behavior, usage guidance, and a limitation without any fluff. Every sentence contributes to orientation or selection.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The presence of an output schema covers return format. The description covers why and when to use the tool, plus a key limitation. It does not explain the effect of filters (kind, repo, pillar) but those are fully documented in the schema. For a tool of this complexity, this is adequately complete, though it could mention that only memory_items are returned for pillar filtering, which is in the schema but not the description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description does not add semantic detail about individual parameters beyond what the schema already provides. It mentions 'opaque cursor' and 'current-truth' conceptually, but these are also covered in the schema descriptions. No extra value added.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action: providing a resume-handoff delta of durable org memory since an opaque cursor, returning current-truth rows newer than the cursor. It explicitly distinguishes itself from memory_recall for keyword search, making its purpose unambiguous to an agent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives direct usage instruction: 'Use this to continue another session's work' and contrasts with 'use memory_recall for keyword search.' It also flags a limitation ('Not a full checkpoint/context pack'), telling the agent when not to use it. This is explicit and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_context_packAInspect

Bounded resume pack for an agent starting or resuming work.

One call returns {cursor, profile_excerpt, work_open, changes, has_more_changes, recall, truncated, token_est, budget_tokens} packed to budget_tokens in that fixed section order. Built from the same paths as memory_profile_get, memory_changes_since and memory_recall (shared brain: no agent filter). Persist cursor and pass it next time to receive only newer changes. Use memory_recall for further search and memory_assemble_context for a task-driven pack.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoNoWorkspace slug. Filters the changes section to this repo and boosts (does not filter) recall hits.
queryNoOptional short keyword anchor. When set, adds a capped, org-shared memory_recall section. Omit to skip recall.
cursorNoOpaque resume cursor from a previous memory_context_pack, memory_changes_since, or a closing context_commit / memory_session_close (handoff_cursor). Pass it back as-is. Omit on first contact to get the newest durable changes plus a cursor to resume from.
githubNoGitHub owner/repo. Filters changes to github:<owner>/<repo> and boosts recall hits.
work_idNoWork item id to include as work_open (compact row)
budget_tokensNoApprox token cap for the whole pack (default 3000)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral disclosure. It states the tool is read-oriented (returns a pack, no agent filter), mentions a 'truncated' flag, and explains cursor semantics. It does not explicitly say it is read-only or describe side effects, but it is implied and sufficient for a context-fetching tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured paragraph that front-loads the purpose and then explains output, source, cursor usage, and alternatives. It is dense but not verbose; every sentence contributes. It is appropriately concise for the complexity of the tool.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the tool's main behavior, cursor mechanics, and relationship to siblings. Since an output schema exists, the return value details are covered. It could mention edge cases like minimum budget or what happens when the budget is exceeded, but it is largely complete for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds some context beyond the schema, such as the fixed section order and token budgeting, and reinforces cursor usage, but it does not significantly expand on individual parameter meaning beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: a 'Bounded resume pack for an agent starting or resuming work.' It distinguishes itself from siblings by naming specific alternatives (memory_profile_get, memory_changes_since, memory_recall) and explicitly says it is built from the same paths, making it clear what it does and how it differs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance: 'Use memory_recall for further search and memory_assemble_context for a task-driven pack.' It also explains the cursor pattern: 'Persist cursor and pass it next time to receive only newer changes,' which tells the agent exactly how to use it incrementally.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_dream_statusAInspect

Latest nightly dream report for this org (what changed overnight).

Returns the most recent leftover-distill + curator report, or found=false when last night wrote nothing. The same note is a normal semantic memory / wiki page memory_recall can find.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral burden and does well by disclosing the found=false fallback when no report was written and clarifying that the content is stored as an ordinary semantic memory page. It does not cover auth, rate limits, or explicit read-only guarantees, but for a parameterless report getter the disclosed behavior is adequate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded, with the key purpose in the first sentence and supporting behavioral details in the next two. Every sentence adds distinct value: scope, return behavior, and relationship to memory_recall. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has no parameters, an output schema is present, and the behavior is simple read-only retrieval, the description covers everything needed to call it correctly. It explains what is returned, what happens when nothing was written, and how the note relates to the broader memory system.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and the schema description coverage is 100%, so there is no parameter information for the description to meaningfully add. The baseline for zero-parameter tools is 4, and the description appropriately focuses on behavior rather than parameter semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb-object pair: 'Latest nightly dream report for this org (what changed overnight)', clearly identifying the resource and scope. It further distinguishes itself by noting the report is also a normal semantic memory page that memory_recall can find, which positions this tool as the dedicated status view rather than a generic recall tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context: use this when you need the most recent nightly dream/change report for the org. It also mentions that the same note is discoverable via memory_recall, which indirectly signals a related alternative, though it does not explicitly state when to prefer one over the other.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_entity_viewAInspect

Roll up wiki, memories, graph neighbors, and work for one entity.

ParametersJSON Schema
NameRequiredDescriptionDefault
slugYesEntity slug (wiki topic slug or ontology entity slug)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description is the only safety/behavior signal. 'Roll up' strongly implies a non-mutating aggregation, but the description does not explicitly state read-only behavior, required permissions, freshness, or pagination; the tool name 'view' helps but is not a substitute for annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with no filler; the action, data sources, and scope are all present and front-loaded. Nothing in the description is redundant with the input schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter tool with a rich output schema, the description says enough about what the caller gets (aggregation of four data categories) to invoke it correctly. It omits deeper selection guidance, but that is already scored under usage guidelines.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the slug parameter is already fully documented as a wiki topic slug or ontology entity slug. The description adds no additional parameter-level meaning beyond reinforcing that the operation is scoped to a single entity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a concrete action ('Roll up') and names the exact resource types included (wiki, memories, graph neighbors, work) plus the entity scope. It conveys a broad aggregated entity view, though it does not explicitly contrast with siblings like memory_recall or memory_assemble_context.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for one entity' implies the intended use case: get an aggregated view of a single entity's related data. However, it never describes when to prefer this over the many memory-looking siblings, nor does it state exclusions or prerequisites beyond having a slug.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_feedbackAInspect

Record +1/-1 feedback on a recalled memory id.

Append-only: writes memory.feedback on the memory:{id} stream. Does not change the memory body — corrections stay on memory_remember (supersede) or memory_forget. Skill/playbook rewrites stay on memory_skill_feedback.

ParametersJSON Schema
NameRequiredDescriptionDefault
noteNoOptional one-line why this vote
voteYes1 if this recall helped, -1 if it hurt
agentNoOverride agent identity
memory_idYesmemory_items UUID from a previous memory_recall hit

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral burden. It discloses that the operation is append-only, writes to the memory:{id} stream, and leaves the memory body unchanged, which is important context beyond the tool's name. It could add details about duplicate votes or failure behavior, but the output schema covers return shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: the first sentence states the action, and the second paragraph adds only essential behavioral guardrails. There is no filler, redundant schema restatement, or vague introductory language.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 4-parameter tool with a full output schema and no annotations, the description covers what the tool does, how it behaves, and which sibling tools handle related cases. An agent has enough information to invoke it correctly without guessing about side effects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and each parameter already has a meaningful description: vote is an enum with semantics, memory_id is traced to a prior memory_recall, and note/agent are optional overrides. The description reinforces that memory_id refers to a recalled memory but adds little new parameter-level meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Record +1/-1 feedback on a recalled memory id', naming a specific action, object, and allowed value range. It also distinguishes itself from nearby memory tools by stating it does not modify the memory body and by pointing corrections and rewrites to memory_remember, memory_forget, and memory_skill_feedback.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly tells the agent when this tool is not the right choice: corrections belong on memory_remember/memory_forget, and skill/playbook rewrites belong on memory_skill_feedback. It clearly implies use after a memory_recall hit when the goal is only to record a positive or negative vote.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_playbook_getDInspect

Alias for memory_procedure_get.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesPlaybook name
versionNoSpecific version
expand_skillsNoInline composed skills

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

D1.4/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of disclosing behavior, but it discloses nothing beyond being an alias. No mention of side effects, read/write characteristics, or any other behavioral traits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short, but that conciseness comes at the cost of any useful information. It is not structured or front-loaded with actionable details; it is a single placeholder sentence that fails to earn its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Even with an output schema present, the description is entirely inadequate. It doesn't explain what the tool does, when to use it, or what the agent can expect. For a tool with this level of complexity, it is completely incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the input schema thoroughly documents all three parameters and their types/defaults. The description adds no extra meaning, but per the rubric, a baseline of 3 is appropriate when schema already covers the parameters fully.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose1/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description only states that it is an alias for memory_procedure_get, without any explanation of what the tool actually does. It fails to identify the verb, resource, or functionality, making it impossible for an agent to understand the tool's purpose without looking up the referenced tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus any alternative, nor any mention of scenarios where it should be preferred or avoided. The description offers no usage context whatsoever.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_playbook_setDInspect

Alias for memory_procedure_set.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesPlaybook name (stable id)
tagsNoTags for discovery
agentNoOverride agent identity
owner_idNoOrg member UUID. Omit inherits last owner or the writer.
steps_mdNoOptional intro markdown before composed skills
descriptionNoOne-line summary
tool_recipeNoOrdered skill list: {"skills": ["lint", "ship-pr"]}
verification_daysNoVerification window: 30, 90, or 180 days; 0 for none. Omit inherits the last window.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

D1.7/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals nothing about side effects, ownership inheritance, verification window semantics, or return behavior. Calling it an alias is renaming, not explanation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The one-sentence description is short, but this is under-specification rather than efficient writing. It spends its only sentence on the alias instead of anything actionable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter tool with no behavioral description, the definition is incomplete. The output schema helps with return shape, but the agent is still missing what the tool does, its side effects, and when to use it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline applies. The description adds nothing about parameters, but the schema fully documents each property, including owner_id inheritance and verification_days options.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description only identifies the tool as an alias for `memory_procedure_set`; it never states the operation or resource. An agent must already know what `memory_procedure_set` does, so the purpose is only indirectly implied.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance is provided, no comparison to siblings like `memory_playbook_get` or `memory_playbooks_list`, and no exclusions. The alias statement offers no decision support for choosing this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_playbooks_listDInspect

Alias for memory_procedures_list.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagNoFilter by tag
limitNo
offsetNo
include_bodyNoInclude full steps_md

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

D1.3/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, and the description discloses no behavioral traits (read-only, destructive, permissions, etc.). 'Alias' gives no information beyond implying identical behavior to memory_procedures_list, which is itself unspecified.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is extremely short but not effectively concise—it omits essential information rather than presenting a focused, structured explanation. It is under-specified rather than efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness1/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 4 parameters and an output schema, the description provides almost no context: no functionality, no usage scenarios, no differentiation from siblings. It is far from complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema documents only tag and include_body (50% coverage), leaving limit and offset undescribed. The description adds nothing about parameters, failing to compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description only says 'Alias for memory_procedures_list' without stating what the tool actually does. It implies shared functionality but leaves the agent to look up the aliased tool, so the purpose is not directly clarified.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines1/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is zero guidance on when to use this tool versus the aliased memory_procedures_list or any of the many other memory_* siblings. The description does not mention context, conditions, or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_profile_getAInspect

Return a token-capped bootstrap profile for this caller.

Compact markdown (soul when the key is account-linked, a few high-signal durable facts preferring kind=preference, and a short recent episodic tail). Empty/minimal when the org has nothing to inject. Also returned on memory_session_ensure as profile and prepended on agent_run_start when non-empty. Does not dump full recall or the skills catalog. namespace= excludes facts/episodes outside that container.

ParametersJSON Schema
NameRequiredDescriptionDefault
namespaceNoWhen set, profile facts/episodes are limited to this org-allowlisted container (exclude other namespaces). Omit for the shared-org digest. Invalid/unknown slugs fail.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It thoroughly explains the output format (compact markdown, soul, preference facts, episodic tail), conditions for empty/minimal results, its appearance in other calls, and its exclusion of full recall/skills. It also discloses namespace filtering behavior. This is exemplary transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the primary purpose, then delivers compact, high-value details about content, edge cases, and cross-tool behavior. Every sentence contributes, and the prose is dense without being verbose. Structure and length are optimal for an agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists (signal indicates yes), the description need not explain return structures. It covers what the profile contains, when it may be empty, how it relates to other calls, and the parameter's effect. No critical information for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already provides 100% coverage for the single 'namespace' parameter, including its meaning and failure behavior. The description adds a redundant phrase ('namespace= excludes...') but no new information beyond the schema. Per the baseline rule, when schema coverage is high, a 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Return a token-capped bootstrap profile') with a clear resource (the caller's memory profile). It also differentiates itself from siblings by explicitly noting it 'Does not dump full recall or the skills catalog', which distinguishes it from memory_recall and memory_tools_catalog. The purpose is unambiguous and actionable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides context on when the tool is used ('Also returned on memory_session_ensure... prepended on agent_run_start'), which implies it's for bootstrapping sessions. It also hints at alternatives by stating what it does not do, but lacks explicit 'use this when' or 'use that instead' guidance. It's clear but not fully prescriptive about selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_recallAInspect

Hybrid recall across memory pillars within the caller's org.

Default scope searches durable pillars only (semantic, episodic, procedural, skill, strategic, work). Pass scope=["working"] to include this chat's open session turns. Shared brain on durable pillars: pass agent="cursor" only to narrow semantic/episodic. For entity/competitor questions use a short keyword anchor in query (e.g. "mex") plus repo / github. Use explain=true; prefer hits with matched_keyword: true. Default recall is current truth (superseded/merged rows are omitted); pass include_superseded=true for the replacement chain. When namespace= is set, semantic/episodic hits outside that container are excluded. Hits include a small provenance citation block (id, kind, timestamps, optional title, optional source) — metadata only, never page HTML. Result includes org + bound.

ParametersJSON Schema
NameRequiredDescriptionDefault
kNoMax records to return
repoNoWorkspace slug of your current repo. When set, durable memories tagged for this repo are boosted (ranked higher); nothing is hidden — cross-repo and un-scoped memories still appear. Pass your workspace slug when recalling for code/repo-specific work.
agentNoOptional filter — restrict semantic/episodic results to this agent's writes. Default (None) is the shared brain: every agent's durable memories in the org are visible. Working memory is always scoped to the caller regardless.
queryYesNatural-language query
scopeNoPillars to search. Default (null): durable pillars only (semantic, episodic, procedural, skill, strategic, work) — working is omitted. Add scope=['working'] when you need this chat's open session turns.
githubNoGitHub repository as owner/repo. Boosts memories tagged github:<owner>/<repo> (portable across checkout paths).
explainNoWhen true, include per-record retrieval attribution in metadata.ranking (base RRF, every factor including bounded feedback trust from memory_feedback votes, relevance floor verdict, diversity penalty; its score equals the hit's score), plus read-only feedback_score when votes exist
filtersNoOptional AND filter. Keys (all optional): kind (fact|preference|event|note|outreach|skill|procedure), pillar (semantic|episodic|procedural|skill|strategic|work), subject (exact, case-insensitive), tags (list; record must include every tag), since / until (ISO datetimes). Example: {"kind": "fact", "subject": "HolderBrief", "tags": ["decision"]}. Nested AND/OR/NOT is not supported — omit a key instead. Empty result retries once without the filter (filter_relaxed=true on the result). Not for named playbook/skill/entity — use get-by-name tools.
verboseNoWhen false, truncate record content and omit metadata
namespaceNoOrg-allowlisted container. When set, semantic/episodic hits outside this slug are excluded (not merely demoted). Omit to keep shared-org recall. repo=/github= still boost. Unknown slugs fail. Example: github:acme/app or project:alpha.
time_rangeNoOptional time bounds for episodic/working hits
deadline_secondsNoOptional wall-clock budget for this call. Finished pillars are returned with errors_by_pillar timeout codes and degraded=true instead of letting the MCP client kill the whole call. Omit to use the server TEAMSHARED_RECALL_DEADLINE_SECONDS default (unset = no cutoff).
include_supersededNoWhen true, also return superseded/merged history. Default is current truth only.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the full burden. It discloses defaults (durable-only scope, current-truth semantics, shared-brain visibility), scoping exclusions (namespace), and the provenance citation block (metadata only, never page HTML), plus result fields (org + bound). It doesn't cover failure modes or rate limits, but for a read-style recall it is substantially transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Around five dense sentences front-load the purpose, then alternate pithy parameter-guidance sentences. Each clause earns its place and there is minimal fluff. Some default-behavior statements mirror schema descriptions, but given the 13-param complexity the repetition is acceptable guidance, not padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The package is strong: detailed per-param descriptions cover the rest (filters, deadline_seconds, k, verbose) and an output schema exists. The description covers the non-obvious semantics (current-truth default, provenance shape, namespace exclusion, working-memory scope) that an agent must know to invoke it correctly. Nothing critical for selection or invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds strategy beyond the schema: entity/competitor queries warrant a short keyword anchor with repo/github, explain=true should be used, and matched_keyword hits are preferred. It also clarifies shared-brain vs agent='cursor' behavior. This pushes it a step above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening line 'Hybrid recall across memory pillars within the caller's org' gives a specific verb (recall), a clear resource (memory pillars), and an org-scoped target. The follow-up enumerates the six durable pillars and the working-memory option, which distinguishes it from other memory_* tools without naming them. It stops short of a crisp one-line contrast with siblings such as memory_entity_view or memory_think, so 4.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear when-to-use tips: pass scope=['working'] for this chat's open session turns, agent='cursor' to narrow semantic/episodic, include_superseded for the replacement chain, and namespace to constrain the container. It also gives query-shaping advice (short keyword anchor + repo/github for entity/competitor questions, explain=true, prefer matched_keyword). It does not explicitly say when not to use this tool in favor of a sibling, so no exclusion statement is present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_rememberAInspect

Write a durable memory into the caller's org.

fact / preference / note / outreach -> semantic pillar. event -> episodic. procedure / skill -> rejected; use memory_procedure_set / memory_skill_set. Routed through the guarded ingestion pipeline (PII, injection screening, near-dupe merge / contradiction supersede) under RLS. When repo / github are given the memory is tagged repo:<slug> / github:<owner>/<repo>. namespace= (or the default from those args) is an org-allowlisted container — recall with the same slug excludes other namespaces. Result includes org + bound.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNofact, preference, event, note, or outreach (not procedure/skill)note
repoNoWorkspace slug of the repo this memory belongs to (e.g. the slug used for memory_state). For code/repo-specific work, pass your current workspace slug so the memory is scoped to this repo (stored as a 'repo:<slug>' tag) and ranks higher when recalled from the same repo. Omit for cross-cutting memories.
tagsNoOptional free-form tags
agentNoOverride agent identity (defaults to bearer-token identity)
githubNoGitHub repository as owner/repo (e.g. teamshared-ai/teamshared). Stored as a 'github:<owner>/<repo>' tag for cross-machine association; use with or instead of workspace repo= when the same GitHub repo is checked out at different paths.
contentYesFree-form text to remember
subjectNoOptional subject/entity this memory is about
namespaceNoOrg-allowlisted container (repo:owner/name, github:owner/repo, or project:name). Tags this write. Omitted: default from github= then repo=. Invalid slugs fail. First write creates the slug on the org allowlist.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description fully carries behavioral disclosure. It mentions the guarded ingestion pipeline (PII, injection screening, near-dupe merge / contradiction supersede) under RLS, and explains tagging and namespace behavior. This is substantial transparency beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well-structured, leading with the core purpose, then kind mapping, pipeline, and tagging. It is not overly verbose and avoids redundancy, though it could be slightly tightened without losing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (8 parameters, output schema exists), the description covers all critical aspects: what it does, kind handling, pipeline behavior, tagging, namespace scoping, and result fields. It is complete enough for an agent to use correctly without further investigation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so parameters are already documented. The description adds semantic meaning (kind→pillar mapping, repo/github tagging behavior, namespace allowlisting) that goes beyond the schema descriptions, providing extra context for parameter selection.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool writes a durable memory into the caller's org, and explicitly distinguishes it from memory_procedure_set and memory_skill_set for procedure/skill kinds. It also lists the accepted kinds, leaving no ambiguity about what this tool does versus its siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit guidance: procedure/skill are rejected and the user is directed to memory_procedure_set / memory_skill_set. It also implies the tool is for semantic and episodic facts, preferences, events, notes, and outreach, which is a clear usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_session_appendAInspect

Append a turn to a working-memory session (self-healing).

When session_id has expired or was closed, a fresh session is opened automatically and the turn lands there; the response then carries the replacement session_id and reopened: true. Pass repo (and optionally github / topic) so reopen preserves workspace scope and updates the state pointer without a manual memory_state_set.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoNoWorkspace slug. When set with an active bearer token, reopen self-healing updates the conversation/active-session pointer.
roleYesuser, assistant, tool, or system
topicNoSession topic used when reopening after expiry
githubNoGitHub owner/repo tag used when reopening a session
contentYesTurn content
session_idYesSession id from memory_session_open

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose behavior. It explicitly describes the self-healing mechanism: if session_id expired/closed, a fresh session opens and the response carries a new session_id and reopened:true. It also explains how repo/github/topic affect scope. This covers the main side effects, though it doesn't mention any potential data loss on the old session or auth/rate limits, which are not critical here.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the primary action. The first sentence is the core purpose; the second expands on self-healing and parameter usage. No wasted words; every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values are already covered. The description explains the self-healing behavior and parameter effects, which is sufficient for an agent to call this tool correctly. It doesn't define 'turn' explicitly, but role and content parameters make it clear. The description is complete for practical use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

All 6 parameters have schema descriptions (100% coverage), so baseline is 3. The description adds meaningful context beyond the schema: it explains that repo 'preserves workspace scope and updates the state pointer', and that topic/github are 'used when reopening after expiry'. This extra context helps the agent understand the parameters' purpose, raising the score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource: 'Append a turn to a working-memory session'. This clearly states the tool's action and distinguishes it from sibling tools like memory_session_open (which opens a new session) or memory_session_close (which closes one). The 'self-healing' hint further clarifies its niche.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use this tool implicitly by describing the self-healing behavior for expired/closed sessions and how to preserve scope via repo/github/topic. It also mentions avoiding a manual memory_state_set, which hints at an alternative. However, it doesn't explicitly compare with memory_session_open or state when NOT to use this tool, though the core usage is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_session_closeAInspect

Close a working-memory session.

If distill is true (default), the transcript is queued for the background worker to summarize into durable org-scoped memories.

ParametersJSON Schema
NameRequiredDescriptionDefault
distillNoEnqueue for distillation into semantic/episodic memory
session_idYesSession id to close

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral disclosure burden. It reveals an important side effect: if distill defaults to true, the transcript is queued for a background worker to create durable org-scoped memories. It does not cover session finality, reversibility, or failure behavior, but the key side effect is transparently disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler: the first states the core action, the second explains the optional distill behavior. It is front-loaded and every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity tool with full schema coverage and an output schema, this description is complete. It covers the required input, the optional flag, and the important side effect, so an agent can invoke the tool correctly without needing additional context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaningful detail for the distill parameter, explaining the background worker and durable org-scoped memory outcome beyond the schema's 'semantic/episodic memory' phrasing, and it clarifies the default behavior. session_id is adequately documented by the schema itself.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states a specific action and object: 'Close a working-memory session.' This clearly identifies what the tool does and contrasts with sibling tools like memory_session_open and memory_session_append through the verb 'Close'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description does not explicitly say when to use this tool versus alternatives such as memory_session_open or memory_session_append. The only conditional guidance is about the distill flag, not about choosing this tool, so usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_session_ensureAInspect

One-call session bootstrap: recover the active session or open one.

Replaces the memory_state_get → memory_session_close → memory_session_open → memory_state_set ritual. Reuses the session in the conversation/active-session state pointer when it is still open and owned by the caller; otherwise closes it (distilling) and opens a fresh one, updating state. Returns {session_id, agent, resumed, soul, soul_linked, profile}. When the bearer is linked to a human account, soul is their private compressed identity block for this org (may be empty string if not yet written). profile is the same token-capped bootstrap object as memory_profile_get. When playbook_slug or a work_id with a bound slug resolves, also returns playbook {name, description, body_md, slug} (its cited skills listed by name and when to load them with memory_skill_get, not their bodies; capped) on both fresh and resume paths. Also org + bound + bound_scope. On unbound /mcp, a user account with several orgs may include warnings suggesting org_bind(slug=...) (CLI teamshared org bind is optional). Optional auto_recall=true attaches budgeted recall hits (same provenance citation shape as memory_recall) when user= and/or topic= is set (thin-client bootstrap). Prefer an explicit memory_recall for keyword anchors.

ParametersJSON Schema
NameRequiredDescriptionDefault
ttlNoSession TTL in seconds (default from server config)
repoYesWorkspace slug (absolute path with leading / removed and / replaced by -). Keys the conversation/active-session state pointer.
userNoSubstantive user request for this turn. When set, appended as the user turn in the same call (replaces a separate memory_session_append).
agentNoOverride agent identity
freshNoForce rotation: close any stored session (queueing distillation) and open a new one. Pass true on the first turn of a new chat.
topicNoWhat this session is about (used when opening a new one)
githubNoGitHub owner/repo; distilled memories inherit the tag
work_idNoWork item UUID. When set and playbook_slug is omitted, that item's playbook_slug is injected on the ensure payload.
namespaceNoLimit the attached profile facts/episodes to this org-allowlisted container. Omit for shared-org profile. Invalid/unknown slugs fail.
auto_recallNoOpt-in thin-client bootstrap: when true and user= and/or topic= is set, run a budgeted memory_recall (k=5, explain=false, short deadline) over that text and attach hits as auto_recall. Missing both skips without error. Default false — existing callers are unchanged. Prefer an explicit memory_recall with a short keyword anchor; this is not a replacement.
playbook_slugNoNamed playbook to attach (wins over the work item's slug). Unset/missing/cross-org omits playbook — never dumps the catalog.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It thoroughly explains side effects: it recovers an existing session if open and owned, otherwise closes it (distilling) and opens a fresh one, updating state. It also details return fields, including playbook, org, bound, and warnings, and describes the auto_recall behavior. Nothing about its mutating nature or side effects is hidden.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence carries information. It front-loads the core purpose, then systematically covers replacement of the ritual, return values, optional behaviors, and parameter nuances. No wasted words; the length is justified by the tool's complexity (11 parameters).

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity, the description is comprehensive. It covers the return object shape, side effects, usage guidance, alternatives, and optional parameters like auto_recall and playbook_slug. The presence of an output schema further reduces the need to explain return structure, but the description already does. Nothing an agent needs to call this tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all parameters. The description adds meaningful usage context beyond the schema: it explains that fresh=true forces rotation for a new chat, that playbook_slug wins over work_id's slug, that auto_recall only runs when user/topic is set, and that repo keys the active-session pointer. This adds value, though the schema already provides baseline definitions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action: 'One-call session bootstrap: recover the active session or open one.' It clearly names the resource (session) and the operation, and distinguishes itself from sibling tools like memory_session_open and memory_session_close by explicitly replacing a multi-step ritual. It is unambiguous about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly explains when to use this tool versus alternatives: it replaces the memory_state_get → memory_session_close → memory_session_open → memory_state_set ritual, and advises preferring an explicit memory_recall for keyword anchors. It also gives concrete guidance like passing fresh=true on the first turn of a new chat, and clarifies that auto_recall is opt-in for thin-client bootstrap, not a replacement for memory_recall.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_session_openCInspect

Open a working-memory session and return a session_id.

ParametersJSON Schema
NameRequiredDescriptionDefault
ttlNoSession TTL in seconds (default from server config)
repoNoWorkspace slug of the repo this session is about. Memories distilled from the session inherit a 'repo:<slug>' tag so they stay scoped to this workspace.
agentNoOverride agent identity
topicNoWhat this session is about (free text)
githubNoGitHub repository as owner/repo. Distilled memories inherit a 'github:<owner>/<repo>' tag.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the full burden of behavioral disclosure. It only says a session is opened and an ID is returned, without explaining session lifecycle, TTL effects, whether opening is destructive or idempotent, or what happens to prior sessions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler or repetition. It is concise and readable, though it is so sparse that it sacrifices useful contextual detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The action is simple, the output schema exists, and there are no required parameters, so the description covers the core call. However, it lacks lifecycle context, such as how the returned session relates to other memory_session_* operations and whether the session needs to be closed or can be reused.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, and each parameter already has a clear meaning in the input schema, including TTL, repo tagging, GitHub tagging, agent override, and topic. The description adds no parameter-level detail, but because the schema is fully self-documenting, the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ("Open"), a clear resource ("working-memory session"), and the return value ("session_id"). It is understandable on its own, but it does not explicitly differentiate from sibling tools like memory_session_ensure or memory_session_append.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to call this tool versus alternatives such as memory_session_ensure, memory_session_append, or memory_session_close. There are no conditions, prerequisites, or exclusions provided, so the agent must infer the appropriate usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_skill_feedbackAInspect

Propose a skill/playbook rewrite from a human-approved draft vs kept.

Stores the pair as a pending rewrite. The live skill does not change. A human accepts or rejects the proposed diff in the console; accept writes vN+1 through the existing setter. Never auto-overwrites.

ParametersJSON Schema
NameRequiredDescriptionDefault
keptNoHuman-kept / live body snapshot. Default: current live body
kindNoskill (atomic how-to) or playbook (composed flow)skill
nameYesSkill or playbook name that ran
noteNoWhy this rewrite, or the human-approved result summary
agentNoOverride agent identity
draftYesProposed new body_md (skill) or steps_md (playbook)
extraNoProposed extras: tool_hints/tags (skill) or tool_recipe/tags (playbook)
originYesWho proposed the rewrite: agent, human, or curator
descriptionNoProposed one-line summary

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden. It discloses the key traits: the operation is non-destructive, the diff is pending human approval, acceptance writes vN+1 through the existing setter, and it never auto-overwrites. This is strong coverage of the tool's most important behavioral risks.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: the purpose appears in the first sentence, followed by the lifecycle and safety guarantees. Every sentence adds necessary information with no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter tool with no annotations and an output schema, the description covers the essential workflow: pending rewrite, human approval, version bump, and no auto-overwrite. The schema covers parameter details and the output schema covers return values. A minor gap is that it never explicitly routes the agent to the direct setter tools for cases where a pending rewrite is not desired.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 9 parameters. The description adds useful framing around the draft-vs-kept relationship and the human-approval flow, but it does not add per-parameter meaning beyond what the schema provides, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Propose a skill/playbook rewrite'. It further clarifies the exact scope by stating the pair is stored as a pending rewrite and that the live skill does not change, which clearly distinguishes it from direct setters like memory_skill_set.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes the usage context clear: it is for human-approved drafts that need human acceptance/rejection, and it explicitly says the live skill does not change. However, it does not name sibling tools like memory_skill_set or memory_playbook_set as the direct-write alternatives, so the when-not-to-use guidance is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_skill_getAInspect

Fetch a stored skill by name (and optionally version).

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesSkill name
versionNoSpecific version (default: latest active)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. 'Fetch' implies a read-only operation with no side effects, but the description does not explicitly state this or mention error behavior (e.g., what happens if the skill does not exist). It adds only basic intent, nothing beyond the obvious read nature.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that directly states the action and scope. Every word is useful, and there is no extraneous content. It is optimally concise for an effective tool description.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the two parameters are fully documented in the schema and an output schema exists (so return format is not needed), the description is sufficient for an agent to call this tool correctly. It does not explain error handling or version semantics beyond what the schema provides, but for a simple fetch operation, this is adequate. A 4 reflects that it is complete with minor optional additions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, with both 'name' and 'version' having descriptions. The description restates 'by name (and optionally version)' which is redundant with what the schema already provides. It adds no extra meaning beyond the schema, so a baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Fetch' and the resource 'stored skill', and specifies the key parameter (name) and optional version. This distinguishes it from siblings like memory_skills_list (which lists all skills) and memory_skill_set (which writes skills). The purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when a specific skill is needed by name, but does not explicitly mention alternatives or exclusions. The context is clear enough for an agent to infer that this tool is for fetching a single skill rather than listing or modifying, though explicit guidance would be better.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_skill_resolveAInspect

Resolve a playbook's tool_recipe.skills refs to full skill records.

ParametersJSON Schema
NameRequiredDescriptionDefault
playbook_nameYesPlaybook whose skills to resolve
playbook_versionNoPin playbook version (default latest)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It only states the action (resolving refs) without mentioning read-only status, side effects, required permissions, or error behavior. This is a significant gap for a tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, well-formed sentence that conveys the purpose directly. The core action is front-loaded with no wasted words, making it easy to parse and understand.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a simple 2-parameter tool and an output schema present, the description plus schema coverage provide enough information to call the tool correctly. It does not need to explain return values due to the output schema. The main gap is behavioral transparency, which is addressed separately, but for execution context it is adequate.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both playbook_name and playbook_version are already well documented. The description adds no extra parameter semantics beyond the schema, which is acceptable given the high coverage. It merely references the playbook's internal field but not the parameters themselves.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb 'resolve' and a specific resource: a playbook's 'tool_recipe.skills' refs, converting them to 'full skill records'. This clearly distinguishes it from sibling tools like memory_skill_get (which fetches a single skill) and memory_playbook_get (which retrieves the playbook itself).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context that the tool resolves skill references from a playbook, but it does not explicitly state when to use it over alternatives or provide exclusions. An agent could infer the use case, but there is no explicit routing guidance compared to tools that name siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_skill_setAInspect

Insert a new version of a skill. Each call creates a new version.

Skills are atomic instruction building blocks. Playbooks compose them via tool_recipe.skills on memory_procedure_set. Routed through the guarded ingestion pipeline; only active skills are visible to recall and memory_skill_get.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesSkill name (stable id)
tagsNoTags for discovery
agentNoOverride agent identity
body_mdYesMarkdown body the agent will read
owner_idNoOrg member UUID. Omit inherits last owner or the writer.
tool_hintsNoOptional structured hints (preferred MCP tools, params)
descriptionNoOne-line summary
verification_daysNoVerification window: 30, 90, or 180 days; 0 for none. Omit inherits the last window.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full behavioral disclosure burden. It discloses that each call creates a new version, routes through a guarded ingestion pipeline, and that only active skills are visible to recall and memory_skill_get. However, it does not describe response behavior, failure conditions, or how a skill becomes active, leaving some important gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences with the primary action front-loaded in the first sentence. The supporting sentences add valuable context about versioning, composition, and visibility without padding or repeating the input schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter setter with a full output schema, the description covers the key behavioral caveats: every call creates a new version, the write goes through a guarded pipeline, and only active skills are visible. Combined with 100% schema coverage, an agent has enough to invoke the tool correctly, though the practical effects of the 'guarded ingestion pipeline' remain vague.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3 because the input schema already documents every parameter. The description adds conceptual context about skills being atomic building blocks but does not provide parameter-specific meaning beyond what the schema already states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with the exact action, 'Insert a new version of a skill,' and reinforces the versioned behavior with 'Each call creates a new version.' This clearly differentiates it from sibling retrieval tools like memory_skill_get, memory_skill_resolve, and memory_skills_list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly indicates when to use the tool: to insert a new version of a skill. It also adds usage-relevant context by noting that only active skills are visible to recall and memory_skill_get, implying when a created skill becomes usable. It does not explicitly list alternatives or exclusion cases, but the context is clear for a distinctly named setter.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_skills_listAInspect

List all skills (latest version of each) in the caller's org.

ParametersJSON Schema
NameRequiredDescriptionDefault
tagNoFilter by tag
limitNo
offsetNoPagination offset
include_bodyNoInclude full body_md and tool_hints

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the disclosure burden. It does add meaningful behavior—only the latest version of each skill is returned, and results are scoped to the caller's org. However, it omits defaults such as include_body=false and limit=50, does not state that it is read-only, and says nothing about response ordering or pagination behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One sentence, front-loaded with the action, no filler. It earns its place, though it sacrifices some behavioral detail for brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple list tool with an output schema and mostly documented parameters, the description is largely sufficient. It lacks explicit guidance on when to choose this over sibling skill tools, but the org/version scoping plus schema defaults give an agent enough to make a correct call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 75%, so the schema already documents tag, offset, and include_body; limit has min/max/default. The description contributes no parameter-level meaning beyond the schema and does not clarify how tag filtering interacts with 'all skills'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with the verb 'List' and identifies the resource ('all skills'), adds the version qualifier ('latest version of each') and scope ('caller's org'). This is specific enough to distinguish from sibling tools like memory_skill_get or memory_skill_set.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is for enumerating skills in the org, but it never states when to prefer it over memory_skill_get, memory_skill_resolve, or memory_skill_set, nor does it mention any exclusions or alternatives. An agent must infer the use case from the name and the generic list behavior.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_soul_getAInspect

Return this person's private soul for the current org.

The soul is a tiny compressed identity block (who they are, style, likes/dislikes, dos/don'ts). Empty when none yet or the API key is not linked to a human account (mint keys from the console while signed in). Also returned on memory_session_ensure as soul.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of explaining behavior, and it does so well: it defines the soul, states that it is empty when absent or when the API key is not linked to a human account, and notes the alternative return channel via memory_session_ensure. It does not explicitly declare side-effect freedom, but 'Return' plus the empty-state behavior strongly imply a read-only operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and information-dense: a clear one-sentence summary, a useful definition of the output, and two important caveats. Every sentence earns its place without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter getter with an output schema, this description is complete: it names the resource, scope, empty-case behavior, authentication caveat, and a related alternative tool. Nothing needed to invoke or interpret the tool is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so there is nothing to document. The description instead explains the returned resource clearly, which is the relevant semantic information for this tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Return this person's private soul for the current org.' It then explains what a soul is, distinguishing this getter from the many memory_* siblings and clarifying the 'get' semantics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description conveys when the tool is appropriate by defining the target resource and scope ('private soul for the current org'). It also names memory_session_ensure as an alternate source of the same data, though it does not fully spell out when one should be preferred over the other.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_soul_setAInspect

Replace this person's private soul for the current org.

Prefer compact structured markdown. Preferences written via memory_remember(kind=preference) also absorb into the soul.

ParametersJSON Schema
NameRequiredDescriptionDefault
agentNoOverride agent attribution label
body_mdYesCompressed soul markdown (identity, role, style, likes, dislikes, dos/don'ts, patterns). Keep short; server caps length.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden. It does disclose the key behavior ('Replace' implies an overwriting mutation, 'private' implies sensitivity, 'for the current org' scopes it). However, it does not mention reversibility, permissions, side effects beyond the soul, or what happens to existing content not included in the new body.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core purpose. The preference-absorption note is useful. Minor redundancy exists: 'Prefer compact structured markdown' largely repeats the schema's 'Compressed soul markdown... Keep short'.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 100% schema coverage and an existing output schema, the description is largely complete. It explains the action, scope, and an interaction with memory_remember. Gaps are minor: no explicit routing to sibling tools and no disclosure of side-effect details like whether the previous soul is lost irreversibly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds only a stylistic preference for compact structured markdown, which overlaps with the schema's 'Keep short' guidance. It does not meaningfully add parameter meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Replace'), resource ('this person's private soul'), and scope ('for the current org'). This clearly distinguishes it from sibling memory_soul_get and other memory setter tools without requiring schema inspection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The note about memory_remember(kind=preference) absorbing into the soul gives indirect context about an alternative mechanism, but there is no explicit when-to-use vs. alternatives such as memory_soul_get, nor any exclusions. Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_statusAInspect

Seat-scoped memory diagnostic (bind, session, capture, queues).

Read-only. Returns org + bound, active session id / resumed / age, profile present / age (no soul body), last capture / last close timestamps if known, a recent recall latency sample or null, and a redacted reason: ok | unbound_multi_org | auth | capture | retrieval | queue_degraded | unknown. Queue degraded uses the same cheap probe as health. Never returns tokens or keys. Complements service-level health; does not mutate memory.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoNoOptional workspace slug. When set with a bound agent, peeks the conversation/active-session pointer for this repo. Omit to use the caller's most recent open session.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers: 'Read-only,' 'Never returns tokens or keys,' and 'does not mutate memory.' It also discloses redaction behavior via the 'reason' enum and the cheap-probe relationship to 'health.'

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but front-loaded with 'Seat-scoped memory diagnostic' and 'Read-only,' and every clause conveys a distinct fact (outputs, redaction, safety, health relationship). The length is justified by the number of status fields it needs to enumerate.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema exists and the one optional parameter is fully documented, the description covers the operation's purpose, safety, and key output semantics. An agent has enough to decide and invoke correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents 'repo' with optional workspace-slug semantics. The description doesn't add much parameter-specific detail beyond that, but it doesn't need to; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the tool as a 'seat-scoped memory diagnostic' and enumerates its exact outputs, making the verb+resource clear. It also distinguishes itself from the sibling 'health' by stating it 'Complements service-level health' and is scoped to seat/memory rather than service-level.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description establishes when to use it: for inspecting seat-scoped memory state (bind, session, capture, queues) without mutation. It explicitly contrasts with service-level 'health' and notes it 'does not mutate memory,' though it doesn't enumerate all alternatives or hard exclusion conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_thinkAInspect

Synthesized answer with citations and gap analysis (GBrain think parity).

Runs durable recall (default scope excludes working), then composes a cited prose answer plus explicit gaps. For named-entity or competitor questions, call memory_recall with a short keyword anchor first — synthesis quality depends on retrieval. Prefer memory_think when you need prose + gaps after recall surfaced hits, or for open strategic questions. Default sources are current truth; pass include_superseded=true to include retired facts.

ParametersJSON Schema
NameRequiredDescriptionDefault
kNoMax source records to retrieve before synthesis
repoNoWorkspace slug; boosts repo-scoped memories in retrieval
queryYesQuestion to answer from team memory
githubNoGitHub owner/repo; boosts github-tagged memories
token_budgetNoApprox token budget for source packing
include_supersededNoWhen true, synthesize from current truth plus superseded/merged history. Default is current truth only.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does substantial work: it discloses the default scope (durable recall, excludes working memory), the current-truth default versus include_superseded behavior, and the dependency of synthesis quality on retrieval. The only gap is that it never explicitly states the operation has no side effects/writes to memory, which is implied by 'runs durable recall... then composes' but not stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four dense sentences, each earning its place: purpose, pipeline and default scope, sibling routing, and the one flag worth flagging. The purpose is front-loaded in the first sentence, and the GBrain parity reference is a compact one-line contextual anchor rather than padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema exists (return values documented) and parameter coverage is 100%, the description covers everything an agent needs to invoke correctly: what it produces, its internal flow, default scoping, the retrieval prerequisite, sibling routing, and the superseded-data option. Nothing material is missing for a read-synthesis tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all six parameters, establishing the baseline of 3. The description re-emphasizes include_superseded in context ('Default sources are current truth; pass include_superseded=true') but adds no parameter meaning beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource+output: 'Synthesized answer with citations and gap analysis' and describes the pipeline (durable recall → cited prose + explicit gaps). It explicitly differentiates from siblings by naming memory_recall and its complementary role, so an agent can distinguish this from memory_recall, memory_remember, and memory_assemble_context without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit routing guidance: for named-entity or competitor questions, call memory_recall with a short keyword anchor first, and prefer memory_think when prose + gaps are needed after recall hits or for open strategic questions. This is exactly the 'when vs. alternative' guidance the rubric demands, and it even notes the prerequisite retrieval step.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

memory_tools_catalogAInspect

Discover teamshared MCP tools for the current turn.

Returns protocol (every-turn loop), chooser (need → tool), never (hard constraints), and grouped tools with when / avoid / copy-paste example. Pass need= when choosing a tool mid-conversation. Also returns tool_recipe_shapes, aliases (procedure_* → playbook_*), and mcp_list (the live tools/list tier gate and how to advertise extended / alias tools).

ParametersJSON Schema
NameRequiredDescriptionDefault
needNoConversation router: short intent (e.g. 'share a file', 'live slack', 'create a task'). Returns matching chooser rows plus those tools' when/avoid/example. Omit to browse.
tierNoOptional filter: core, extended, or human
scopeNomemory, work, or all tool groupsall

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the disclosure burden. It does so by detailing what is returned, including the need-to-tool chooser mapping, hard constraints, the live tools/list tier gate, and alias normalization. It does not explicitly say 'read-only,' but 'Discover... Returns' makes the non-mutating query nature reasonably clear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short paragraphs with the one-line purpose front-loaded, followed by a compact inventory of return fields and a usage cue. Every sentence contributes information, with no filler or redundant restatement of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema exists, the description need not explain all return values, and it covers purpose, invocation mode, and key output structures well. Minor gaps are unexplained terms such as tool_recipe_shapes and 'advertise extended / alias tools,' but these are left for the output schema to clarify.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema already fully explains need, tier, and scope. The tool description adds only a light usage note about need= and does not meaningfully extend the schema's parameter semantics, keeping this at the baseline 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb and resource: 'Discover teamshared MCP tools for the current turn,' then enumerates concrete returned categories (protocol, chooser, never, grouped tools with when/avoid/example). This clearly differentiates it from the many sibling memory_*/work_* tools, since it is the catalog/discovery tool among them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides an explicit usage trigger: 'Pass need= when choosing a tool mid-conversation,' which tells an agent when to supply the key parameter. It does not name exclusions or alternative tools, but the sibling list contains no direct rival for this catalog function, so the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

org_bindAInspect

Bind this MCP conversation or workspace to an authorized org slug.

Subsequent memory/work/file/context calls on this streamable-HTTP session resolve under that org. Other MCP sessions are unchanged. Membership is fail-closed. Does not change the OAuth token default. Path-bound /o/{slug}/mcp mounts cannot be overridden.

ParametersJSON Schema
NameRequiredDescriptionDefault
slugYesOrg slug this account belongs to (live slug or alias).
scopeNoconversation (this MCP session only) or workspace (account + workspace= slug). Default conversation.conversation
workspaceNoWorkspace / repo slug. Required when scope=workspace. Optional extra key when scope=conversation.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does so thoroughly. It discloses that the binding is session-scoped, leaves other sessions unchanged, is fail-closed, does not change the OAuth token default, and cannot override path-bound mounts—all valuable behavioral context beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is five dense sentences with no wasted words. The core purpose is front-loaded, and every subsequent sentence adds distinct information about scope, failure behavior, or constraints.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given full schema coverage, an output schema, and a description that covers side effects, scope, failure semantics, and limits, the definition is complete for an agent to invoke the tool correctly. Nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents slug, scope, and workspace. The description adds no parameter-level meaning beyond the phrase 'authorized org slug', so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Bind') and resource ('this MCP conversation or workspace to an authorized org slug'), making the tool's purpose unambiguous. It is clearly distinguishable from siblings like org_unbind and org_list based on the binding action and the session-scope detail.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context on when the binding takes effect ('Subsequent memory/work/file/context calls... resolve under that org') and includes an exclusion ('Path-bound /o/{slug}/mcp mounts cannot be overridden'). However, it does not explicitly mention alternatives such as org_unbind or org_list, so it stops short of full alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

org_context_getAInspect

Current org plus bind precedence for this MCP session.

Layers: path (/o/{slug}/mcp) > conversation (Mcp-Session-Id) > workspace (account + repo slug) > OAuth / tsk_ token default. Same snapshot as org_list.

ParametersJSON Schema
NameRequiredDescriptionDefault
workspaceNoOptional workspace / repo slug. When set, include that workspace bind in the snapshot.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It discloses meaningful behavior: the layered precedence order (path > conversation > workspace > OAuth/token default) and that the result is a snapshot equivalent to org_list. It also references authentication context via 'OAuth / tsk_ token default'. It does not explicitly state read-only/no side effects, but the 'get' name and 'snapshot' language strongly imply a safe read operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: the core purpose is stated first, followed by a concise layered breakdown and a useful equivalence to org_list. Every sentence contributes meaningful information without fluff or repetition of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has one optional parameter, an output schema, and a focused purpose, the description covers the essential context: what is returned, how precedence works, and its relationship to org_list. It could be more explicit about when to use it versus org_list, but for a read-only context lookup the provided information is largely sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds value by explaining that 'workspace' corresponds to a binding layer composed of account + repo slug, and by placing it in the precedence hierarchy. This goes beyond the schema's generic 'Optional workspace / repo slug' and helps an agent understand when and why to set the parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool as returning the current org and bind precedence for the MCP session. It names the resource ('current org', 'bind precedence') and references sibling org_list with 'Same snapshot as org_list', which helps situate it. It lacks an explicit verb phrase, relying on the tool name, but the intent is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use the tool: to inspect the current org context and binding precedence for the session. It also references org_list by saying the snapshot is the same, which hints at a relationship but does not explicitly state when to choose this tool over org_list or other org tools. The usage context is implied rather than stated with clear exclusions or alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

org_listAInspect

List memberships, current org, and bind layers for this MCP session.

Does not switch orgs. Use org_bind to bind this conversation or a workspace. tsk_ keys see only their seat org.

ParametersJSON Schema
NameRequiredDescriptionDefault
workspaceNoOptional workspace / repo slug. When set, include that workspace bind in the snapshot.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral disclosure burden. It does well by disclosing side-effect-free behavior ('Does not switch orgs') and a meaningful access limitation ('tsk_ keys see only their seat org'). While it does not exhaustively describe authentication or formatting behavior, it covers the most important semantic traits for safe invocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is tight and front-loaded: the first sentence immediately answers what the tool returns, and the following two sentences add only high-value caveats and routing guidance. Every sentence earns its place, with no filler or redundant elaboration.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity list tool with one fully documented optional parameter and an output schema, the description is complete. It covers the operation's scope, the key side-effect boundary, the alternative tool, and an important access restriction. An agent has enough context to invoke org_list correctly without further inference.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already documents the single optional workspace parameter at 100% coverage, so the description does not need to add parameter-level detail. The description's mention of 'bind layers' aligns with the schema's 'include that workspace bind in the snapshot,' but it does not go beyond what the schema already states. This is the appropriate baseline when schema coverage is complete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function with a specific verb and resource: 'List memberships, current org, and bind layers for this MCP session.' It distinguishes itself from sibling tools like org_bind and org_unbind by immediately noting 'Does not switch orgs,' so an agent knows exactly what org_list does and does not do.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly tells the agent when to use an alternative: 'Use org_bind to bind this conversation or a workspace.' It also sets a clear boundary with 'Does not switch orgs' and provides a useful scoping constraint for certain key types: 'tsk_ keys see only their seat org.' This gives the agent actionable routing guidance beyond the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

org_unbindAInspect

Clear a conversation or workspace org bind for this MCP session.

Later calls fall back to the next precedence layer. Does not change path-bound /o/{slug}/mcp mounts or the OAuth token default.

ParametersJSON Schema
NameRequiredDescriptionDefault
scopeNoconversation, workspace, or all. Default conversation. Token default is never cleared.conversation
workspaceNoWorkspace / repo slug to clear when scope=workspace or all.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Since no annotations are provided, the description carries the full burden. It discloses the key behavior: clearing a bind, the effect on subsequent calls (fallback to next precedence), and what is NOT affected (path-bound mounts and OAuth token default). This is substantial and goes beyond a simple 'clears a bind.' It does not mention side effects like reversal or permissions, but the core behavior is well-transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no fluff. The primary action is front-loaded, and the second sentence adds essential clarifications about precedence and exclusions. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with only two optional parameters, full schema coverage, and an output schema present, the description is complete. It explains the effect on precedence and what is not changed, giving an agent all necessary context to invoke it correctly without requiring additional knowledge.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters (scope and workspace) already have detailed descriptions. The tool description adds no extra parameter-specific syntax or format details, but it does reference 'conversation or workspace' which aligns with scope. Since the schema covers the parameters, baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Clear a conversation or workspace org bind') and the scope ('for this MCP session'). It is specific about the resource and distinguishes from sibling tools like org_bind (which would set) and org_context_get (which reads) by focusing on the clearing action.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use the tool: to remove an existing bind so later calls fall back to the next precedence layer. It also states exclusions (path-bound mounts, OAuth token default) which helps avoid misuse. However, it does not explicitly contrast with org_bind or org_context_get, but the purpose is implied strongly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

process_run_advanceAInspect

Advance a Process run after the current step is finished.

Illegal when the run is completed/cancelled, finished_step is not the current step, or the process is unknown.

ParametersJSON Schema
NameRequiredDescriptionDefault
agentNoOverride agent attribution label
run_idYesProcess run UUID
finished_stepYesStep the accountable seat just finished (must match current_step)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden of behavioral disclosure. It reveals important state restrictions, but does not describe what 'advance' actually changes in the run state, any side effects, or idempotency expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two tight sentences: the first states the action and timing, the second lists illegal conditions. There is no filler or repetition of schema details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The core preconditions and invalid states are stated, and an output schema exists, so return-value detail is not required. It could be slightly more complete by clarifying how this differs from process_run_complete and what state the run is left in after advancing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema fully documents run_id, finished_step, and agent. The description mostly reinforces the schema's finished_step/current_step constraint without adding new parameter-level meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Advance') with a specific resource ('Process run') and states the timing ('after the current step is finished'). It clearly conveys step-wise progress rather than creation or completion, though it does not explicitly name or contrast a sibling tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear trigger condition for use (after a current step finishes) and explicit negative conditions ('Illegal when the run is completed/cancelled, finished_step is not the current step, or the process is unknown'). It does not name alternatives like process_run_complete, but the when/not-when guidance is strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

process_run_completeAInspect

Close a Process run with outcome criteria and result.

Does not require finishing the last step. Emits process.completed on run:{id}. A run that already has an outcome is refused.

ParametersJSON Schema
NameRequiredDescriptionDefault
agentNoOverride agent attribution label
run_idYesProcess run UUID
outcome_resultYesWhat actually happened (met / missed / notes)
outcome_criteriaYesWhat success looks like for this case (frozen on close)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It does so well: it states the emitted event ('process.completed' on 'run:{id}'), clarifies that completing the last step is not required, and discloses the refusal behavior for runs that already have an outcome.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: the core action appears in the first sentence, followed by two terse sentences that add high-value behavioral details. No sentence is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the fully documented schema, the presence of an output schema, and the behavioral details in the description (event emission, not requiring the last step, refusal on already-outcome runs), an agent has enough information to invoke this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema already documents all four parameters. The tool description adds little beyond naming 'outcome criteria and result,' which gives the baseline score of 3 rather than requiring extra compensation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Close a Process run with outcome criteria and result.' It also adds a distinguishing nuance, 'Does not require finishing the last step,' which separates it from process_run_advance and other process-run siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when this tool applies: closing a run with outcome criteria and result, even without finishing the last step. It also notes a firm exclusion ('A run that already has an outcome is refused'), though it does not explicitly name alternative sibling tools for other scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

process_run_createAInspect

Create a Process run from a catalog process or a launchable playbook.

Board Pickup (process or playbook_slug=board-pickup) starts at Scan. Any playbook with tool_recipe.process_steps or a skill sequence can start a run without a hardcoded catalog entry. The run freezes playbook/skill versions + steps, current step, status, and stream_id=run:{id}. Human/gate steps open a Work item and link it. Not a Camunda / Zeebe / LangGraph workflow.

ParametersJSON Schema
NameRequiredDescriptionDefault
agentNoOverride agent attribution label
processNoOptional process slug or name (board-pickup / Board Pickup). Omit when playbook_slug names a launchable playbook.
work_idNoOptional work UUID (writes WorkItem --case_of--> current ProcessStep)
project_idNoOptional project UUID to bind this run (writes Project --runs--> Process)
playbook_slugNoPlaybook to launch as a run. tool_recipe.process_steps drive the walk when present; otherwise the skill sequence is used. Enough on its own — no catalog process required.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden and does so well: it discloses that the run freezes versions/steps/status and stream_id, that human/gate steps open and link a Work item, and that this is not a Camunda/Zeebe/LangGraph workflow. It does not cover permissions or failure modes, but the main side effects are stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, with the core purpose first and supporting behavioral details after. Each sentence adds distinct information; the 'Not a Camunda / Zeebe / LangGraph workflow' note is the only borderline sentence but is short and useful for disambiguation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description, combined with 100% schema parameter coverage and an output schema, gives an agent enough to invoke the tool correctly. It explains the two launch modes and the run-freezing behavior; only explicit guidance on when to prefer standing-draft/commit siblings is absent, which is a minor gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds value by explaining the process/playbook_slug relationship ('board-pickup' starts at Scan) and the conditions under which playbook_slug alone is sufficient, which goes beyond the schema's field-level descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence names a specific verb ('Create'), a specific resource ('Process run'), and the two source types ('catalog process' or 'launchable playbook'). This clearly separates it from sibling process_run_* tools that advance, complete, or read runs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explains when a run can be started without a catalog entry ('Any playbook with tool_recipe.process_steps or a skill sequence') and calls out the Board Pickup special case. It does not explicitly name alternative tools for existing runs, but the 'Create' framing and sibling names make the intended use clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

process_run_eventsAInspect

List the event stream for a Process run (#343 primary event stream).

Returns append-only stream_events for run:{id} in timeline order. The run is the event stream; this is the raw log.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
offsetNo
run_idYesProcess run UUID

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the disclosure burden. It meaningfully reveals append-only semantics, timeline ordering, and that the output is the raw log rather than a derived view. This is genuine behavioral context beyond what the schema or output schema would imply, though it does not mention pagination behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and front-loaded with the core purpose. The second sentence adds useful conceptual context ('the run is the event stream') but the parenthetical '#343 primary event stream' is slightly cryptic and could be clearer.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple list operation with an output schema and only one required parameter, the description plus schema provides enough to call the tool correctly. It captures the essential identity of the tool and the read-only append-only nature, though it would benefit from explicit guidance vs sibling event tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33%, and the description adds no parameter detail. It does not clarify run_id format beyond the schema's 'Process run UUID', and limit/offset are not mentioned at all when they could have been briefly described.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states a specific verb ('List'), a specific resource ('the event stream for a Process run'), and clarifies it is the primary/raw event stream. This meaningfully distinguishes it from generic siblings like event_list or process_run_get, even without naming them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use is implied: call this to read the raw, append-only event log for a Process run. However, it does not explicitly state when to prefer this over alternatives such as event_list or process_run_get, and it gives no exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

process_run_getAInspect

Fetch one Process run (snapshot, current step, standing, work, stream_id).

standing is a projection over run:{id} events (true / changed / next / owner / blocked / since / due / timeout). It is not a writable column — draft with process_run_standing_draft and commit with process_run_standing_commit. Pass source=events to fold the run object from the log alone (no process_runs read).

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYesProcess run UUID
sourceNorow (default) reads process_runs; events reconstructs the run from stream_events on run:{id} only

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses that standing is a projection over events, is not writable, and explains the source parameter's behavior. It also lists the fields returned. This is substantial and helpful, though it could mention potential errors or authorization needs, but for a read tool this is adequate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured. The first line states the main purpose, followed by a crucial explanation of standing and the source parameter. No unnecessary fluff; every sentence adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is an output schema (though not shown), and the description explains the key behavioral aspects: what standing is, how to use source, and what the run contains. It does not cover error handling or permissions, but given the presence of an output schema, the essential information is present. It is complete enough for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds meaning beyond the schema: it clarifies what source=events does (reconstructs from log only) and explains the nature of standing as a projection. This is valuable context that the schema alone does not provide.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb and resource: 'Fetch one Process run' and enumerates the fields returned (snapshot, current step, standing, work, stream_id). It distinguishes itself from sibling tools like process_run_list and process_run_events by focusing on a single run and explicitly calling out the standing projection, which is a unique concept.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly explains the source parameter ('Pass source=events to fold the run object from the log alone') and clarifies when to use this vs. the draft/commit tools for standing (standing is not writable, use process_run_standing_draft/commit). This gives clear guidance on selection and alternatives, even naming specific siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

process_run_listAInspect

List Process runs in this org. Each row includes stream_id and work_id.

Pass source=events to fold every run from the append-only event log — the primary event stream — without touching process_runs.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
offsetNo
sourceNo'row' (default) reads process_runs; 'events' reconstructs the list from stream_events alone (#343 cutover)
statusNoFilter: active, completed, cancelled
processNoFilter by process slug (board-pickup)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the transparency burden. It discloses the data source behavior (`process_runs` vs. the append-only event log), the fact that `source=events` avoids touching `process_runs`, and the included row fields. It does not explicitly state that the operation is read-only, but the word 'List' plus the source distinction gives a solid behavioral picture.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler. It front-loads the core purpose, then adds the most important behavioral distinction (`source=events`) in a compact, readable way. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a list tool with an output schema, the description covers the core purpose, row fields, and the key source-behavior distinction. Filtering and pagination are left to the schema, which documents them adequately. It is missing only an explicit pointer to sibling tools for alternative use cases, but that is a minor gap given the schema and output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 60%, so the schema already documents `source`, `status`, and `process` reasonably well. The description adds useful meaning to `source` by explaining the event-log folding behavior and the 'without touching `process_runs`' aspect, but it does not add anything to `limit`, `offset`, `status`, or `process` beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb and resource: 'List Process runs in this org.' It also names the row fields (`stream_id` and `work_id`) and distinguishes the `events` source path, but it does not explicitly contrast itself with sibling tools like `process_run_events` or `agent_run_list`, so it falls just short of full differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides a concrete usage condition for `source=events` ('fold every run from the append-only event log... without touching `process_runs`'), which helps an agent decide between the default and event-derived paths. However, it gives no guidance on when to choose this tool over sibling tools such as `process_run_get` or `process_run_events`, and it states no exclusions or alternative routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

process_run_standing_commitAInspect

Commit standing on a run stream (append-only event).

Appends standing.committed. Omit every field to promote the latest draft. Partial fields overlay the projection; they do not PATCH a row.

ParametersJSON Schema
NameRequiredDescriptionDefault
dueNoISO due timestamp for the next action
nextNoNext action (visible on the run)
agentNoOverride agent attribution label
ownerNoWho owns the next action
sinceNoISO timestamp when this standing started
run_idYesProcess run UUID
blockedNoWhat is blocking, if anything
changedNoWhat just changed
timeoutNoISO timeout for the current wait
what_is_trueNoWhat is currently true on this run

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden and does well: it discloses that this is an append-only event, that it appends standing.committed, and that partial fields overlay rather than PATCH. It does not cover error cases or idempotency, but the core side-effect semantics are transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences with no filler; the core operation is front-loaded and each sentence adds useful information about event type, promotion behavior, and overlay semantics.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the high parameter count and no annotations, the description is reasonably complete: it covers the operation, the event appended, and the two main invocation modes. It could mention prerequisites such as an existing draft or contrast with process_run_standing_draft, but the output schema exists and the core call contract is clear.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already covers all 10 parameters, so the baseline is 3. The description adds meaningful group-level semantics: omitting every field promotes the latest draft, and partial fields overlay the projection. This clarifies how the parameters collectively behave beyond their individual schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action ('Commit standing') and a specific resource ('run stream'), and identifies the exact event appended ('standing.committed'). It is clear enough to be distinguished from generic event tools, though it does not explicitly differentiate from the sibling process_run_standing_draft.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives concrete usage guidance: omit every field to promote the latest draft, and partial fields overlay the projection rather than patching a row. It provides clear context for how to invoke the tool, though it does not explicitly state when to prefer this over alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

process_run_standing_draftAInspect

Draft standing fields on a run stream (does not become live).

Appends standing.drafted on run:{id}. Read live standing from process_run_get — standing is not a separately mutated field.

ParametersJSON Schema
NameRequiredDescriptionDefault
dueNoISO due timestamp for the next action
nextNoNext action (visible on the run)
agentNoOverride agent attribution label
ownerNoWho owns the next action
sinceNoISO timestamp when this standing started
run_idYesProcess run UUID
blockedNoWhat is blocking, if anything
changedNoWhat just changed
timeoutNoISO timeout for the current wait
what_is_trueNoWhat is currently true on this run

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the behavioral burden. It discloses that the operation is non-destructive (only drafts), that it appends an event named 'standing.drafted', and that the live state is read via process_run_get. This is transparent about the effect. It does not mention permissions or rate limits, but those are not critical for typical use; the core behavior is well-covered.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the essential purpose and the key differentiator (does not become live). Every clause adds value: the event name, the pointer to process_run_get, and the clarification about standing mutation. Zero waste and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has an output schema, return values are covered. The description explains the effect (drafting, appending event) and how to read live standing. It could have explicitly named process_run_standing_commit as the counterpart to make live, but the 'does not become live' phrasing strongly implies its existence. Minor omission, but overall sufficient for an agent to use it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema covers 100% of parameters with clear descriptions for each (due, next, agent, etc.). The tool description adds no additional meaning beyond the schema—it does not explain how parameters relate to 'standing fields' or provide usage hints. Per the baseline rule, when schema coverage is high, a score of 3 is appropriate even with no parameter info in the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Draft standing fields on a run stream' with the critical qualifier '(does not become live)'. It specifies the resource (run stream) and the exact action (appends an event), and explicitly contrasts it with reading live standing via process_run_get, effectively distinguishing it from siblings like process_run_standing_commit. This is specific and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit guidance on when to use this tool: to draft (not make live) standing fields, and tells the agent to use process_run_get to read live standing. It also clarifies that standing is not a separately mutated field, implying this is the appropriate mutation path. This directly informs tool selection and refers to an alternative, leaving no ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

project_createBInspect

Create a project (Asana-style task container).

ParametersJSON Schema
NameRequiredDescriptionDefault
nameYesProject name
agentNoOverride agent identity
colorNoOptional color label
statusNoLifecycle: backlog or active (default active)
team_idNoOwning team UUID
priorityNourgent, high, normal, or low
start_dateNoStart date YYYY-MM-DD
owner_emailNoLead (owner) member email
target_dateNoTarget date YYYY-MM-DD
default_viewNolist, board, timeline, or calendarlist
initiative_idNoStrategic initiative UUID for roll-up
description_mdNoOptional markdown body

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description must carry the burden of behavioral disclosure. It only says 'Create a project' and does not mention side effects, required ownership, idempotency, permission needs, or what the created resource looks like. The name alone already implies creation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one short sentence with no filler. The clarifying parenthetical adds useful context without repeating schema information, and the key action is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 12-parameter creation tool with an output schema and no annotations, the one-line description is thin on behavioral and usage context. However, the schema is self-documenting for parameters and defaults, so the core invocation details are present; only the higher-level lifecycle context is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema documents all 12 parameters and their defaults. The description adds no parameter-level meaning beyond the parenthetical 'Asana-style task container,' which is a baseline contribution.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource ('Create a project') and clarifies what a project is ('Asana-style task container'), which distinguishes this creation tool from sibling read/update tools like project_get, project_list, and project_status_post. The resource and operation are unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given on when to choose this tool over alternatives such as work_create for tasks or project_list for finding existing projects. The phrase 'Asana-style task container' hints at the domain, but the description never states conditions, prerequisites, or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

project_getAInspect

Fetch a project with its sections, latest status, attached files, and optionally its tasks.

ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYesProject UUID
include_itemsNoInclude the project's tasks (board view)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the full burden. The verb 'Fetch' clearly implies a read-only operation, and the description lists what is returned. However, it does not disclose additional behavioral traits such as error behavior for invalid UUIDs, authorization requirements, rate limits, or pagination. For a simple read tool this is adequate, but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, tightly written sentence front-loads the verb and resource, then lists inclusions in a natural order. Every word earns its place; there is no fluff, repetition, or irrelevant detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with two parameters and an output schema, the description adequately covers the primary behavior. It does not need to describe return values because an output schema exists. Minor omissions like error handling or performance characteristics are not critical for correct invocation, and the sibling list provides enough surrounding context for tool selection.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for both parameters, so the baseline is 3. The description adds value by specifying exactly what is always returned (sections, latest status, attached files) and that tasks are optional, reinforcing the role of include_items. This context goes beyond the schema's terse 'Include the project's tasks (board view)'.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description begins with a specific verb 'Fetch' and a clear resource 'a project', then enumerates the full payload: sections, latest status, attached files, and optionally tasks. This distinguishes it from project_list (which lists projects) and work_get (single work item) without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly conveys the tool's scope: it returns a project with its related entities, so an agent knows to select it when needing project detail. It also hints at the include_items toggle by saying 'optionally its tasks', but it does not explicitly mention alternatives like project_list or exclusions (e.g., when to prefer a lighter list).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

project_listCInspect

List projects with health, progress (done/total), lead and member count.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
statusNoOnly this lifecycle: backlog, active, completed, archived
team_idNoFilter by team UUID
initiative_idNoFilter by initiative UUID
include_archivedNoInclude archived projects

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description must carry the full burden of behavioral disclosure. It does not mention default limits (limit=100), default exclusions (include_archived=false), or that results can be filtered by status/team/initiative. These are significant behavioral traits that are only discoverable via the input schema, but the description adds no context beyond a bare listing of output fields.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, clear sentence that front-loads the verb and resource. It contains no fluff or redundancy. Every word contributes to understanding the tool's purpose, making it highly concise and well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

While an output schema exists (so return values are covered), the description fails to mention important behavioral aspects like pagination (limit), default filtering, and the fact that archiving is excluded by default. These are critical for an agent to know when invoking the tool effectively. The description is too minimal to be considered complete for a list tool with multiple optional filters.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 80% (4 of 5 parameters have descriptions). The description does not add any semantic information about parameters; it only lists output fields. Since the schema already covers most parameters well, the description adds little value. Baseline of 3 is appropriate given the high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states what the tool does: list projects, and specifies key output attributes (health, progress, lead, member count). This is distinct from sibling tools like project_get (which fetches a single project) or work_list (which lists work items). However, it doesn't explicitly mention that it supports filtering, so it's not as specific as it could be, but the core purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives like project_get or work_list. There is no mention of 'use this when you need all projects' or any exclusions. An agent must infer from the name and schema, which is insufficient for optimal decision-making.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

project_status_postAInspect

Post a project status update (on-track / at-risk / off-track banner).

ParametersJSON Schema
NameRequiredDescriptionDefault
agentNoOverride agent identity
statusNoAlias for state_label
body_mdNoStatus note (markdown ok)
project_idYesProject UUID
state_labelNoon_track, at_risk, or off_track

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.5/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It only conveys a mutating 'Post' action and banner values; it does not disclose whether existing status is overwritten, what permissions are required, visibility implications, or any side effects beyond the banner.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with the verb and resource first. It is appropriately sized with no redundant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists and all parameters are documented in the input schema, so return values and parameter details don't need to be repeated. However, the description lacks explicit usage context and behavioral side effects, making it adequate but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all five parameters. The description adds no parameter-specific meaning beyond the status banner context, which is already partially captured by the enum descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Post') and resource ('project status update'), and explicitly enumerates the status banner values (on-track / at-risk / off-track). It is unambiguous and clearly distinct from sibling tools like project_create or work_update.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit when-to-use instructions or alternatives are provided, but the tool's purpose is self-evident among the sibling list since no other tool posts status updates. The usage context is implied rather than stated, with no prerequisites or exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

request_gated_approvalAInspect

Create an item_type=approval child and block the parent.

Does not send, publish, spend, delete, or grant. Human or manager marks the child done, then performs the gated action. Agent seats must call this with gate=publish instead of file_publish / storage_publish, or gate=destructive instead of memory_forget.

ParametersJSON Schema
NameRequiredDescriptionDefault
gateYesNamed gate: external_send, spend, destructive, taste, publish. Does not perform the action.
agentNoOverride agent identity
titleYesApproval child title
parent_idYesParent work item UUID to block
description_mdYesProposed action + draft pointer (handoff packet)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses that the tool blocks the parent, does not perform the action, and requires a human or manager to mark the child done before the action occurs. This is comprehensive behavioral disclosure for a gating tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short paragraphs: the first states purpose and non-actions, the second gives usage guidance. Every sentence is purposeful, front-loaded with the core function, and contains no fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex gating tool with 5 parameters and an output schema, the description covers purpose, usage, non-behavior, and routing. It explains the approval workflow sufficiently for an agent to call it correctly without additional context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds value by explaining the 'gate' parameter's practical usage (e.g., gate=publish instead of file_publish), which enriches the schema description. It doesn't detail other parameters but they are already well-covered in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it creates an approval child and blocks the parent, and explicitly lists actions it does NOT perform (send, publish, spend, delete, grant). It distinguishes itself from siblings like file_publish, storage_publish, and memory_forget by specifying when to use this tool instead.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use: 'Agent seats must call this with gate=publish instead of file_publish / storage_publish, or gate=destructive instead of memory_forget.' This provides direct routing guidance against named alternatives, leaving no ambiguity about selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_toolsAInspect

Search for tools using natural language.

Returns matching tool definitions ranked by relevance, in the same format as list_tools.

ParametersJSON Schema
NameRequiredDescriptionDefault
queryYesNatural language query to search for tools

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden. It states that it returns matching tool definitions ranked by relevance, which is a behavioral trait. It also mentions the format is same as list_tools, which is useful. However, it doesn't disclose any side effects, rate limits, or other behavioral details. Given the tool is a search operation, this is adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, two sentences, and front-loaded with the purpose. Every sentence adds value: the first states what it does, the second clarifies the return format. No waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple with one parameter and an output schema. The description explains the return format (same as list_tools) and ranking by relevance. Given the simplicity and the presence of an output schema, the description is complete enough. It could mention that it's a read-only operation, but that's not critical.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents the single parameter 'query' as a natural language query. The description adds no additional meaning beyond that. Baseline 3 is appropriate since the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: searching for tools using natural language. It specifies the action (search) and the resource (tools), and distinguishes it from siblings like list_tools by mentioning the return format. However, it doesn't explicitly differentiate from other sibling tools like call_tool, but the purpose is clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage: when you need to find tools by natural language query. It mentions the return format is same as list_tools, which gives some context. However, it doesn't explicitly state when to use this vs alternatives, nor does it provide exclusions or alternative tool references. The guidance is minimal but not misleading.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

storage_attachAInspect

Attach an existing Storage object to a work item and/or project.

Idempotent on (org, work|project, drive_object). Does not copy bytes. Canonical artifact attach for blobs and documents. Transcripts stay on file_* / work_item_files.kind=transcript. Pass at least one of work_id / project_id. drive_attach is a deprecated alias of this tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
agentNoOverride agent identity
file_idYesStorage object UUID to attach
work_idNoWork item UUID to attach to (kind=artifact)
project_idNoProject UUID to attach to (kind=artifact)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses idempotency (a key retry-safe behavior), that it does not copy bytes (so it is a reference operation), and the canonical usage scope. It does not mention permissions or error conditions, but the core behavioral traits are well covered, earning a 4.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is four sentences, each carrying distinct information: purpose, idempotency/no-copy, canonical usage with transcript exclusion, and the parameter requirement. It is front-loaded with the main verb and object, and there is zero fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a moderate tool with an output schema and 4 parameters, the description covers the essential operational semantics: idempotency, non-copy behavior, canonical usage, transcript routing, and the minimum parameter requirement. It does not explicitly address what happens if both work_id and project_id are provided, but the 'and/or' phrasing implies simultaneous attachment, and the schema already documents each parameter. Overall, the description is complete enough for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds value by clarifying that at least one of work_id/project_id must be passed (a constraint not enforced by the schema, which only requires file_id) and by framing file_id as 'existing Storage object'. This goes beyond the schema descriptions, so a 4 is warranted.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (attach), the object (existing Storage object), and the targets (work item and/or project). It differentiates from siblings by labeling itself the 'canonical artifact attach for blobs and documents' and explicitly excluding transcripts, which distinguishes it from file-related tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context: it is the canonical attach for blobs/documents, and it explicitly says transcripts should go to file_* or work_item_files.kind=transcript. It also states the 'at least one of work_id/project_id' requirement. However, it does not name specific sibling alternatives (like storage_move or storage_publish) for comparison, so it falls short of a perfect 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

storage_folder_createAInspect

Create a Storage folder (metadata only; no bucket bytes).

Folders are drive_objects rows with kind=folder. Root is parent_id IS NULL. Archive refuses a folder that still has active children — empty it first. drive_folder_create is a deprecated alias.

ParametersJSON Schema
NameRequiredDescriptionDefault
agentNoOverride agent identity
titleYesFolder name
parent_idNoParent folder UUID (omit or empty = org root)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden and does well: it discloses that no bucket bytes are created, folders are drive_objects rows with kind=folder, root semantics, and the Archive refusal constraint for non-empty folders. This goes beyond a simple 'create' statement and surfaces the most important operational caveat.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, each earning its place: the core operation and scope, the data-model/root semantics, and the critical prerequisite/alias. Information is front-loaded and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple metadata-only creation tool with 3 params, full schema coverage, and an output schema, the description covers the operation, data model, root behavior, and a key failure condition. Nothing essential is missing for an agent to call this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents title and parent_id. The description adds useful context for parent_id by explaining 'Root is parent_id IS NULL', but it does not add substantial meaning beyond the schema's own descriptions. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Create a Storage folder', and immediately distinguishes it from byte-level operations with '(metadata only; no bucket bytes)'. It also names the deprecated alias drive_folder_create, making the tool's scope unambiguous relative to storage_* siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context: this creates metadata-only folder rows, root is parent_id IS NULL, and Archive requires emptying active children first. It does not explicitly enumerate when-not-to-use scenarios versus storage_attach or file_create, but the metadata-only qualifier and folder-specific model make the intended use clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

storage_getAInspect

Fetch Storage metadata and a download pointer (not the blob body).

Returns download_url for an authenticated GET of the bytes. When published and the bucket publisher is configured, also includes a fresh signed_url (presigned bucket GET; primary for agents) plus signed_url_expires_in / signed_url_expires_at. Durable share_url / public_url still point at /d/{share_token} (Storage's public raw-bytes route; distinct from shared-file /s/{slug}). Includes work_ids / project_ids for current attachments. drive_get is a deprecated alias.

ParametersJSON Schema
NameRequiredDescriptionDefault
file_idYesStorage object UUID

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the transparency burden and does well: it discloses that download_url requires an authenticated GET, that signed_url is conditional on publish/bucket-publisher config, that share_url/public_url use the /d/{share_token} route, and that drive_get is deprecated. It does not explicitly state no side effects or auth scopes, so a strong 4 rather than 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The opening sentence front-loads the core purpose and exclusion, and each subsequent sentence adds a distinct fact about returned URLs or aliases. Slightly dense, but no filler; a high 4.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter read tool with an output schema, the description covers the important nuances: pointer-not-body, conditional signed_url, route distinction from shared-file /s/{slug}, attachment IDs, and the deprecated alias. It does not cover error conditions, but that is a minor gap given the schema and simple input.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%: the single file_id is documented as 'Storage object UUID'. The description adds no extra parameter semantics, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Fetch') and object ('Storage metadata and a download pointer') and explicitly delimits scope with '(not the blob body)', which distinguishes it from blob-fetching siblings like file_get. Also notes the deprecated alias drive_get, further disambiguating.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes clear this is the metadata/pointer accessor, not the body download ('not the blob body'), and signals the deprecated alias 'drive_get'. It lacks an explicit 'when not to use' pointer to file_get or storage_list, so one point off.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

storage_listAInspect

List active Storage objects in the caller's org, newest first.

Each row is metadata plus download_url (bearer GET /v1/drive/{id}/content). Published objects also include a fresh signed_url (presigned bucket GET, primary for agents) plus signed_url_expires_in / signed_url_expires_at, and durable share_url / public_url (/d/{share_token}, console/human fallback). /d/{token} is Storage's public raw-bytes route. Pass work_id or project_id to list attachments (project_id wins if both are set; folder filters are ignored on those joins). Pass parent_id / folder_id to list one folder (root = org root). Omit both to list every active object. Folders sort first. Archived objects stay joined but are omitted. work_id / project_id list every Storage join (blobs and document-capable objects). Transcripts stay on file_list(work_id=). drive_list is a deprecated alias of this tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
queryNoOptional case-insensitive title substring
work_idNoOnly Storage objects attached to this work item UUID
folder_idNoAlias of parent_id
parent_idNoFolder UUID to list, or 'root' / empty for the org root. Omit to list everything.
project_idNoOnly Storage objects attached to this project UUID

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it delivers: sorting (newest first, folders first), archived-object omission, URL semantics for download_url/signed_url/share_url/public_url, route details, and parameter precedence. This goes well beyond the schema and gives a complete behavioral picture.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long and dense, but nearly every sentence carries operational value: filtering modes, URL semantics, exclusions, and sibling routing. It is front-loaded with the core purpose and organized by behavior, though it could be slightly tightened without losing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity and the presence of an output schema, the description is remarkably complete: it covers output row composition, authentication-bearing URLs, filter modes, precedence, exclusions, and deprecated aliases. An agent has everything needed to call this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is already high at 83%, but the description adds significant meaning beyond the schema: it explains that folder filters are ignored on work/project joins, that 'root' means org root, and how the URL fields relate to each other. This materially improves parameter understanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a precise verb+resource: 'List active Storage objects in the caller's org, newest first.' It also distinguishes itself from related tools by noting that drive_list is a deprecated alias and that transcripts belong on file_list, so an agent can tell it apart from siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit filter-mode guidance: pass work_id/project_id for attachments, pass parent_id/folder_id for a single folder, or omit both to list everything. It also states precedence ('project_id wins if both are set') and points to file_list for transcripts, making when-to-use and alternatives clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

storage_moveAInspect

Move a Storage file or folder into another folder (or to root).

Refuses cycles (a folder cannot become its own descendant) and non-folder destinations. drive_move is a deprecated alias.

ParametersJSON Schema
NameRequiredDescriptionDefault
agentNoOverride agent identity
file_idYesStorage object or folder UUID to move
parent_idNoDestination folder UUID, or empty / 'root' for the org root

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral disclosure burden. It discloses concrete error constraints: cycles are refused and non-folder destinations are refused. This goes beyond a simple 'move' statement, though it does not detail auth or side effects on descendants.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the primary action and followed by precise constraints. Every sentence earns its place and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is well-covered: all parameters are documented, the output schema exists, and the description explains key failure modes. Minor omissions like permission requirements or descendant behavior when moving a folder keep it from being fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already fully documents file_id and parent_id. The description adds the conceptual 'or to root' behavior but does not provide substantial parameter semantics beyond what the schema already offers.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Move a Storage file or folder into another folder (or to root).' It clearly distinguishes this from sibling tools like work_move by explicitly naming Storage objects.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly communicates when to use the tool: relocating a storage file or folder. It also routes away from the deprecated alias drive_move. It does not explicitly contrast with work_move or other storage tools, but the 'Storage' scope makes the intended use clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

storage_publishAInspect

Publish a Storage object and return a presigned bucket GET URL.

Idempotent: returns the existing token if already published. Does not re-upload bytes (they already live in the bucket). Agents should use signed_url (TTL TEAMSHARED_STORAGE_SIGNED_URL_TTL_SECONDS, default 3600s) plus signed_url_expires_in / signed_url_expires_at. share_url / public_url remain the durable /d/{share_token} fallback for console/humans — not /s/{slug} (shared-file HTML). Fails closed if the object-storage bucket is unconfigured (cannot mint a signed URL).

Agent seats (org tsk_ / agent_run) are refused — call request_gated_approval(gate="publish") instead. Console humans (ts_session) still publish. After the approval child is done, a human seat calls this tool (same completion path as file_publish). drive_publish is a deprecated alias of this tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
file_idYesStorage object UUID to publish

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does so richly. It discloses idempotency, that bytes are not re-uploaded, the signed-URL TTL and expiration fields, the failure mode when the bucket is unconfigured, and the agent-seat refusal behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence adds operational value: purpose, idempotency, URL semantics, TTL details, auth gating, and deprecated alias. The core purpose is front-loaded and the supporting details are logically grouped.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a publish token tool with auth constraints, failure modes, and URL semantics, the description is exceptionally complete. It covers the output fields to use, the approval flow, the fallback behavior, and the deprecated alias, leaving no critical operational gap for an agent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single parameter file_id is already described as 'Storage object UUID to publish' in the schema. The description adds useful context about existing bucket bytes and idempotent behavior, but does not add significant new param-specific meaning beyond what the schema states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Publish a Storage object and return a presigned bucket GET URL.' It distinguishes the tool from siblings like file_publish and drive_publish by clarifying that drive_publish is a deprecated alias and noting the same completion path as file_publish.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use and when-not-to-use guidance: agent seats must call request_gated_approval(gate='publish') instead, while console humans can publish. It also directs agents to prefer signed_url fields and identifies the durable /d/{share_token} fallback versus the /s/{slug} HTML path.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

storage_upload_requestAInspect

Mint a one-time uploader script for a private Storage blob (any file type).

Storage is the org's binary store: bytes go to the Railway bucket, metadata to drive_objects. Objects stay private until a human seat calls storage_publish (agent seats must request_gated_approval(gate=publish)). Document MIME types (Markdown, HTML, plain text, diagram JSON/YAML) are stored as blobs; versioned edit / /s/ publish still use file_upload_request. Pass work_id and/or project_id to attach after ingest (kind=artifact; transcripts stay on shared files). Pass parent_id to place the file in a folder. Batch uploads are multiple grants (one file each) — console multi-select uses sequential ingest, not a multi-file grant. Returns upload_url, upload_token, max_bytes, and a self-deleting Python script. Token is single-use and expires in ~10 min. drive_upload_request is a deprecated alias of this tool.

ParametersJSON Schema
NameRequiredDescriptionDefault
agentNoOverride agent identity
titleYesStorage object title
work_idNoAttach the uploaded Storage object to this work item UUID
filenameNoOptional local filename (script default path + MIME sniff)
parent_idNoFolder UUID to land the file in (omit or empty = org root). One grant per file.
project_idNoAttach the uploaded Storage object to this project UUID (kind=artifact)
content_typeNoOptional MIME type hint stored on the grant
upload_base_urlNoOptional server origin (e.g. https://teamshared.com). Defaults to settings.public_url.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden — and it delivers: it discloses that the token is single-use with ~10 min expiry, the Python script is self-deleting, objects remain private until published, and batch uploads are strictly one-grant-per-file with console multi-select using sequential ingest. This is rich security-and-lifecycle context well beyond a bare 'upload' claim.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long, but every sentence earns its place — storage model, privacy flow, sibling routing, parameter semantics, return contract, token lifecycle, and deprecation notice. It is front-loaded with a crisp one-line summary and organized so each sentence adds non-redundant information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter, high-stakes upload tool with no annotations, the description is remarkably complete: it covers the storage architecture, the publish/gating workflow, sibling alternatives, parameter semantics, return values, token expiry, and alias deprecation. The output schema exists and the description still adds return-field context, leaving no material gap for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds genuine meaning beyond the schema: work_id/project_id attach with kind=artifact plus the caveat that transcripts stay on shared files, parent_id places the file in a folder, and the one-grant-per-file granularity constraint. It does not merely restate the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific verb ("Mint"), a precise resource ("one-time uploader script for a private Storage blob"), and explicitly broadens scope ("any file type"). It further distinguishes itself from siblings by naming file_upload_request for versioned edit //s/ publish flows and noting drive_upload_request as a deprecated alias.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly delineates when to use this tool (binary/any-type blobs into Storage) versus when not to (document MIME types needing versioned edit //s/ publish → file_upload_request). It also clarifies the post-ingest publish flow (human calls storage_publish; agent seats need request_gated_approval) and warns off the deprecated alias.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

versionAInspect

Report server + memory-rule version and whether the rule needs updating.

Returns {server_version, rule_version, installed_rule_version, rule_path, rule_update, update_available}. rule_path is the canonical rule in the teamshared Cursor plugin; Cursor loads it from there. When update_available is true (the installed rule is missing or behind), tell the user to update the teamshared plugin. Never write a local .cursor/rules/teamshared.mdc copy. rule_markdown is also returned for clients without the plugin, which may put it in their instructions slot (e.g. AGENTS.md). See the rule's "Staying current".

ParametersJSON Schema
NameRequiredDescriptionDefault
installed_rule_versionNoThe `version` from your installed teamshared rule's frontmatter (Cursor: the teamshared plugin's rules/teamshared.mdc). Omit if your rule has no version marker.

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does meaningfully more than a bare read: it explains the canonical rule location, the required user-facing action when update_available is true, an explicit prohibition on writing a local copy, and the fallback role of rule_markdown for clients without the plugin. It stops short of explicitly labeling the call as a safe read-only operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded, then return keys, then action guidance, then the critical 'never write a local copy' warning. The only mild redundancy is enumerating return keys, which the output schema already covers.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the description need not restate returns, yet it does and also supplies post-call decision logic and a safety prohibition. For a single optional-parameter diagnostic tool this is nearly complete; the only gap is the absence of an explicit invocation trigger.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the sole parameter (installed_rule_version) is fully documented in the schema, including its origin and the omit-if-absent case. The description adds no further syntax or format detail beyond that, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Report') and two resources (server version and memory-rule version) plus the derived question of whether an update is needed. An agent immediately understands this is a diagnostic/version-read tool distinct from any sibling, since no other tool covers version reporting.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives procedural guidance for interpreting results (if update_available, tell the user to update the teamshared plugin) and warns never to write a local rule copy, but it never states when an agent should call this tool versus not calling it. Usage is implied rather than framed as when/when-not.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

work_add_to_projectBInspect

Add a task to a project (tasks can belong to multiple projects).

Result includes org + bound.

ParametersJSON Schema
NameRequiredDescriptionDefault
agentNoOverride agent identity
work_idYesWork item UUID
project_idYesProject UUID
section_idNoOptional section UUID

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses that the result includes 'org' and 'bound' (some output context), and notes multi-project membership, but it does not reveal permissions, idempotency, side effects, or error behavior. For a mutation tool, this is a significant transparency gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise—two sentences with no fluff. The main action is front-loaded. The second sentence about the result is cryptic but not verbose. It earns a high score for efficiency, though the cryptic result note could be clearer.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained. However, the tool has four parameters (two optional) and no annotations; the description does not explain when to use section_id or agent override, nor does it address idempotency or failure modes. For a mutation tool in a rich context, this is incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

All four parameters have schema descriptions (coverage 100%), so the schema already documents them. The tool description adds no extra meaning about parameters like agent override or section_id. The baseline 3 is appropriate because the schema covers the semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a clear verb and resource: 'Add a task to a project.' It also adds the key semantic that tasks can belong to multiple projects, which differentiates it from the sibling work_remove_from_project. The purpose is unambiguous and specific.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage (add a task to a project) and the multi-project note hints at when it's applicable, but it does not explicitly contrast with the sibling work_remove_from_project or state when not to use it. The guidance is implied rather than explicit, so it falls short of a clear when/when-not statement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

work_closeAInspect

Mark a work item done or cancelled. Result includes org + bound.

If the item's playbook sets artifact_required, closing done is refused (handoff_required) until the item has an attached artifact or a comment with a pointer (file:<id>, drive:<id>, a PR URL).

ParametersJSON Schema
NameRequiredDescriptionDefault
agentNoOverride agent identity
work_idYesWork item UUID
work_statusNodone or cancelleddone

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and does disclose non-obvious behavior: the artifact/handoff gate and the refusal code, plus a note that the result carries org and bound. Missing is permission/auth requirements and whether the transition is reversible.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action in one sentence, then the gating rule in a short second paragraph. The mention of ``org`` + ``bound`` in the result is slightly redundant given an output schema exists, but nothing else is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers the key state-transition precondition and relies appropriately on the output schema for return values. What is left implicit is the relationship to work_update and the identity semantics of the optional agent override.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so work_id, work_status and agent are already documented in the schema. The description confirms only the two terminal statuses and adds no format or interaction detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (mark a work item done/cancelled) and names the two terminal statuses. However, it does not distinguish itself from the sibling work_update, which plausibly also changes status, so the agent must infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives one concrete precondition (closing 'done' is refused with handoff_required when the playbook sets artifact_required), which is useful context. But there is no explicit when-to-use-this vs work_update guidance or statement about when cancellation is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

work_comment_addCInspect

Add a comment to a work item. Result includes org + bound.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyYesComment text (markdown ok)
agentNoOverride agent identity
work_idYesWork item UUID

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It only hints at the result format ('org' + 'bound') but does not explain what these mean, nor does it mention permissions, idempotency, or side effects. The verb 'Add' implies a mutation, but the description does not explicitly confirm write behavior or its consequences.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact—two sentences—with no filler. The first sentence clearly states the action. The second sentence about result is concise but cryptic, slightly reducing clarity without adding bulk. Overall, it is efficient, though the result note could be phrased more meaningfully.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a simple operation (3 params, 2 required) with 100% schema coverage, so the description need not elaborate on parameters. An output schema exists (not shown but flagged), which likely explains return fields, reducing the need for description detail. However, the cryptic 'org' and 'bound' reference is a gap—if the output schema does not clarify it, agents may be confused. The description is adequate for a low-complexity tool but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema already documents all three parameters (work_id, body, agent) with descriptions. The tool description adds no parameter-specific information; it only mentions 'org' and 'bound' in the result, which is not about parameters. Per the baseline, a 3 is appropriate given full schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear action: 'Add a comment to a work item.' It specifies the verb (add) and resource (work item), which is unambiguous. However, it does not differentiate from sibling tools like work_comment_list, and the cryptic 'Result includes ``org`` + ``bound``' adds confusion rather than clarity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus alternatives. The natural sibling is work_comment_list, but the description never mentions it, nor any conditions, prerequisites, or exclusions. An agent is left to infer that this tool is for adding, not listing, comments.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

work_comment_listAInspect

List comments on a work item (oldest first).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
work_idYesWork item UUID

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the disclosure burden. It states the read-only nature ('List') and the ordering ('oldest first'), which is useful. It does not disclose pagination/limit behavior or whether the listing is flat or hierarchical, though the output schema and limit parameter partially cover operational details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single compact sentence that front-loads the core action, names the resource, and adds the key ordering detail. Every word earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list tool with one required parameter and an output schema, the description plus schema is nearly sufficient. It identifies the target resource, the ordering, and the only required input. A small gap is the lack of any mention of pagination behavior beyond the schema's limit field.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 50%: work_id is described as 'Work item UUID', while limit only has type/range/default constraints. The description adds little beyond the schema, tying the action to 'a work item', but the limit parameter's purpose is reasonably inferable from its name and constraints. No severe gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'List comments on a work item'. It also adds the meaningful ordering detail '(oldest first)', and it is clearly distinguishable from siblings like work_comment_add (creating comments) and work_list (listing work items).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use is obvious: call this when you need the comments attached to a work item. However, the description gives no explicit when/when-not guidance and does not point to alternatives such as work_comment_add for adding comments or work_get for retrieving the work item itself.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

work_createAInspect

Create a work item. Created active immediately for humans and agents (no approval queue).

Creating-agent path: omitted priority → skill task-priority-assessment; omitted due_at → skill task-time-required. The distiller does not fill these. Console blanks default to priority=normal only. Result includes org + bound.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoNoOptional workspace slug tag
tagsNoOptional tags
agentNoOverride agent identity
titleYesShort task title
due_atNoOptional due datetime. If omitted, fetch skill task-time-required and pass the computed due_at
githubNoOptional owner/repo tag
part_ofNoOntology entity this beat belongs to (campaign Project name)
priorityNourgent, high, normal, low. If the user/packet did not name one, fetch skill task-priority-assessment first — do not silently take the normal defaultnormal
start_atNoOptional start datetime
item_typeNotask, milestone, or approvaltask
parent_idNoParent task UUID (makes this a subtask)
project_idNoAdd the task to this project UUID
section_idNoPlace in this project section UUID
assignee_idNoAssignee UUID
descriptionNoAlias for description_md
work_statusNoInitial workflow statustodo
assignee_typeNoAssignee type (user)
initiative_idNoOptional strategic initiative UUID
playbook_slugNoNamed playbook to inject when agent_run_start(work_id=) spawns (e.g. teamshared-manager). Unset = no inject
assignee_emailNoAssign to org member by email
description_mdNoOptional markdown body
assigned_to_entityNoOntology Person or Agent name (graph assigned_to; not assignee_email)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It is genuinely informative: it states that created items are active immediately, there is no approval queue, the distiller does not fill omitted priority/due_at, and the result includes 'org' + 'bound'. This goes well beyond a typical 'create X' description, though it does not address permissions or reversibility.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and well-structured: a one-sentence core purpose, then a focused bullet list for the non-obvious agent-path behavior. Every sentence earns its place. The only minor weakness is the cryptic 'Result includes ``org`` + ``bound``' phrase, which is concise but under-explained.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 22-parameter creation tool with no annotations, the description covers the most decision-critical behaviors: immediate activation, no approval queue, and the skill-invocation rules for priority and due_at. The output schema presumably documents return values, so the brief mention of result fields is acceptable. It is not exhaustive, but it gives an agent the essential context needed to call this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds some value by highlighting the special behavior of omitted priority and due_at parameters and clarifying that console blanks only default priority to normal. However, it does not add meaningful semantics for the other 20 parameters, which remain fully documented only in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Create a work item.' It immediately distinguishes this tool from siblings like work_update and work_close by stating the created item becomes active immediately for both humans and agents with no approval queue. This is unambiguous and tells an agent exactly what operation this tool performs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context about when this tool applies — creating work — and adds important behavioral context about the creating-agent path versus console entry: omitted priority/due_at trigger assessment skills, and console blanks default to priority=normal only. It does not explicitly name alternative tools or state when not to use it, but the creation purpose and contrast with the update/close siblings are clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

work_getAInspect

Fetch one work item by id.

ParametersJSON Schema
NameRequiredDescriptionDefault
work_idYesWork item UUID

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral disclosure burden. It conveys a read-only operation via 'Fetch', but does not mention error handling, idempotency, permissions, or other behavioral traits. Adequate for a simple get tool but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence with no filler, front-loaded with the action and resource. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter read tool with an output schema, the description covers the core behavior, the parameter schema fully documents the input, and no essential information is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and already documents work_id as a UUID. The description adds no extra meaning beyond restating 'by id', so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states a specific verb ('Fetch'), resource ('work item'), and identifier ('by id'), clearly distinguishing it from siblings like work_list and work_create. The purpose is immediately unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage when a single work item ID is known, but does not explicitly mention alternatives or when not to use it (e.g., vs. work_list). Context is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

work_listBInspect

List org work items (shared task queue for humans and agents).

ParametersJSON Schema
NameRequiredDescriptionDefault
tagNoOnly items carrying this label
mineNoOnly items assigned to the caller (human or agent)
sortNoSort keyupdated_at
limitNo
offsetNoPagination offset
has_dueNoTrue: only dated items; False: only undated
assigneeNoFilter by user email, agent name, actor id, or 'unassigned'. Names that match nobody return no items
due_fromNoInclusive start of due_at window (UTC)
sort_dirNoasc or descdesc
due_untilNoExclusive end of due_at window (UTC)
item_typeNoFilter: task, milestone, or approval
project_idNoFilter to work linked to this project UUID
work_statusNoFilter by workflow status
closed_sinceNoAlso include items closed at or after this instant (UTC)
initiative_idNoFilter to tasks linked to a strategic initiative UUID
work_statusesNoMatch any of these statuses (e.g. todo, in_progress, blocked)
exclude_closedNoOmit done/cancelled items (default true)
include_cancelledNoFalse drops cancelled items (default true)

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must disclose behavioral traits. It only says 'List', implying a read operation, but does not mention pagination behavior, default filters (e.g., exclude_closed defaults true), or any potential side effects. The shared queue context is mentioned but not elaborated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with no filler, front-loading the core purpose. It is highly concise and immediately clear.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema exists and parameter descriptions are detailed, so much of the contextual burden is handled. However, the description does not mention that this is a shared queue for humans and agents beyond a parenthetical, nor does it explain any unique behaviors like default filtering. For a tool with 18 parameters, a bit more context would help but is not critical given the schema richness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 94%, so most parameters are already well documented in the input schema. The description adds no additional meaning or context for parameters. Baseline 3 is appropriate because the schema carries the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'List' and the resource 'org work items', and adds context that this is a 'shared task queue for humans and agents'. This distinguishes it from single-item tools like work_get and create tools like work_create without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It does not mention work_get for single items, work_create for creating, or other list tools. There is no explicit context for when this is the appropriate choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

work_moveBInspect

Move a task to a section and/or reorder it within a project.

Result includes org + bound.

ParametersJSON Schema
NameRequiredDescriptionDefault
agentNoOverride agent identity
work_idYesWork item UUID
project_idYesProject UUID
section_idNoTarget section UUID
sort_orderNoFractional rank within section

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the core operation but does not mention side effects, permissions, reversibility, or what happens when optional parameters like section_id or sort_order are omitted. The cryptic 'Result includes org + bound' adds little clarity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and front-loaded, with the primary action stated in the first sentence. The second sentence about 'org + bound' is cryptic but not verbose; overall, the description is appropriately sized with no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a complete input schema and an output schema present, the description does not need to explain return values. However, the lack of annotations and absence of usage guidance or behavioral caveats leave some gaps for a mutation tool, especially regarding when to choose this over sibling tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema already documents all parameters. The description adds minimal semantic value beyond mentioning 'section' and 'reorder,' which map to section_id and sort_order, but it does not compensate or extend the schema's parameter explanations.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action ('Move a task to a section and/or reorder it within a project') with a specific verb and resource. It distinguishes the tool from siblings like work_add_to_project and work_remove_from_project by its focus on moving/reordering, though it does not explicitly name alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use the tool: when a task needs to be moved to a section or reordered within a project. However, it provides no explicit guidance about when not to use it or which sibling tools might be better suited for related operations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

work_remove_from_projectBInspect

Remove a task from a project. Result includes org + bound.

removed is false when the task was not in that project.

ParametersJSON Schema
NameRequiredDescriptionDefault
work_idYesWork item UUID
project_idYesProject UUID

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations at all, the description carries the full behavioral burden. It does disclose a genuinely useful trait — the non-error outcome where 'removed' is false if the task was not in the project — but says nothing about authorization needs, side effects, or whether the task itself is altered beyond the project association.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with the core action front-loaded and no wasted wording. The residual RST-style double-backtick markup is slightly noisy but does not obscure meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the description is not obliged to explain return values (though it usefully clarifies the 'removed' flag semantics). For a mutation tool with zero annotation coverage, however, it omits permissions, idempotency, and the relationship to work_add_to_project, leaving real gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and both UUID parameters are documented in the schema, so the description owes nothing extra. It adds no syntax, format, or validation detail beyond what the schema already supplies, which is the expected baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Remove a task from a project') so the operation is unambiguous. It does not, however, name or differentiate itself from the obvious inverse sibling work_add_to_project, leaving the routing implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use guidance, no prerequisites (e.g. required permissions or whether the project must exist), and no mention of the alternative work_add_to_project for the reverse operation. Usage is only inferable from the tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

work_updateCInspect

Update a work item (status, assignee, priority, parent, due_at, etc.).

Result includes org + bound.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoNoWorkspace slug tag
tagsNoReplace tags
agentNoOverride agent identity
titleNoNew title
due_atNoDue datetime. Omit to leave unchanged; use clear_due to unset
githubNoowner/repo tag
work_idYesWork item UUID
priorityNourgent, high, normal, low
clear_dueNoClear due_at. Omit due_at to leave the date unchanged
parent_idNoParent task UUID (reparent as subtask)
assignee_idNoAssignee UUID
work_statusNoWorkflow status
assignee_typeNoAssignee type (user)
initiative_idNoStrategic initiative UUID
playbook_slugNoNamed playbook injected on agent_run_start(work_id=). Empty string clears. Omit to leave unchanged
assignee_emailNoAssign to user by email
blocked_reasonNoWhy blocked (when status=blocked)
description_mdNoNew markdown body

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full burden. It only says 'Update a work item' and that the result includes org + bound; it does not disclose permissions, side effects, reversibility, or behavior when only work_id is supplied. This is thin disclosure for a mutating tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and front-loaded, with the core purpose in the first sentence. The second sentence about org + bound is somewhat cryptic and may be redundant given the output schema, so it doesn't fully earn a top score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex 18-parameter mutation tool with no annotations, this description is under-specified: it lacks usage conditions, exclusions, and side-effect disclosure. The schema covers parameters and output, but the description doesn't complete the behavioral picture an agent needs to choose and invoke the tool safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description's field list is a useful summary but adds little beyond the schema; it does not clarify interactions between assignee_id/assignee_email/assignee_type or due_at/clear_due, which are already documented in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the action (update), the resource (work item), and examples of mutable fields (status, assignee, priority, parent, due_at). It is clear at a glance, but it doesn't explicitly distinguish this from overlapping siblings such as work_close or work_move, so it stops short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to use this tool versus alternatives like work_close, work_move, or work_create, nor any mention of prerequisites or special cases such as clearing fields. The description only states the action and result; an agent must infer the appropriate context from the tool name and sibling names.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool update
    • Changedagent_run_start1 field changed
      • addedInput schema / properties / agent_slug
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Agent record to run as (agent-<label> or <label>; see agent_list). The run inherits the record's model and tools: an explicit model= wins, an explicit tools= can only narrow the agent's. Unknown, paused or archived agents are refused."
        +}
  2. 1 tool update
    • Addedobject_search
  3. 1 tool update
    • Changedversion1 field changed
      • changedInput schema / properties / installed_rule_version / description
        Previous value: -"The `version` from your installed teamshared rule's frontmatter (e.g. the value in ~/.cursor/rules/teamshared.mdc). Omit if your rule has no version marker."New value: +"The `version` from your installed teamshared rule's frontmatter (Cursor: the teamshared plugin's rules/teamshared.mdc). Omit if your rule has no version marker."
  4. 3 tool updates
    • Changedproject_create5 fields changed
      • changedInput schema / properties / owner_email / description
        Previous value: -"Owner member email"New value: +"Lead (owner) member email"
      • addedInput schema / properties / priority
        Added value: +{
        +  "anyOf": [
        +    {
        +      "enum": [
        +        "urgent",
        +        "high",
        +        "normal",
        +        "low"
        +      ],
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "urgent, high, normal, or low"
        +}
      • addedInput schema / properties / start_date
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Start date YYYY-MM-DD"
        +}
      • addedInput schema / properties / status
        Added value: +{
        +  "anyOf": [
        +    {
        +      "enum": [
        +        "backlog",
        +        "active",
        +        "completed",
        +        "archived"
        +      ],
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Lifecycle: backlog or active (default active)"
        +}
      • addedInput schema / properties / target_date
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Target date YYYY-MM-DD"
        +}
    • Changedproject_list1 field changed
      • addedInput schema / properties / status
        Added value: +{
        +  "anyOf": [
        +    {
        +      "enum": [
        +        "backlog",
        +        "active",
        +        "completed",
        +        "archived"
        +      ],
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Only this lifecycle: backlog, active, completed, archived"
        +}
    • Changedwork_list6 fields changed
      • changedInput schema / properties / assignee / description
        Previous value: -"Filter by agent name or user email"New value: +"Filter by user email, agent name, actor id, or 'unassigned'. Names that match nobody return no items"
      • addedInput schema / properties / closed_since
        Added value: +{
        +  "anyOf": [
        +    {
        +      "format": "date-time",
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Also include items closed at or after this instant (UTC)"
        +}
      • addedInput schema / properties / include_cancelled
        Added value: +{
        +  "default": true,
        +  "description": "False drops cancelled items (default true)",
        +  "type": "boolean"
        +}
      • changedInput schema / properties / sort / enum
        Previous value: -[
        -  "updated_at",
        -  "priority",
        -  "work_status",
        -  "created_at"
        -]New value: +[
        +  "updated_at",
        +  "priority",
        +  "work_status",
        +  "created_at",
        +  "due_at"
        +]
      • addedInput schema / properties / tag
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Only items carrying this label"
        +}
      • addedInput schema / properties / work_statuses
        Added value: +{
        +  "anyOf": [
        +    {
        +      "items": {
        +        "enum": [
        +          "backlog",
        +          "todo",
        +          "in_progress",
        +          "blocked",
        +          "done",
        +          "cancelled"
        +        ],
        +        "type": "string"
        +      },
        +      "type": "array"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Match any of these statuses (e.g. todo, in_progress, blocked)"
        +}
  5. 3 tool updates
    • Addedcall_tool
    • Addedmemory_context_pack
    • Addedsearch_tools
  6. 60 tool updates
    • Removedagent_run_child
    • Removedagent_run_trace_get
    • Removedcontext_retrieve
    • Removeddrive_attach
    • Removeddrive_folder_create
    • Removeddrive_get
    • Removeddrive_list
    • Removeddrive_move
    • Removeddrive_publish
    • Removeddrive_unlink
    • Removeddrive_unpublish
    • Removeddrive_upload_request
    • Removedfile_archive
    • Removedfile_unpublish
    • Removedfile_version_delete
    • Removedmemory_action_apply
    • Removedmemory_action_log_list
    • Removedmemory_capture_inventory
    • Removedmemory_distill_scorecard
    • Removedmemory_episodes_list
    • Removedmemory_forget
    • Removedmemory_forget_procedure
    • Removedmemory_forget_skill
    • Removedmemory_graph_relate
    • Removedmemory_graph_related
    • Removedmemory_ontology_link_type_set
    • Removedmemory_ontology_list
    • Removedmemory_ontology_merge_entities
    • Removedmemory_ontology_object_kind_set
    • Removedmemory_ontology_propose_entity
    • Removedmemory_ontology_rekind_entity
    • Removedmemory_procedure_get
    • Removedmemory_procedure_set
    • Removedmemory_procedures_list
    • Changedmemory_recall1 field changed
      • changedInput schema / properties / explain / description
        Previous value: -"When true, include per-record retrieval attribution in metadata.ranking (base RRF, every factor, relevance floor verdict, diversity penalty; its score equals the hit's score), plus read-only feedback_score when votes exist"New value: +"When true, include per-record retrieval attribution in metadata.ranking (base RRF, every factor including bounded feedback trust from memory_feedback votes, relevance floor verdict, diversity penalty; its score equals the hit's score), plus read-only feedback_score when votes exist"
    • Removedmemory_session_get
    • Removedmemory_state_get
    • Removedmemory_state_set
    • Removedmemory_strategic_entity_get
    • Removedmemory_strategic_initiative_set
    • Removedmemory_strategic_key_result_set
    • Removedmemory_strategic_objective_set
    • Removedmemory_strategic_plan_get
    • Removedmemory_strategic_plan_list
    • Removedmemory_strategic_plan_set
    • Removedmemory_strategic_statement_get
    • Removedmemory_strategic_statement_set
    • Removedproject_archive
    • Removedproject_section_add
    • Removedproject_section_list
    • Removedproject_update
    • Removedstorage_unlink
    • Removedstorage_unpublish
    • Removedwork_dependencies_list
    • Removedwork_dependency_add
    • Removedwork_dependency_remove
    • Removedwork_follower_add
    • Removedwork_follower_remove
    • Removedwork_followers_list
    • Removedwork_subtasks_list
  7. 2 tool updates
    • Changedagent_run_child1 field changed
      • addedInput schema / properties / durable_memory
        Added value: +{
        +  "default": false,
        +  "description": "Keep the child's own capture session and run memory, for a specialist that must remember its work. Default false.",
        +  "type": "boolean"
        +}
    • Changedagent_run_start1 field changed
      • addedInput schema / properties / durable_memory
        Added value: +{
        +  "default": false,
        +  "description": "Nested runs are ephemeral: no capture session or run memory of their own, the parent's work thread is the record. Set true for a specialist that must keep its own memory.",
        +  "type": "boolean"
        +}
  8. 1 tool update
    • Changedagent_run_start1 field changed
      • addedInput schema / properties / parent_run_id
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Parent agent run UUID when this is a nested child run; defaults to the run that invoked this tool"
        +}
  9. 1 tool update
    • Addedagent_run_trace_get
  10. 1 tool update
    • Addedagent_run_child
  11. 1 tool update
    • Addedagent_run_hitl_decide
  12. 1 tool update
    • Changedmemory_recall1 field changed
      • changedInput schema / properties / explain / description
        Previous value: -"When true, include per-record retrieval attribution in metadata, plus read-only feedback_score when votes exist"New value: +"When true, include per-record retrieval attribution in metadata.ranking (base RRF, every factor, relevance floor verdict, diversity penalty; its score equals the hit's score), plus read-only feedback_score when votes exist"
  13. 1 tool update
    • Changedmemory_session_ensure1 field changed
      • addedInput schema / properties / auto_recall
        Added value: +{
        +  "default": false,
        +  "description": "Opt-in thin-client bootstrap: when true and user= and/or topic= is set, run a budgeted memory_recall (k=5, explain=false, short deadline) over that text and attach hits as auto_recall. Missing both skips without error. Default false — existing callers are unchanged. Prefer an explicit memory_recall with a short keyword anchor; this is not a replacement.",
        +  "type": "boolean"
        +}
  14. 1 tool update
    • Changedagent_run_start1 field changed
      • changedInput schema / properties / model / description
        Previous value: -"Router model id; omit to use TEAMSHARED_AGENT_RUN_MODEL"New value: +"Router model id (prefixed, e.g. openrouter/...); omit to use TEAMSHARED_AGENT_RUN_MODEL. An id the router would refuse (no routable prefix) is replaced by that default and the result carries model_substitution={requested, model, reason}."
  15. 1 tool update
    • Changedagent_run_start1 field changed
      • addedInput schema / properties / tools
        Added value: +{
        +  "anyOf": [
        +    {
        +      "items": {
        +        "type": "string"
        +      },
        +      "type": "array"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "description": "Tool name/prefix patterns this run may bind (e.g. ['work_*','memory_recall']). Narrows the box allowlist; never widens it. Omit for the deployment default."
        +}

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    Persistent shared memory for AI coding agents that turns a folder of markdown files into searchable memory across sessions, repos, and machines.
    5 npm
    Functional Source , Version 1.1, MIT Future
  • A
    license
    A
    quality
    C
    maintenance
    Persistent shared memory for AI coding agents. Stores facts as entity/key/value triples with hybrid semantic search, task checkpoints, and conflict resolution — shared across Claude Code, Codex CLI, and GitHub Copilot.
    16
    235 npm
    5
    AGPL 3.0
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.