Skip to main content
Glama

Server Details

Autonomous dev team steered from chat: plain-English requests in, tested merged PRs out.

Ownership verified
Status
Healthy
Uptime
99.8% over 42 days
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL
Repository
jamie7893/keelen-mcp
GitHub Stars
0
Server Listing
keelen-mcp

TDQS

A3.9/5.0

Scored across 41 tools

Disambiguation4/5

Most tools target clearly distinct resource+action pairs (projects, roadmap items, tasks, reviews), and the densest cluster — resolve_escalation vs retry_blocked_task vs replan_task vs rearm_roadmap_item vs resolve_platform_policy_conflict — is explicitly cross-differentiated in each description (e.g. acknowledgement-only vs retry-eligible vs unretryable). Minor residual overlap between archive_project/delete_project and between the several get_*_findings tools, but descriptions disambiguate well.

Naming Consistency4/5

Consistent snake_case with a strong verb_noun convention throughout (create_project, list_roadmap, cancel_roadmap_item, get_request_status). A few minor deviations: project_status reads as a noun-first status call rather than get_project_status, and signup/verify_email are single-token verbs outside the pattern.

Tool Count3/5

41 tools is heavy and sits in the over-large band, but the domain is genuinely broad (auth/onboarding, projects, roadmap, tasks, escalations, three review types, billing, dashboard) and nearly every tool maps to a distinct operation. It is defensible but will strain an agent's selection budget.

Completeness4/5

Strong lifecycle coverage: create/import/archive/delete projects, full roadmap ordering/pinning/cancel, task close/replan, escalation recovery paths, security/legal/control-gap reviews with read-back, plus account and billing flows. The main gap is that there is no direct task read/list tool (tasks surface only via project_status and list_escalations) and no roadmap-item content editing beyond ordering.

Available Tools

41 tools
answer_requestA
Destructive
Inspect

Answer a thread's clarifying questions (status must be awaiting_answers).

`answers` is a list of {"idx": <int from get_request_status>, "answer_md":
<str, 1..2000 chars>}. Answer EVERY question exactly once. Flips the thread
back to intake_pending; poll get_request_status again.
ParametersJSON Schema
NameRequiredDescriptionDefault
answersYes
thread_idYes
project_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses a concrete side effect: 'Flips the thread back to intake_pending'. It also instructs to poll get_request_status again, which reveals the post-condition and expected follow-up. This goes beyond the annotations, which only indicate that the operation is not read-only and may be destructive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the core action and precondition. Every sentence adds value: the answers format, the exactly-once requirement, and the state transition. No filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity, the description covers the key requirements: how to format answers, the mandatory status condition, the effect on thread state, and the next step. An output schema exists, so return-value documentation is not the description's burden. Could add error-handling behavior for invalid idx values, but not essential.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, but the description compensates by detailing the `answers` parameter structure with `idx` and `answer_md`, including types, source, and length constraints. It also enforces 'Answer EVERY question exactly once'. Project_id and thread_id are not described, but their purpose is evident from their names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'Answer a thread's clarifying questions' and immediately specifies the required state ('status must be awaiting_answers'). This makes the tool's purpose clear, though it does not explicitly contrast it with sibling tools like refine_request or get_request_status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly states when to use the tool: the thread must be in awaiting_answers status. It also gives operational guidance, like answering every question exactly once and polling get_request_status afterward. However, it doesn't name alternatives or explicitly state when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

archive_projectA
Destructive
Inspect

Archive a project (reversible shelve) — frees a project slot in-tool.

Stops any in-flight machine, then flips the project to `archived`: it drops
out of the per-tier project cap (freeing a slot for a new project) and the
loop stops dispatching it, but the project + its history are kept and can be
restored from the dashboard project page. Stopping a machine DESTROYS the
running process state — restoring the project later cannot bring that
execution back, so this is destructive despite being reversible. A replay on
an already-archived project changes no row, but is not idempotent end to end.
Prefer this over `delete_project` unless you specifically want the project
gone. Owner-scoped (an MCP key is owner-only); a project not in the
workspace 404s.
ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds significant behavioral detail beyond the annotations: it explains that stopping a machine destroys running process state and that restoration cannot recover it, that the operation is destructive despite being reversible, and nuances about replay not being idempotent. It also covers access scoping and 404 behavior. These are not present in the annotations and are critical for the agent to avoid misuse.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence contributes. It opens with a clear one-line purpose, then logically elaborates on effects, destructiveness, idempotency, and alternatives. No fluff or redundancy; the structure supports quick parsing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the single parameter, the annotations covering destructive and idempotent hints, and an output schema (which is assumed to define the return), the description covers all necessary operational aspects: side effects on machine state, slot freeing, restoration path, destructive nature, replay behavior, and access restrictions. An agent can confidently decide when and how to call this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema coverage, the description must clarify the parameter. It implies project_id refers to a project in the workspace and adds that it must be owner-scoped and that a non-workspace project 404s. Though it does not explicitly define project_id as the identifier, the context makes it clear. This adds useful constraints beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action: 'Archive a project (reversible shelve)' and explains the specific outcome of freeing a project slot. It distinguishes this from delete_project by explicitly naming the alternative and contrasting the behaviors. No ambiguity or tautology.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly instructs when to use this tool over delete_project: 'Prefer this over delete_project unless you specifically want the project gone.' It also provides operational constraints: owner-scoped and a 404 if the project is not in the workspace, which helps the agent decide if the call is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

cancel_roadmap_itemAInspect

Cancel / close a single roadmap item (get its id from list_roadmap).

Use this to close a DELIVERED or duplicate item that keeps re-parking: when
the work already shipped, every expand produces no dev-ready tasks and files a
recurring `roadmap_item_parked` escalation you have to keep acking. Cancelling
drops the item out of the expand queue AND resolves any open expand-lane
escalation for it. Business-level replay guard: an already expanded/cancelled
item changes no state, and the dashboard can restore a cancelled item.
Refuses (409) while the item is actively being expanded (retry once that iter
ends).
Returns {id, status, changed, next_step}.
ParametersJSON Schema
NameRequiredDescriptionDefault
item_idYes
project_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses far more than the annotations: cancellation removes the item from the expand queue, resolves open expand-lane escalations, acts as a no-op on already expanded/cancelled items, supports dashboard restore, and refuses with 409 during active expansion. This is rich operational context that the annotations alone do not provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every clause carries operational weight: trigger, effect, escalation resolution, guardrail, conflict behavior, and return shape. The purpose is front-loaded in the first sentence, and inline code values make it easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers triggers, effects, replay safety, conflict/retry behavior, and return shape, and an output schema exists for the return value. The only completeness gap is the undocumented project_id parameter, which keeps it from a top score.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for both parameters. It tells the agent to get item_id from list_roadmap, but project_id is never explained or associated with the item. Both parameters remain opaque strings, leaving a significant semantic gap for a correct call.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence names a specific verb and resource ('Cancel / close a single roadmap item') and tells the agent where to obtain the id from list_roadmap. Scoping to a single item clearly distinguishes it from bulk or inverse operations such as rearm_roadmap_item. No sibling ambiguity remains.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Use this to close a DELIVERED or duplicate item that keeps re-parking' is an explicit trigger condition, and the description explains the underlying symptom (recurring roadmap_item_parked escalations). It gives clear context but does not explicitly name alternative tools or when-not-to-use cases, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

clear_horizon_pinAInspect

Clear a queued roadmap item's horizon pin (now/next/later → none).

A horizon pin dominates the queue sort, so `reorder_roadmap` cannot move a
pinned item out of its band — a stale `now` pin on a delivered/duplicate item
clogs the front of the queue. This unpins it and reprices the queue so the
item follows plain priority order again (and reorder_roadmap can then move it).
Only queued items carry a settable pin (in-flight / shipped items derive
theirs), so this refuses (422) on a non-queued item — use cancel_roadmap_item
to close a delivered item. Returns {id, previous_pin, horizon_pin, next_step}.
ParametersJSON Schema
NameRequiredDescriptionDefault
item_idYes
project_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are mostly false and provide no safety detail, so the description carries the full behavioral burden. It discloses the side effect (repricing the queue so the item follows plain priority order), the 422 refusal on non-queued items, and the expected return fields. No annotation contradiction exists.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core action is front-loaded and the supporting rationale earns its place by explaining why reorder_roadmap won't work. It is slightly long, but each sentence adds usage or behavioral value, and the structure moves from action to consequence to alternative to response.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the operation, side effects, error conditions, alternatives, and return contract. With an output schema present and detailed contextual guidance, nothing critical is missing for an agent to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains that the item is a queued roadmap item and that unpinning re-enables reordering, which gives semantic meaning to item_id. However, it never explicitly names project_id or item_id nor explains their roles, leaving some mapping to inference.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Clear a queued roadmap item's horizon pin' and explains the state transition (now/next/later → none). It clearly differentiates from sibling tools by referencing reorder_roadmap and cancel_roadmap_item, so an agent can immediately tell what this tool does and what it is not.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the exact scenario when to use the tool: a stale pin on a delivered/duplicate item clogs the queue, and reorder_roadmap cannot move a pinned item. It explicitly names the alternative for closing delivered items (cancel_roadmap_item) and notes that non-queued items are refused. This gives clear when-to-use and when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

close_taskA
Destructive
Inspect

Close a task that should not be built — a duplicate, or work already shipped.

Use this when a task on the board is obsolete: the change already landed in
another PR, a sibling task covers it, or the user changed direction. The
task is marked `cancelled` and keeps ALL of its history (acceptance
criteria, QA steps, iterations) — nothing is deleted.

`reason` is REQUIRED and is recorded on the audit trail; say why in one
line. Pass `superseded_by_pr_number` (or `superseded_by_task_id`) when the
work was genuinely delivered somewhere else — that records verified
provenance instead of a bare abandon. Closing does NOT claim the content is
on the default branch, so any task that declared a dependency on this one
keeps waiting; deliver or re-plan those separately.

Refuses with 409 while the task is being worked on by a running iteration
(stop the machine first, or wait for it to finish). A task in another
workspace 404s. Business-level replay guard: closing an already-closed task
changes no task state. The call is not idempotent end to end — an
authenticated request also persists credential-use state.

The response echoes `open_tasks`: how many tasks are still open on the
project, counted after the close commits. Check it — if it did not drop,
the ticket was already closed and this call changed nothing. It is the same
count `project_status` returns, and neither counts a closed task as open.

Prefer this over leaving a dead task on the board: unfinished tasks count
against the project's planning capacity, so stale duplicates quietly stop
new roadmap items from being expanded.
ParametersJSON Schema
NameRequiredDescriptionDefault
reasonYes
task_idYes
project_idYes
superseded_by_task_idNo
superseded_by_pr_numberNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes far beyond the annotations: it explains that the task is marked cancelled with all history preserved, that nothing is deleted despite destructiveHint=true, that reason is recorded on the audit trail, that dependencies keep waiting, that 409/404 responses occur in specific states, and that the call is not idempotent end-to-end due to credential-use state persistence. No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but information-dense, and every section earns its place by covering semantics, error behavior, side effects, response interpretation, and usage rationale. The opening sentence front-loads the core purpose, and subsequent paragraphs are structured around distinct behaviors rather than repeating the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity, the description is unusually complete: it explains required vs optional parameters, error conditions, side effects, the open_tasks response count, how to verify the close actually happened, and why closing is preferred over leaving stale tasks. Since an output schema exists, the description does not need to enumerate the full response shape.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the parameter-semantics burden. It adds meaningful guidance: reason is REQUIRED, one line, and recorded on the audit trail; superseded_by_pr_number and superseded_by_task_id should be passed when work was genuinely delivered elsewhere to record verified provenance. This is essential context the schema alone does not provide.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Close a task that should not be built' and clarifies the intended meaning with concrete examples ('a duplicate, or work already shipped'). It clearly targets board tasks rather than projects, roadmap items, or requests, so it is distinguishable from sibling tools without opening their schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says when to use the tool: when a task is obsolete because the change landed elsewhere, a sibling task covers it, or the user changed direction. It also gives a when-not condition: a 409 occurs while a running iteration is working on the task, so stop the machine or wait. It even advises preferring this over leaving a dead task on the board.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

connect_githubAInspect

Get the link that connects a GitHub account, so work can reach real repos.

Returns an `install_url` — send it to the user to open in a browser. They
pick the GitHub account/org, approve the install, and land on a "connected"
page; then poll get_onboarding_status() until github_connected is true. The
link expires in 10 minutes — call connect_github() again for a fresh one.
ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide no positive hints (all false), so the description carries the burden of behavioral disclosure. It covers the return value, browser-based approval flow, polling requirement, and the 10-minute link expiration. It could explicitly mention whether calling repeatedly invalidates previous links or whether there are side effects beyond issuing a link, but the provided details are useful.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and well-structured. The first sentence states the purpose, the second explains the return value and user flow, and the third covers expiration and re-invocation. Every sentence earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter tool with an output schema, the description is complete. It explains the returned install_url, the browser flow, how to detect successful connection via get_onboarding_status, and the expiration behavior. Nothing essential is missing for an agent to invoke and handle the result correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the description does not need to explain parameter meaning. The description still adds value by clarifying what the return value represents and the external action it triggers, which is more than necessary for a parameterless tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: getting a link that connects a GitHub account so work can reach real repos. It clearly differentiates from siblings like list_github_repos and get_onboarding_status by focusing on initiation of GitHub connection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly explains when and how to use the tool: send the returned install_url to the user, wait for them to approve in the browser, then poll get_onboarding_status until github_connected is true. It also covers link expiration and the need to call connect_github() again, which is clear operational guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

control_schedulerA
Destructive
Inspect

Start, pause, or resume a project's autonomous work.

`action` is one of: "enable" | "disable" | "resume" | "process_now".
- enable/disable flip scheduler_enabled (the loop dispatches only enabled,
  status='active' projects).
- resume clears a pause (peak/backoff/manual) so the project dispatches again.
- process_now durably prioritizes and immediately attempts the next intake
  batch, independent of the dev scheduler switch. A gate returns a typed
  reason and recovery step instead of a spawn promise.
- enable/resume can execute pending refinement that DELETES tasks, and
  process_now can start an intake batch, so treat this tool as destructive
  even when the echoed scheduler state looks unchanged.

Returns the resulting scheduler state.
ParametersJSON Schema
NameRequiredDescriptionDefault
actionYes
project_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the destructiveHint annotation by specifying exactly what can be destroyed: enable/resume may execute pending refinement that DELETES tasks, and process_now may start an intake batch. It also warns that destructive effects can occur even when the echoed scheduler state looks unchanged, and clarifies the gate returns a typed reason/recovery step instead of a spawn promise.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose line is front-loaded and the bullets are efficient, with no filler around the operations. A few phrases like 'spawn promise' and 'dev scheduler switch' assume internal familiarity, which keeps it from a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter tool with a rich annotations block and an output schema, this covers the action semantics, destructive side effects, and return state. It does not discuss permissions, idempotency, or project existence, but none of those are implied as required by the schema or annotations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With schema coverage at 0%, the description carries the burden; it fully defines the action values and their distinct effects. project_id is only implied by 'a project's autonomous work,' but as one of the two required fields it needs no deep explanation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence names a specific verb set—start, pause, resume—with a clear resource ('a project's autonomous work'), and the bullet list disambiguates four action modes. It does not explicitly differentiate from sibling tools such as project_status or refine_request, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The action bullets give concrete conditions: enable/disable govern scheduler_enabled for active projects, resume clears peak/backoff/manual pauses, and process_now applies independently of the dev scheduler switch. This is solid context for choosing an action, though it never names a sibling as the alternative, so there are no exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create_projectA
Destructive
Inspect

Start building something new: creates a GitHub repo and begins work on it.

Use this ONLY when the user wants a NEW repo scaffolded. If they already
have a repo, use import_project(repo_full_name) instead — this tool would
create a second, empty one beside theirs (list_github_repos() browses what
the workspace can see).

Scaffolds a new GitHub repo, a bootstrap-mode project, and submits
`build_description` as the project's first Roadmap Request. `name` is a
concise GitHub short repo slug (no owner); `project_kind` is REQUIRED and
one of library | node_library | python_library | service | cli | web_app |
godot_game | roblox_game; `preview_command` is required iff
`project_kind == 'web_app'`.
`engine` is
OPTIONAL — one of claude_code | codex | glm | kimi | grok (defaults to
claude_code); codex, glm, kimi, and grok require the workspace to have a
matching connected credential.
`org` is OPTIONAL — a GitHub organization login to create the repo inside
(e.g. your company org); omit it to land the repo on a member's personal
account. `private` defaults to True.
`ci_runs_on` is OPTIONAL — the CI runner labels for the scaffolded workflow,
e.g. ["self-hosted", "linux", "x64", "my-fleet"]. Omit it to inherit the
workspace default (ubuntu-latest if unset). Labels no registered org runner
carries are rejected, because GitHub would queue such a job forever rather
than fail it.
`framework` is OPTIONAL and `web_app`-only — one of vite | next (defaults to
vite). It picks the scaffolded frontend rails: `vite` a vanilla-TypeScript
SPA, `next` a Next.js app-router app. Passing it with any other
`project_kind` is an error.

The repo is created on the GitHub account of a workspace member with
repo-create OAuth access (this path has no specific caller user), so the
returned `repo` owner is whichever member's token resolved (or the chosen
`org`). If no member has repo-create access — or the resolving member can't
create in `org` — the call returns an actionable error.

Returns {project_id, repo, thread_id, next_action, poll_after_seconds,
next_step}; follow next_step (poll get_request_status with the returned
thread_id). On the rare arm where the first Request failed to submit,
next_action is "call_tool" with next_tool="submit_request".

If scaffolding fails after the repository has been created, best-effort
compensation deletes that just-created repository so a retry can reuse the
name — this tool can therefore remove external state it created moments
earlier.
ParametersJSON Schema
NameRequiredDescriptionDefault
orgNo
nameYes
engineNo
privateNo
frameworkNo
ci_runs_onNo
project_kindYes
preview_commandNo
build_descriptionYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already declare destructiveHint and readOnlyHint, and the description goes further by disclosing the OAuth-based account resolution, the possibility of actionable errors, the fallback next_action, and the best-effort compensation that deletes a just-created repo on failure. It also warns that labels not registered on an org runner will be rejected, which is valuable non-obvious behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well-structured: purpose and routing are front-loaded, followed by parameter semantics, then behavior and response handling. Every sentence delivers a distinct constraint or useful fact, with no filler or repetition of schema fields that already have obvious meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex, mutating 9-parameter tool, the description covers purpose, alternatives, parameter constraints, auth path, error behavior, compensation, and next-step guidance. The output schema already documents the return object, so the description appropriately focuses on how to consume next_step rather than restating fields.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Because the schema has 0% description coverage, the description carries the full burden and fully compensates. It explains every parameter, provides allowed values for project_kind, engine, and framework, clarifies conditional requirements like preview_command for web_app, and documents defaults such as private=True and ci_runs_on inheritance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a clear action: "Start building something new: creates a GitHub repo and begins work on it." It also explicitly contrasts with import_project and mentions list_github_repos, so an agent can distinguish it from closely related sibling tools without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use this tool ("ONLY when the user wants a NEW repo scaffolded") and when not to (existing repo should use import_project instead). It also gives precise conditional rules for project_kind, preview_command, engine credential requirements, org usage, and framework restrictions, leaving no ambiguity about valid invocation contexts.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

delete_projectA
Destructive
Inspect

Soft-delete a project (the harder option) — frees a slot and hides it.

Stops any in-flight machine, then flips the project to `deleted`: it
disappears from `list_projects`, drops out of the project cap, and the loop
stops dispatching it. Deleting ALSO permanently purges the project's stored
review artifacts. The row is retained for audit and re-importing the repo
restores that deleted row in place (unlike `archive_project` there is no
in-tool restore), but the purged artifacts are gone for good and the stopped
workers' process state cannot be recreated. A replay on a deleted project
changes no row; the call is not idempotent end to end. Owner-scoped; a
project not in the workspace 404s.
ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes far beyond the annotations: it discloses permanent purge of review artifacts, irrecoverable worker state, stopping in-flight machines, dropping out of project cap, list visibility changes, replay behavior, non-idempotency, and owner-scoping with 404s. This is exactly the behavioral depth needed for a destructive operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence adds necessary operational detail. It is front-loaded with the core meaning, then covers side effects, restore semantics, idempotency, and scoping without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, non-idempotent tool with nuanced side effects, the description covers everything an agent needs: what is destroyed, what is retained, behavior on replay, alternative tool, and failure case. An output schema exists, so return-value documentation is not required.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description carries the parameter-semantics burden. It compensates by stating that the project must be owner-scoped and that a project not in the workspace 404s, which tells the agent what values are valid. The single project_id parameter is otherwise self-evident from the tool name.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'Soft-delete a project' and immediately states the core effects (frees a slot, hides it). It clearly distinguishes this tool from archive_project by contrasting restore behavior, so an agent can tell them apart.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear decision-relevant context by naming archive_project as the alternative and explaining the key difference: delete has no in-tool restore while archive_project does. It does not phrase an explicit 'use this when...' instruction, but the contrast makes the appropriate use case inferable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_billingAInspect

Billing status + a Stripe checkout link when a NEW subscription is needed.

`plan` is one of starter | pro | agency (default starter). When the
workspace has NO live subscription and needs one (open-signup unpaid,
churned, or converting from a free/trial tier), returns a `checkout_url`
with next_action "browser" — send it to the user to open in a browser (the
one setup step that can't happen in chat). Compute unlocks automatically
once payment completes (a Stripe webhook flips the workspace to active);
you do not need to block on it. A past_due workspace gets NO checkout —
the fix is a card update in the dashboard billing page (a new checkout
would create a second subscription); follow next_step. Subscribed or
suspended-with-subscription states return checkout_url=None with an
explanatory next_step.
ParametersJSON Schema
NameRequiredDescriptionDefault
planNostarter

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

All annotations are false, so the description carries the full burden and delivers richly: it discloses that payment completion flips the workspace to active asynchronously via a Stripe webhook, warns that a past_due workspace must not receive a checkout to avoid creating a duplicate subscription, and specifies per-state response shapes (checkout_url vs None with next_action/next_step). No contradiction with readOnlyHint=false or idempotentHint=false.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with a one-line summary, and the remaining roughly 140 words are dense with non-redundant operational details. Every sentence addresses a distinct billing state or hazard; there is no filler or repetition of schema content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with an output schema present, the description covers every agent-relevant decision: when to send checkout links to a browser, the async webhook unlock behavior, the past_due exclusion, and subscribed/suspended return behavior. The only minor omission — error handling for an invalid plan value — is negligible given the documented enum.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% with no enums, so the description must compensate and does: it enumerates valid values ('starter | pro | agency') and the default ('default starter'). It stops short of explaining how plan maps to checkout pricing or limits, but for a single well-named parameter the semantic role is clear from context.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific verb and resource ('Billing status + a Stripe checkout link') with a precise condition ('when a NEW subscription is needed'). It distinguishes from all sibling tools — none other handles billing — and the remaining text clarifies boundary cases (past_due, subscribed, suspended) rather than obscuring the core purpose.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit conditional routing: generate a checkout_url and send it to the user for open-signup unpaid, churned, and free/trial states; never issue a checkout for past_due workspaces, with the rationale that it would create a second subscription; follow next_step for subscribed/suspended states. It also tells the agent when not to act — 'you do not need to block on it' — and describes the dashboard card-update alternative by action rather than by tool name.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_control_gap_findingsAInspect

Read a project's security-control observations and review boundaries.

The result groups records by evidence class without a total or an overall
framework outcome. Read `not_determinable` and `outside_review_scope`
before describing any observation. An empty group does not establish that
a control is in place, and `disclaimer_md` must reach the user.

`limit` defaults to 25 (max 100). `include_all` includes resolved,
dismissed, and out-of-scope records. This read-only tool has no compute
quota and remains available after the trigger closes for a project.
ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
project_idYes
include_allNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description explicitly calls itself 'read-only' and says it 'remains available,' but the annotations set readOnlyHint to false, implying it may modify state. This is a direct contradiction between the description and the structured annotation, which is a serious inconsistency for an agent deciding safety.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with purpose, then groups behavioral caveats and parameter details in separate sentences. Every sentence provides distinct, necessary information with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite the contradiction, the description is impressively complete: it explains grouping by evidence class, absence of totals, that empty groups do not prove controls, that disclaimer_md must reach the user, and the effects of include_all. Since an output schema exists, it need not describe return values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description carries the full burden and delivers: it explains that limit defaults to 25 with a max of 100 and that include_all broadens the result set to resolved, dismissed, and out-of-scope records. This adds meaning the schema lacks.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Read a project's security-control observations and review boundaries,' which names a specific verb and resource. It clearly differentiates from the sibling tool 'run_control_gap_review' by emphasizing this is a read operation, and the 'read-only' statement reinforces the separation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides practical context: the tool is read-only, has no compute quota, and remains available after the trigger closes. This implies when to use it, but it never explicitly contrasts it with alternatives like 'run_control_gap_review' or states when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_onboarding_statusAInspect

Check what is set up so far and what to do next to get building.

Reports engine_connected / github_connected / project_count / payment_status
and a `next_action` string you should follow VERBATIM, polling this tool
between steps. `next_action` is "call_tool" (call the tool named in
`next_tool`, following `next_step`) until onboarding is complete, then
"done". (`next_action_detail` echoes the pre-2026-07-29 dict shape and is
DEPRECATED — it is removed 2026-10-29; read next_action/next_tool instead.)
- engine step: send the user the dashboard /login link. Engine subscriptions
  (Claude / Codex / GLM) are connected in the DASHBOARD for security —
  NEVER ask for or paste engine credentials in this chat.
- github step: call connect_github() for an install link.
- project step: create_project(...) for a new repo, or
  import_project(repo_full_name) for an existing one.
- launch step: poll get_provisioning_status(project_id) until ready.
Re-checking is YOUR job — the server does not push.
Also returns a `usage` block (pool / daily / machine-hours counters + tier
caps) for capacity-aware automation clients.
ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds substantial behavior beyond the all-false annotations: it warns never to ask for or paste engine credentials, explains that engine subscriptions connect through the dashboard, marks next_action_detail as deprecated with a removal date, and documents the usage block for capacity-aware automation. Nothing here contradicts the annotations; the polling/check framing is consistent with a status inspection tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but tightly structured: purpose first, then next_action protocol, per-step routing, deprecation notice, security directive, and polling responsibility. Each sentence earns its place and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter polling tool with an output schema, the description covers the return fields, the action protocol, step-specific tool routing, deprecation timing, security constraints, usage block, and server push behavior. An agent has everything needed to call and interpret this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and the schema documents 100% of nothing, so the baseline is 4. The description instead clarifies the output contract, especially the next_action/next_tool/next_step protocol, which is the meaningful semantic content an agent needs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific purpose: 'Check what is set up so far and what to do next to get building.' It then enumerates the exact fields reported (engine_connected, github_connected, project_count, payment_status, next_action), which clearly distinguishes this from sibling tools like get_provisioning_status or connect_github.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is explicit: poll the tool between steps, follow next_action VERBATIM, and branch to connect_github, create_project, import_project, or get_provisioning_status depending on the current step. It also states that re-checking is the client's job because the server does not push, removing any ambiguity about when and how often to call it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_product_goalAInspect

Read a project's current product goal (the outcome work is aimed at).

READ THIS BEFORE YOU CALL `set_product_goal`: the setter REPLACES the whole
document rather than appending to it, so writing without reading first
silently discards whatever the user already recorded. To add a line, read
the current text, edit it, and set the full result back.

Returns {project_id, product_goal_md, updated_at}. `product_goal_md` is
None when no goal has been set. Tenant-scoped: a project not in the
caller's workspace 404s.
ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds excellent behavioral detail: it documents the return shape, the None-when-unset case, the 404 on tenant-scoped access, and the setter's replace-not-append behavior. However, the annotations declare readOnlyHint=false while the description explicitly says 'Read,' which is a direct contradiction with the structured metadata. Per the rubric, any such contradiction drops this dimension to 1.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is purpose-built with no filler: the core read action comes first, the critical setter warning is clearly highlighted, and the return/error behavior is summarized in compact sentences. Every paragraph earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter read tool, the description covers the default response, the missing-goal case, tenant scoping, error behavior, and the key interaction with `set_product_goal`. An output schema also exists, so the description does not need to further explain return values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema provides no description for `project_id`, so the description partially compensates by noting that a project outside the caller's workspace 404s, clarifying the meaning and failure behavior of the parameter. It does not spell out type or format, but the parameter name and the returned `project_id` make the semantics unambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Read a project's current product goal,' and even defines what a product goal is ('the outcome work is aimed at'). It is clearly distinguishable from the sibling `set_product_goal` (read vs. write) and from `get_product_vision` (goal vs. vision).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly instructs the agent to read before calling `set_product_goal`, establishing a clear and important when-to-use relationship. It does not enumerate other alternatives or exclusion conditions, but the read-vs-set contrast plus the tenant-scoping note gives sufficient usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_product_visionAInspect

Read a project's current product vision (what the product is for).

READ THIS BEFORE YOU CALL `set_product_vision`: the setter REPLACES the
whole document rather than appending to it, so writing without reading
first silently discards whatever the user already recorded. To add a line,
read the current text, edit it, and set the full result back.

Returns {project_id, product_vision_md, updated_at}. `product_vision_md` is
None when no vision has been set. Tenant-scoped: a project not in the
caller's workspace 404s.
ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide no hints (all false), so the description carries the full burden. It discloses the return shape, the None case, tenant scoping, and even the setter's destructive behavior. This goes beyond minimal disclosure and gives an agent everything needed to reason about side effects and safety.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence earns its place. The purpose is front-loaded, the warning about the setter is critical and concise, and the return/tentant details are compact. The multi-sentence structure is warranted by the volume of important context; nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple getter with one parameter and an existing output schema, the description covers all necessary behavior: what it returns, potential None, tenant scoping, and the critical relationship to the setter. Nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With a single required parameter (project_id) and 0% schema coverage, the description does not explain the parameter's format or origin. However, the name is self-explanatory and the tool is a getter; the description adds no literal parameter meaning, but the context of 'project' is inherent. A 3 is appropriate for not adding explicit value, though the parameter's purpose is obvious from the tool name.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a clear verb+resource: 'Read a project's current product vision (what the product is for).' It immediately distinguishes itself from the sibling setter by name, and the parenthetical clarifies the meaning of 'product vision.' No ambiguity about what this tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly warns against calling set_product_vision without reading first, explaining the replace-vs-append behavior and the risk of data loss. Also states tenant scoping and 404 behavior. This is direct, actionable guidance for when to use the tool and what to expect.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_provisioning_statusAInspect

Check whether a new project has finished setting up and is ready to build.

Returns overall (provisioning | ready | errored), a 7-stage checklist, and a
user-facing error_kind when a stage failed. next_action is "wait" with
poll_after_seconds (~10s) while provisioning or errored, and "done" when
overall is 'ready' — then steer the loop with submit_request(project_id,
text). Reaching 'ready' also PERSISTS the owner's onboarding-completion
state, so this is not a pure read. Tenant-scoped: a project not in the
caller's workspace 404s.
ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the annotations by disclosing the non-read side effect: 'Reaching ready also PERSISTS the owner's onboarding-completion state, so this is not a pure read.' It also exposes tenant-scoping behavior with 404s and explains polling semantics, giving agents an accurate behavioral model.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: purpose first, then return values, then the actionable next step, then the important side effect and scoping caveat. Every sentence earns its place and the structure supports quick agent parsing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter polling tool with an output schema, the description covers the essential operational details: return statuses, error conditions, polling interval, terminal action, side effects, and tenant scoping. Nothing critical is missing for an agent to invoke and act on this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema provides only a bare project_id string with no description, so the description carries the interpretive burden. It adds meaningful semantics by clarifying that a project outside the caller's workspace causes a 404 and that the tool applies to newly created projects. It could go further by explaining how to obtain project_id, but the name and context make it largely self-evident.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'Check whether a new project has finished setting up and is ready to build.' It also enumerates the exact return shape (overall status, 7-stage checklist, error_kind), making the tool's purpose unambiguous and easy to distinguish from sibling status tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit control-flow guidance: poll every ~10s while provisioning or errored, and call submit_request(project_id, text) once overall is 'ready'. It clearly states when to use the result, though it does not explicitly name alternative tools or state when not to use this one.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_request_statusBInspect

Check what happened to a request, and read any questions it asked back.

Intake is async (~5min cadence) - poll periodically. Read `next_action`:
"wait" (still processing), "answer_questions" (call answer_request with one
answer per question), "done" (see generated_roadmap_item_ids), "cancelled"
(terminal, no items), "failed" (see intake_failure_reason).
ParametersJSON Schema
NameRequiredDescriptionDefault
thread_idYes
project_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description uses read-oriented language ('Check', 'read') and describes a status retrieval, which implies a read-only operation. However, the annotations set readOnlyHint: false, indicating the tool may modify state. This is a direct contradiction. Additionally, the description provides no other behavioral disclosures beyond the async cadence and next_action semantics, which are useful but do not override the inconsistency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise: two sentences plus a compact enumeration of next_action values. It leads with the core purpose, then adds necessary operational details. The structure is efficient and easy to parse, though the list of values adds some length – still acceptable given the benefit.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description explains the asynchronous nature, polling cadence, and the meaning of each next_action value, which is valuable even with an output schema. However, it fails to document the two required parameters, leaving a significant gap that could cause incorrect invocation. The output schema likely details return fields, so the description's focus on behavior is partially sufficient, but the parameter gap undermines completeness.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description does not explain project_id or thread_id at all. It mentions 'a request' but provides no context for what these identifiers are or how they relate. Since there is no parameter documentation in the schema either, the agent has no guidance on how to correctly supply these required fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Check what happened to a request, and read any questions it asked back' – a specific verb and resource. It goes further to enumerate the possible values of next_action, which distinguishes it from sibling tools like answer_request (which handles responding to questions) and refine_request. The purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit usage context: 'Intake is async (~5min cadence) - poll periodically' advises when to call. It also gives conditional guidance on what to do based on next_action, including referencing the sibling tool answer_request for the 'answer_questions' case. While it doesn't explicitly say when *not* to use this tool, the polling guidance and follow-up actions make the intended usage clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

import_projectA
Destructive
Inspect

Connect an EXISTING GitHub repo as a Keelen project.

This is the counterpart of create_project: use THIS tool when the user
already has a repo, and create_project only to scaffold a brand-new one.

`repo_full_name` is "owner/repo" — it MUST be visible to the workspace's
GitHub connection (list_github_repos() to browse; a non-visible repo 404s).
`engine` is OPTIONAL — one of claude_code | codex | glm | kimi | grok
(defaults to claude_code); codex/glm/kimi/grok require a matching connected
credential.
`build_description` is OPTIONAL but STRONGLY recommended — a plain-language
"what should Keelen build first?" submitted as the project's first Request so
the loop has work; an imported project with no Request sits idle until you
call submit_request(project_id, ...).

`project_kind` is OPTIONAL — one of library | node_library | python_library |
service | cli | web_app | godot_game | roblox_game | unknown. Omit it and the
kind is auto-detected. PASS IT when the repo is a MONOREPO (apps in
subdirectories), a stack with no standard root manifest (Java, Ruby, PHP,
.NET, Elixir), or when you want a classification detection cannot infer —
in those cases detection yields "unknown", which BLOCKS the dev lane until
someone overrides it. A value you pass is authoritative and is never
overwritten by later auto-detection. `stack` is the OPTIONAL language axis
(python | node | rust | go | cpp) for a language-agnostic kind.
`preview_command` is REQUIRED when project_kind is "web_app" (the command
that serves the app locally, e.g. "npm run dev") and optional otherwise,
where it overrides the detected one.

Re-importing the same live repo is replay-guarded (returns the existing
project with already_exists=True), and re-importing a SOFT-DELETED repo
RESTORES that deleted row in place — it does not create a fresh project.
Restored queued work can include refinement that deletes tasks, and the call
is not idempotent end to end. On a plan with no scheduled-project allowance
the project is still created but with the loop OFF — next_step then steers to
get_billing(). Otherwise follow next_step and poll
get_provisioning_status(project_id).
ParametersJSON Schema
NameRequiredDescriptionDefault
stackNo
engineNo
project_kindNo
repo_full_nameYes
preview_commandNo
build_descriptionNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations provide only the basic safety hints (destructive=true, not idempotent, open world), and the description substantially extends them: replay-guard returning already_exists=True, soft-deleted repos being RESTORED in place rather than recreated, restored queued work that 'can include refinement that deletes tasks,' and the no-scheduled-project-allowance branch that creates the project with the loop OFF. It also surfaces the 404 error condition for non-visible repos and the credential requirement for non-default engines. No contradiction — destructiveHint=true is consistent with the disclosed task-deletion behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded in the first sentence and every sentence earns its place — parameter semantics, the create_project contrast, and the non-idempotency warning are all load-bearing for correct invocation. However, the content is a dense prose block without labeled sections or bullets; given the 6 parameters and several conditional branches, scannability would be improved by formatting, so it falls just short of a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex, destructive-capable import tool with 6 parameters, no schema descriptions, and a rich interaction surface, the description is essentially complete: full parameter semantics, error behavior (404), replay-guard and soft-delete restore edge cases, billing-plan branch, and the next_step/get_provisioning_status follow-up protocol. Nothing an agent needs to invoke it correctly or handle its outcomes is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% — the schema exposes bare titles and defaults only — so the description carries the entire burden and fully compensates. Every one of the 6 parameters gets meaning: repo_full_name's 'owner/repo' format plus visibility requirement, engine's full enum with default and credential constraint, project_kind's full enum with the auto-detection fallback caveat, build_description's role as the first Request, stack's language-axis semantics, and preview_command's conditional requirement for web_app. The enum values are enumerated inline even though the schema has no enums declared.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource — 'Connect an EXISTING GitHub repo as a Keelen project' — and immediately distinguishes itself from the sibling create_project: 'use THIS tool when the user already has a repo, and create_project only to scaffold a brand-new one.' The scope (existing repo, not scaffolding) is unambiguous and no other sibling could be confused with it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names create_project as the alternative and states which condition selects each tool. It also gives conditional usage rules: when to pass project_kind (monorepos, non-standard root manifests) versus omit it (auto-detection), when preview_command is mandatory (web_app), and when build_description is strongly recommended. Follow-up routing to submit_request, get_billing, and get_provisioning_status is spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_escalationsAInspect

See what the work is stuck on and waiting for a human decision about.

`project_status` only COUNTS open escalations; this returns each one with its
kind, reason, detail_md, recommended_action, and (when task-scoped) the
blocked task's title + PR url. A "forever-paused" project with no open
`project_pause` row is surfaced as a synthetic `orphan:<project_id>` row.
Each row carries a server-derived `task_retry_available` flag and an exact
`next_tool`: eligible recovery blocks route to `retry_blocked_task`, roadmap
parks to `rearm_roadmap_item`, platform conflicts to their structured
resolver, stale work to `replan_task`, and only safe auxiliary cards to
`resolve_escalation`.
ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses non-obvious behaviors beyond the sparse annotations: synthetic orphan:<project_id> rows for forever-paused projects, a server-derived task_retry_available flag, and exact next_tool routing rules. This gives the agent meaningful expectations about edge cases and side-channel behavior. There is no explicit contradiction with readOnlyHint=false, though the description does not promise side-effect-free behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The opening sentence front-loads the core purpose, and the remaining sentences add dense, useful detail: return fields, the synthetic-row edge case, and next_tool routing. Every sentence earns its place, and the structure is logical and scannable with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that an output schema exists, the description appropriately focuses on semantics rather than re-listing return types. It explains the important edge case, the server-derived flag, and how the output guides follow-up actions. For a single-parameter, data-returning tool, the description is complete enough for an agent to invoke it correctly and interpret its results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one required parameter, project_id, and schema description coverage is 0%. The description never explains what project_id values are valid, where the agent gets it, or any constraints/ownership semantics. It mentions <project_id> only as part of an output example, which does not compensate for the missing schema-level documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('returns each one') and clearly identifies the resource: escalations for a project. It explicitly differentiates from project_status by noting that project_status only counts escalations while this tool returns the full rows with their fields. An agent can immediately tell what this tool does and how it differs from a sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when project_status is insufficient ('only COUNTS open escalations') and describes how the returned next_tool field routes to retry_blocked_task, rearm_roadmap_item, replan_task, and resolve_escalation. It gives clear context for downstream tool selection, though it does not explicitly list exclusions such as 'do not use this when you already know the escalation ID and want to resolve it directly.'

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_github_reposAInspect

List repos the workspace's GitHub connection can see (for import_project).

Each entry has full_name, default_branch, private, language, pushed_at. Pass
a `full_name` to import_project(repo_full_name) to connect it. Reading the
list can MINT a short-lived GitHub installation token, so this is not a pure
read. Returns 409 if GitHub isn't connected yet — call connect_github()
first.
ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

It discloses the non-obvious side effect of minting a short-lived installation token, which is genuinely valuable because the annotations only say readOnlyHint is false without explaining why. It also documents the 409 failure mode and the prerequisite connection state, making the behavior transparent beyond structured annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and well-organized: purpose first, then output field details, then usage caveats and error handling. Every sentence adds necessary information and none is redundant with the schema or annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema already covers the return structure, so the description wisely focuses on what the schema cannot: side effects, prerequisites, error conditions, and downstream use. Nothing an agent needs to call this tool successfully is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters, so the schema leaves no input semantics to document. The description mentions full_name as an output field and points out how to feed it into import_project, which is useful context even though it is not a parameter explanation. The baseline for a no-parameter tool is satisfied.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the exact operation ('List repos') and the resource scope ('the workspace's GitHub connection can see'), and it anchors the purpose to import_project. This distinguishes it from sibling list tools and connects it to downstream use without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly frames usage as a precursor to import_project, warns that the call is not a pure read because it can mint a GitHub installation token, and gives a concrete prerequisite: call connect_github() first if a 409 is returned. This fully guides the agent on when and how to use the tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_projectsAInspect

List the projects in this workspace, with what each one is building.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description implies a read-only listing operation, but readOnlyHint is false, indicating the tool is not read-only. This is an annotation contradiction. The phrase 'with what each one is building' adds some return-content context, but the conflicting annotation is a serious transparency issue.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler. It efficiently conveys the action, scope, and the useful detail of what each project is building.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless list tool with an output schema available, the description provides the essential invocation context: list projects in the workspace and their purpose. No pagination, filtering, or return-format details are needed because there are no arguments and the output schema can describe the response.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and the schema fully documents that with an empty properties object and 100% coverage. There is nothing for the description to clarify about parameters, so the zero-parameter baseline applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Clearly states the action ('List') and resource ('projects in this workspace'), and previews the content ('what each one is building'). This distinguishes it from sibling project-management tools like create_project, archive_project, delete_project, and list_roadmap.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description makes the use case evident: get an overview of projects and their purpose in the current workspace. It does not explicitly state when not to use it or point to alternatives for status/roadmap details, but for a simple list tool the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list_roadmapBInspect

See what is planned for a project and in what order.

Queued rows also say whether the expand picker would elect them
(`electable`) and, when not, why it skips them (`skip_reason`:
crash_capped | noop_capped | crash_cooldown | noop_cooldown |
awaiting_intake_thread | held_on_open_pr | other). A skipped item ahead of
yours does not delay it.
ParametersJSON Schema
NameRequiredDescriptionDefault
statusNoqueued
project_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.2/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description says 'See what is planned,' implying a purely read-only operation, while the annotations mark readOnlyHint=false. That is a direct contradiction, so the rubric requires a score of 1. The additional detail about electable and skip_reason is valuable, but it cannot override the fundamental inconsistency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the core purpose, then adds dense but relevant behavioral detail about electability, skip reasons, and the practical effect of skipped items. Every sentence earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Because an output schema exists, the description doesn't need to restate return fields. It adds crucial behavioral nuance—electable, skip_reason, and the reassurance that a skipped item ahead doesn't delay yours. It falls just short of complete by not addressing what other status values might be passed or how they change the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must help explain parameters. It clarifies the status default by focusing on 'queued rows' and explains what such rows include, which partially documents the status parameter. However, it does not explicitly explain project_id or enumerate possible status values beyond queued, so compensation is only partial.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a clear verb and object: 'See what is planned for a project and in what order.' This immediately tells an agent what list_roadmap does and distinguishes it from mutation-focused siblings like reorder_roadmap, cancel_roadmap_item, and rearm_roadmap_item. It doesn't explicitly name sibling alternatives, but the core purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies this is the tool to use for viewing planned project work, especially queued items, and the detailed queued-row behavior orients an agent to the default status. However, it never states when not to use this tool or points to alternative tools such as project_status or list_projects, so usage guidance is mostly implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

open_dashboardAInspect

Get a one-click, pre-authenticated dashboard sign-in link for the owner.

The engine-connect step — and any dashboard task (billing card update, a
project page) — needs a signed-in browser. Because this server has already
authenticated the workspace owner, this mints a single-use magic-link login
token and returns a `/login?token=…` deep link: opening it signs the user
straight into the dashboard (no email round-trip, no password) and lands
them where onboarding left off. Send the user the returned `login_url`; it
works once and expires in 15 minutes — call again for a fresh one.
ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses important behavioral traits beyond the annotations: the token is single-use, expires in 15 minutes, bypasses email/password authentication, requires a pre-authenticated server context, and returns a deep link that lands the user at the proper onboarding location. This fully compensates for the sparse annotation hints.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Although the description is several sentences long, each sentence earns its place. The first sentence states the main purpose, and subsequent sentences provide needed context on when to use it, how the token works, the expiry, and how the result is consumed. No fluff or repetition exists.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no parameters, the description is exceptionally thorough: it explains the prerequisite (already authenticated owner), the behavior (single-use token, 15-minute expiry), the output (`/login?token=...`), when to use it, and how the user should receive it. This covers all practical usage scenarios despite minimal structured schema/annotation data.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters, so the input schema is trivially complete and the description needs no parameter clarification. The description adds meaning by explaining the output field (`login_url`), its characteristics, and how it should be used, which is valuable since there are no parameters to document.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific, action-oriented verb: 'Get a one-click, pre-authenticated dashboard sign-in link for the owner.' It clearly identifies the resource (a login link/token) and differentiates this tool from sibling dashboard/onboarding tools by focusing on the magic-link minting mechanism.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states when to use the tool: before dashboard tasks such as the engine-connect step, billing card updates, or accessing a project page. It also explains that the returned link should be sent to the user and that a fresh call is needed after expiry, providing clear contextual usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

project_statusAInspect

Check how a project is doing: what is in flight, what shipped, what is stuck.

Returns lifecycle status, open task count, queued roadmap item count, runs
today, verification pass rate, open blockers, pause state, and freshness.
Zero `open_tasks` does not mean an empty roadmap: `queued_roadmap_items`
counts planned work awaiting expansion, including work held while the
scheduler is disabled. Use list_roadmap to inspect those items.

`glm_peak_paused` is NOT a fault and sets no pause columns: it is the
ephemeral GLM peak-hours skip. True means clean PRs hold and runs stop
until `glm_peak_resumes_at`. It is an intentional cost gate, so report it
as "waiting for off-peak", never as a failure.

For `web_app` projects the optional `ui_review` block can mint a GitHub
installation token and read the default branch's `ui-review.json`, so this
status call is not guaranteed to be a purely local read.
ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations carry no safety hints (all false), so the description carries the burden. It discloses the non-local nature for web_app projects (token minting and reading ui-review.json), and clarifies that glm_peak_paused is an intentional cost gate, not a fault. This goes beyond the structured fields and is not contradicted by any annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is moderately long but each sentence adds unique value: purpose, return list, two clarifications, and a warning. It is front-loaded with the core purpose and progressively adds nuance. A slight tightening of the field list would improve it, but it remains efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers all non-obvious behaviors: the meaning of zero open_tasks, the glm_peak_paused state, and the web_app ui_review side effect. Since an output schema exists, the return format need not be restated. The description fully equips an agent to call and interpret the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the description never mentions project_id explicitly. It uses 'project' implicitly but gives no guidance on format, source, or how to obtain it. For a single required parameter, the description should at least acknowledge it; it does not, so it fails to compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Check') and resource ('how a project is doing'), then enumerates the returned facets. It clearly distinguishes from siblings like list_projects (which lists projects) and list_roadmap (which inspects roadmap items), so an agent can tell them apart.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete routing instruction: 'Use list_roadmap to inspect those items' when queued_roadmap_items is non-zero. It also explains how to interpret glm_peak_paused (report as 'waiting for off-peak', not failure). It lacks an explicit 'when not to use' for the general case, but the context is sufficient.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

rearm_roadmap_itemAInspect

Run the dashboard's replay-guarded same-item dependency re-arm.

Pass an ``expand_produced_nothing`` or ``roadmap_item_parked`` escalation
id. Keelen verifies delivered structural prerequisites, grants one bounded
expand retry, preserves the roadmap item id and ``depends_on_item_ids``, and
resolves the matching cards in one transaction. A bounded live GitHub check
may be used when the local merge state is stale, so a call can reach out
even when the re-arm itself changes nothing.
ParametersJSON Schema
NameRequiredDescriptionDefault
escalation_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description richly discloses behavior beyond the annotations: it is transactional, preserves key identifiers, grants a bounded retry, and may perform a bounded live GitHub check that reaches out even when the re-arm itself changes nothing. This aligns with openWorldHint and idempotentHint false and provides meaningful operational context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the core purpose, then explains the key behaviors. It uses some internal jargon like 'replay-guarded' and 'delivered structural prerequisites,' but every sentence adds meaningful detail without excessive length.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has a single parameter, an output schema, and annotations, the description is largely complete: it covers the trigger condition, the actions performed, transactional guarantees, preserved fields, and potential external network access. It leaves some details vague, such as what 'delivered structural prerequisites' means and what the matching cards are, but those are not critical for invoking the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema only defines escalation_id as a string with no description. The tool description compensates by stating the accepted escalation id types: expand_produced_nothing or roadmap_item_parked. This gives essential meaning that the schema lacks, though it does not explain how to obtain or format such an id.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific operation: running the dashboard's dependency re-arm, and details what it does: verifies prerequisites, grants one bounded expand retry, preserves the roadmap item id and depends_on_item_ids, and resolves matching cards. It is clear it is a specialized re-arm action, though it does not explicitly contrast itself with sibling tools like resolve_escalation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a clear usage condition: pass an expand_produced_nothing or roadmap_item_parked escalation id. This tells an agent when this tool is appropriate, but it does not explicitly state when not to use it or which alternative to choose for other escalation types.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

refine_requestA
Destructive
Inspect

Say what is wrong with what a request produced, and have it reworked.

Status must be "done". `feedback_md` is 1..2000 chars. Moves the thread to
refine_pending; poll get_request_status to see the revised items. Refinement
re-plans the thread's roadmap items and can EDIT OR HARD-DELETE the task rows
it generated earlier (the machine patch-tasks endpoint deletes them
outright), so treat this as destructive: up to five refinements per thread
each carry that power.
ParametersJSON Schema
NameRequiredDescriptionDefault
thread_idYes
project_idYes
feedback_mdYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark the tool as destructive, but the description goes well beyond that by specifying exactly what gets affected: 'can EDIT OR HARD-DELETE the task rows it generated earlier,' explains the deletion mechanism ('the machine patch-tasks endpoint deletes them outright'), warns 'treat this as destructive,' and adds a limit ('up to five refinements per thread each carry that power'). It also discloses the state transition ('Moves the thread to refine_pending') and the re-planning side effect. This is substantial behavioral disclosure that annotations alone do not provide, and it is consistent with destructiveHint=true.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is about 80 words, front-loads the purpose in the first sentence, and then packs in only high-value information: the status precondition, feedback length, state transition, follow-up polling, side effects (re-planning, editing/deleting tasks), destructive warning, and limit. Every sentence earns its place; the technical aside about the patch-tasks endpoint explains the mechanism without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with an output schema and destructive side effects, the description covers the essential operating context: when it can be used (status done), the input constraint (feedback length), what happens to the thread state, how to observe results (poll get_request_status), the scope of side effects (re-plans roadmap, edits/deletes tasks), and the limit of five refinements. There is no requirement to explain return values because an output schema exists, so the description is complete for correct invocation and expectation-setting.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does add meaning for feedback_md by specifying the character range ('1..2000 chars') and the status precondition, but it does not explain project_id or thread_id beyond their names. These are standard identifiers and likely self-evident, but with zero schema descriptions, a full compensation would require at least a sentence clarifying their role. The description partially fills the gap, earning a 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb-resource pair: 'Say what is wrong with what a request produced, and have it reworked.' This clearly distinguishes refine_request from siblings like submit_request (initial submission) and answer_request (responding), and the later mention of polling get_request_status differentiates its purpose from status-checking tools. The phrase 'have it reworked' plus the detailed mechanics make the purpose unambiguous and distinct.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives an explicit precondition: 'Status must be "done".' It also tells the agent what to do after invocation: 'poll get_request_status to see the revised items,' implicitly naming the alternative for checking status. It does not explicitly say 'use this only when status is done and avoid manual editing tools,' but the condition and outcome guidance are clear enough to route the agent correctly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

reorder_roadmapAInspect

Change what gets built first.

`ordered_ids` is the desired front-to-back order of queued roadmap-item ids
(get them from list_roadmap). The first id becomes the highest priority —
the cadence expands the lowest-priority_int queued item next. Horizon pins
still dominate: a pinned-later item stays at the back and a pinned-now item
at the front, regardless of position in `ordered_ids`. Ids that are unknown
or no longer queued are skipped; duplicates are rejected. Returns {updated,
queue, next_step}.
ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYes
ordered_idsYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses several non-obvious behaviors beyond the annotations: horizon pins override the supplied order, unknown or no-longer-queued ids are silently skipped, duplicates are rejected, and the response shape is provided. This gives the agent a accurate mental model of side effects even though the write-related annotations are minimal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized and front-loaded with the core purpose. Each sentence earns its place by adding behavioral detail. The minor typo 'lowest-priority_int' slightly interrupts readability, but the structure and focus remain strong.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity, the description covers ordering rules, edge cases, and return value shape. The annotations and output schema cover safety and result structure, while the description fills the behavioral gaps an agent needs to call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description thoroughly explains `ordered_ids`, including ordering semantics, priority direction, pin interactions, duplicate handling, and unknown-id behavior. `project_id` is not elaborated, but it is a common identifier parameter and appears in the required schema. For a tool with 0% schema description coverage, this is strong compensation overall.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource ('Change what gets built first') and then explains the effect on queued roadmap items. It clearly distinguishes this tool from siblings like list_roadmap and cancel_roadmap_item by focusing on ordering rather than listing, creating, or removing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly implies when to use the tool: when the agent needs to change build priority order. It also tells the agent where to get the required ids ('get them from list_roadmap'). It does not explicitly name exclusions or alternatives such as rearm_roadmap_item or cancel_roadmap_item, so it stops short of a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

replan_taskA
Destructive
Inspect

Replace an unretryable blocked task without losing its lineage.

Creates a fresh, explicitly planned task under the same roadmap item and attempt lineage, transfers prerequisites and downstream dependents, and cancels the stale task with supersession provenance. The old PR, task, failures, criteria, and evidence remain in the audit history; nothing is marked delivered.

ParametersJSON Schema
NameRequiredDescriptionDefault
titleYes
body_mdYes
task_idYes
project_idYes
decision_mdYes
acceptance_criteriaYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already flag destructiveHint=true, but the description adds substantial context beyond that: lineage preservation, transfer of prerequisites and downstream dependents, cancellation with supersession provenance, and the critical guarantee that old artifacts remain in audit history while "nothing is marked delivered." This materially refines what the destructive hint means and prevents false assumptions about history loss.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with no filler: purpose, mechanism, and preservation guarantee each earn their place. The final sentence is high-value because it clarifies what a destructive operation does NOT do (mark delivered, erase history), which an agent would otherwise have to guess.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex destructive mutation with six required parameters, the description covers the core contract well: trigger condition, artifacts created, relationships transferred, and preserved state. The output schema covers return values, so that omission is acceptable. A brief note on prerequisites (e.g., what happens if the task is not actually blocked) would push it to complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the burden for parameter meaning. It implicitly maps most parameters — task_id is the blocked task being replaced, title/body_md/acceptance_criteria describe the new task, project_id anchors the roadmap item — but does not explicitly tie parameters to the described behavior. decision_md's role is especially left to inference.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource+qualifier: "Replace an unretryable blocked task without losing its lineage." It then details the exact mechanics — creating a fresh planned task, transferring prerequisites and dependents, and cancelling the stale task with supersession provenance — which clearly differentiates it from siblings like retry_blocked_task and close_task.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The qualifier "unretryable blocked task" in the first sentence gives clear context for when to use this tool, and contrasts naturally with the sibling retry_blocked_task for retryable cases. However, it never explicitly names the alternative or states a when-not-to-use condition, leaving the routing slightly implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resolve_escalationA
Destructive
Inspect

Acknowledge a handled ask only when no blocked task becomes invisible.

`decision_md` is a required short note (why/how it was resolved), appended to
the escalation's detail_md as an audit trail. Resolving a project_pause /
orphan_pause RESUMES the project (clears the pause). A blocked task's final
task_block/operator_action cannot be acknowledged: use its typed retry,
platform-policy resolution, replan, close, or supersede operation instead.
The acknowledgement itself is replay-guarded (a repeat changes nothing), but
the call is not idempotent end to end: resuming a paused project lets queued
refinement run, and that refinement can DELETE tasks. Accepts a real
escalation UUID or a synthetic `orphan:<project_id>`.
ParametersJSON Schema
NameRequiredDescriptionDefault
decision_mdYes
escalation_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the annotations: it explains that decision_md is appended as an audit trail, that resolving a project_pause/orphan_pause resumes the project, that the acknowledgement is replay-guarded but the call is not idempotent end to end, and that queued refinement can DELETE tasks. This materially enriches the destructiveHint and idempotentHint flags without contradicting them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence earns its place: the opening states the primary guard, the middle covers parameter semantics and side effects, and the closing defines accepted ID formats. The description is dense but not padded, and critical warnings are front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's destructive, non-idempotent, and conditional behavior, the description covers the key decision criteria, parameter usage, side effects, alternatives, and valid ID forms. It is complete enough for an agent to know when to call it, how to fill the arguments, and what consequences to expect.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry parameter meaning, and it does. It defines decision_md as a required short note explaining why/how it was resolved and where it goes (detail_md). It defines escalation_id by its accepted forms: a real escalation UUID or a synthetic 'orphan:<project_id>'. Both parameters receive meaning the schema alone does not provide.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action ('Acknowledge a handled ask') on a specific resource (an escalation) and states the precise condition under which it applies ('only when no blocked task becomes invisible'). It also differentiates from siblings by explicitly saying a blocked task's final task_block/operator_action must be handled via retry, platform-policy resolution, replan, close, or supersede operations instead.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives an explicit when-to-use gate ('only when no blocked task becomes invisible') and spelling out alternatives by operation type, which map to sibling tools like retry_blocked_task, resolve_platform_policy_conflict, replan_task, and close_task. It also clarifies when not to acknowledge and what side effects to expect for paused projects.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

resolve_platform_policy_conflictAInspect

Resolve a platform-policy dead-end with a structured plan decision.

``decision`` is one of ``reuse_existing_evidence``, ``split_task``,
``raise_project_cap``, or ``remove_scenario``. ``raise_project_cap`` also
requires ``project_cap`` (6..32). The decision becomes a new task-plan
section; the same task lineage is requeued before the escalation resolves.

On a project that runs the ui-review evidence lifecycle (``ui-review.json``
is a retained evidence catalog), ``raise_project_cap`` changes the
per-run capture budget -- how many scenarios one capture run holds -- not
catalog capacity: the catalog has no scenario limit, so no decision is
needed to make room in it.
ParametersJSON Schema
NameRequiredDescriptionDefault
decisionYes
decision_mdYes
project_capNo
escalation_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations mark this as a non-readOnly, non-idempotent, open-world mutation but disclose nothing about side effects. The description usefully adds that the decision becomes a new task-plan section and that the same task lineage is requeued before the escalation resolves, plus the capture-budget vs catalog-capacity distinction for raise_project_cap. It stops short of stating permission requirements or reversibility.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded in the first line, followed by decision semantics and the ui-review caveat in a logical order with no filler. The final paragraph is dense and specialized but carries genuine disambiguation value rather than redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be explained, and the description covers the decision space and the ui-review lifecycle caveat well. However, against a 4-parameter mutation tool with 0% schema coverage, the unexplained escalation_id and decision_md leave an agent guessing at two required inputs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does document the decision enum values (absent from the schema) and the project_cap range 6..32 with its conditional requirement, but escalation_id and decision_md receive no explanation, leaving two required parameters undocumented in both places.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: resolve a platform-policy dead-end with a structured plan decision, and it enumerates the allowed decision values. It is clearly distinguished from a generic escalation resolver by the 'platform-policy dead-end' framing, though it never names the sibling resolve_escalation to draw the boundary explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implies the triggering scenario (a platform-policy dead-end) and enumerates the valid decisions, but gives no explicit when-to-use vs when-to-use-an-alternative routing against siblings like resolve_escalation or replan_task. The one 'when not' nuance ('no decision is needed to make room in the catalog') is about semantics of a specific decision value rather than tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

retry_blocked_taskAInspect

Grant one fresh attempt after fixing a recovery-budget task block.

Pass the task_block id returned by ``list_escalations``. Eligibility,
tenant, task, failure class, task status, and competing blockers are all
derived server-side. This is distinct from ``resolve_escalation``, which
remains acknowledgement-only for task blocks. Business-level replay guard:
an already-retried block is not retried again; the call is not idempotent end
to end, because an authenticated request also persists credential-use state.
ParametersJSON Schema
NameRequiredDescriptionDefault
decision_mdYes
escalation_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide idempotentHint=false and readOnlyHint=false, but the description adds valuable behavioral context: the call persists credential-use state, making it non-idempotent end-to-end, and includes a business-level replay guard (an already-retried block is not retried again). It does not contradict annotations. The additional detail about server-side derivation of eligibility, tenant, etc., also enhances transparency, though it doesn't explicitly mention destructive or safety implications beyond what annotations imply.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but efficient; it front-loads the primary purpose, then provides usage context, differentiation, and important behavioral caveats. Every sentence carries information, and it avoids redundancy. Slightly long, but each clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the purpose, usage, distinction, and non-idempotency, and there is an output schema so return details are not required. However, it does not explain the role of decision_md, and the parameter name mismatch (task_block id vs escalation_id) is a potential source of confusion. It is mostly complete but leaves one parameter and naming consistency unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains that 'task_block id' should be passed (mapping to escalation_id), but it gives no information about the second required parameter, decision_md. It also uses a different name ('task_block id' vs 'escalation_id') which could cause confusion. The description adds minimal value for the parameters beyond what the schema already lists.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Grant one fresh attempt') and resource ('recovery-budget task block'), and explicitly distinguishes itself from 'resolve_escalation', making it clear what it does and how it differs from a sibling. The purpose is unambiguous and action-oriented.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly tells the user when to use it: after fixing a recovery-budget task block. It also names the alternative (resolve_escalation) and clarifies that the alternative is acknowledgement-only, giving a clear routing decision.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

rollback_roblox_placeA
Destructive
Inspect

Roll a Roblox project's place back to a previously-published version.

For a `roblox_game` project, re-publishes the RETAINED build artifact for
`version_number` (get published versions from the web Roblox Publishing card)
— it never rebuilds from source, so rollback is fast + deterministic. This
mints a NEW Roblox version pointing at the old build. Unknown version → 404;
a version with no retained artifact → 422; a place open in Studio /
rate-limited → 409 (retry); an invalid or unscoped Open Cloud key → 409 /
403. Returns {version_number, env, published_at, status, next_step}.
ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYes
version_numberYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description thoroughly discloses behavioral traits: it 'never rebuilds from source', 'mints a NEW Roblox version pointing at the old build', and lists specific error codes (404, 422, 409) with meanings. Annotations indicate destructiveHint=true and idempotentHint=false, which the description aligns with. It goes above annotations by detailing rate limits, invalid key handling, and the return payload. This is excellent transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-organized: a lead sentence stating purpose, then a clause on how it works, then error scenario details, then the return shape. It is concise yet packed with necessary details, with no filler or fluff. Each sentence serves a distinct role, and the error code list is efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool complexity (destructive operation, external platform interaction, error handling), the description is remarkably complete. It covers usage constraints, failure modes, retry guidance, and the output schema (though also present separately). With annotations and output schema, the description provides all necessary context for an agent to invoke it safely and correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 2 parameters, and schema description coverage is 0%. The description only mentions 'version_number' implicitly (it references getting published versions) but does not explain the meaning or constraints of 'project_id' or the exact syntax of 'version_number' (e.g., that it's an integer existing in the list). Since the schema itself provides only names and types, the description fails to add sufficient semantic elaboration to make parameter usage clear.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's action: rolling back a Roblox project's place to a previously-published version. It specifies the resource ('place'), the operation ('roll back'), and adds context (for 'roblox_game' projects) and distinguishes it from rebuilding from source. It clearly differentiates from sibling tools like update or publish by focusing on rolling back to a retained artifact.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides clear usage context: it's for roblox_game projects, and tells the agent to get published versions from the web Roblox Publishing card. It also infers when to use it (when rolling back is needed) but does not explicitly mention alternatives or exclusions (e.g., 'use this only when you need to revert to a published version, not for normal updates'). However, given the limited sibling tool set, this is adequate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_control_gap_reviewAInspect

Start a source-bound security control gap review for one project.

The review records bounded engineering observations under the user's
selected Cyber Essentials or CMMC Level 1 or Level 2 context. It does not determine framework
standing, and it does not make changes. The `disclaimer_md` field must be
repeated to the user before any observation is summarised.

The project needs a saved framework profile, an eligible plan, and a place
on the operator allowlist. Billable; one review is allowed in flight per
project. Refusals return `ok: false` with a next step and do not start work.
Rate-limited per workspace.
ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations carry only negative hints (readOnlyHint: false, destructiveHint: false), so the description shoulders the behavioral burden — and it delivers richly: it states the review 'does not make changes,' is billable, rate-limited per workspace, limited to one in-flight review per project, requires the disclaimer_md to be repeated, and that refusals return ok: false without starting work. None of this contradicts the annotations; it substantially exceeds them.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact yet information-dense; every sentence earns its place — purpose, non-effect, disclaimer obligation, prerequisites, billing, concurrency, refusal behavior, and rate limiting. It is front-loaded with the core purpose before the constraints, and no sentence is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Because the tool has an output schema, the description need not document return values, and it correctly omits them. Everything an agent needs to invoke it correctly — purpose, scope, prerequisites, side effects, limits, and the mandatory disclaimer step — is present. Nothing material is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and there is one parameter, project_id, whose semantics are essentially self-evident. The description compensates by defining the project eligibility prerequisites (saved framework profile, eligible plan, allowlist) and the source-bound, single-project scope, which tells the agent what kind of project_id is expected. This goes beyond a bare schema but stops short of documenting the field format or type nuances.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb-resource-scope construction: 'Start a source-bound security control gap review for one project.' It disambiguates from the sibling tools run_security_review and run_legal_exposure_review by pinning the subject to 'control gap' and naming the concrete framework contexts (Cyber Essentials, CMMC Level 1/2). An agent can tell exactly what this tool does and what it is not.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description lays out the precondition set explicitly — a saved framework profile, an eligible plan, and operator allowlist placement — so an agent knows when the call is valid. It does not name the sibling alternatives it competes with or give explicit when-not-to-use guidance, but the prerequisite list plus the framework context effectively constrains the intended use case.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_security_reviewAInspect

Find security problems in a repository: a deep, whole-codebase review.

Spawns a one-shot audit that scans the repo across a kind-aware taxonomy
(secrets + git history, vulnerable/abandoned deps, injection, SSRF, path
traversal, deserialization, crypto, info-leak, plus web authz/session/CORS,
library API-misuse, game client-trust, or infra/CI as applicable) and posts
findings to the project's Security review for human triage. You review the
findings, then send the ones worth fixing into the loop as Requests; no fix
is applied automatically. Billable; one audit in-flight per project.
(Triggering is disabled while the feature is hardened for production: a
project not on the operator allowlist — empty by default — returns a message
instead of spawning; earlier results stay visible.) Requires a paid plan;
a free or trial workspace gets a message telling the user to upgrade.
ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations, the description discloses substantial behavioral detail: findings are posted for human triage, no fixes are applied automatically, triggering is disabled on non-allowlisted projects, earlier results remain visible, and free/trial workspaces receive an upgrade message. This gives the agent an accurate picture of side effects and failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with a clear summary, followed by relevant operational details. The long taxonomy list adds specificity but makes the text dense; still, every sentence contributes useful information and the structure is logical.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity, the description is unusually complete: it covers prerequisites (paid plan), concurrency limits, disabled states, billing, human triage flow, and the fact that no fix is automatic. An output schema exists, so return-value documentation is not required from the description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has one required parameter, project_id, with 0% description coverage. The tool description implies the parameter identifies the project/repository to audit and mentions per-project constraints, but it does not explicitly describe the format, provenance, or how to resolve a project_id. It partially compensates but leaves some ambiguity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Find security problems in a repository: a deep, whole-codebase review.' This clearly distinguishes it from sibling review tools like run_control_gap_review and run_legal_exposure_review by naming the security scope and taxonomy.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides strong contextual guidance: it is one-shot, billable, limited to one in-flight audit per project, requires a paid plan, and may be disabled by an operator allowlist. It does not explicitly name alternative tools or state 'use X instead,' but the scope and constraints are clearly communicated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_product_goalA
Destructive
Inspect

Set the outcome to aim at right now, so work is prioritised toward one thing.

`goal_md` is free-form markdown and must be non-empty. Tenant-scoped: a
project not in the caller's workspace 404s. Returns {project_id,
product_goal_md, updated_at, next_step}.
ParametersJSON Schema
NameRequiredDescriptionDefault
goal_mdYes
project_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With destructiveHint=true already carrying the mutation/destructive signal, the description adds useful behavior beyond annotations: tenant-scoped access control, a 404 for out-of-workspace projects, non-empty enforcement for goal_md, and the exact return shape. It does not explicitly state that an existing goal is overwritten, but the destructive annotation covers that risk.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences: one for purpose, one for parameter/scoping constraints, one for the return contract. Information is front-loaded and no sentence is redundant with the annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter mutation with output schema and annotations, the description covers purpose, parameter constraints, authentication scope (tenant), error behavior (404), and return values. Nothing critical is missing for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It explains goal_md is free-form markdown and must be non-empty, and the tenant-scoped note clarifies that project_id must reference a project in the caller's workspace. This is useful, though project_id's type/format is left implicit.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb-resource pair: 'Set the outcome to aim at right now,' which clearly identifies the product-goal resource and distinguishes it from sibling read tool get_product_goal and long-term set_product_vision. It also explains the intended effect ('work is prioritised toward one thing'), making the tool's purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear context: this is for setting the current outcome/focus, and it warns that goal_md must be non-empty and that tenant-scoped projects outside the caller's workspace will 404. It does not explicitly name alternatives or say when not to use it, but the 'right now' framing plus sibling names makes the selection context fairly clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_product_visionA
Destructive
Inspect

Say what this product is for, so every run knows what it is building toward.

`vision_md` is free-form markdown and must be non-empty. Tenant-scoped: a
project not in the caller's workspace 404s. Returns {project_id,
product_vision_md, updated_at, next_step}.
ParametersJSON Schema
NameRequiredDescriptionDefault
vision_mdYes
project_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds meaningful context beyond the annotations: vision_md must be non-empty, the call is tenant-scoped, and out-of-workspace projects return 404. It also discloses the return object. It does not contradict the destructiveHint flag, though it could have explicitly mentioned that the existing vision is overwritten.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact, front-loaded with the core purpose, and each sentence adds necessary detail: the vision statement, the markdown constraint, tenant behavior, and return fields. There is no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-parameter mutation tool with annotations already present, this description is nearly complete: it defines the core purpose, input constraints, tenant scoping, and return shape. It lacks only explicit guidance about overwriting behavior or when to use it relative to siblings, but those are minor gaps given the tool's simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With schema description coverage at 0%, the description carries the burden of explaining parameters. It clarifies that vision_md is free-form markdown and must be non-empty, which adds meaningful guidancechers. However, project_id is only indirectly described through the tenant-scoping noteفق and lacks direct semantic detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly explains what the tool does: it sets the product vision by stating what the product is for and what it builds toward. It uses a specific verb and resource, but it does not explicitly distinguish itself from closely related siblings like set_product_goal or get_product_vision.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no explicit guidance on when to use this tool instead of alternatives. It mentions a tenant-scoping constraint and 404 behavior, but not when setting vision is appropriate or how it differs from setting the product goal.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

set_ui_review_scenario_capAInspect

Set this web_app project's ui-review scenario budget, or reset it.

``scenario_cap`` is an integer from 6 through 32, or ``null`` to reset to
the platform default of 12. The screenshot cap derives from it and moves
with it, so the two can never starve each other.

Raising is always allowed, including from a completely full manifest.
LOWERING is refused when the default branch already declares more scenarios
or screenshots than the smaller budget allows, and is also refused when that
manifest cannot be read — both caps are enforced when keelen pushes and not
in your CI, so an over-cap manifest fails every push while CI stays green.

On a project with the retained evidence catalog (the response carries
``evidence_lifecycle: true``) the same parameter is a CAPTURE-RUN setting:
it limits how many scenarios and screenshots one capture run holds, a pull
request that needs more is split into capture batches, and the catalog
keeps every scenario whatever the value. The response then also reports
``catalog`` (the retained counts) separately from ``run_limits``, and a
lower is not refused for catalog size once the catalog is migrated.

Read ``project_status.ui_review`` first to see the live occupancy. Repeated
calls with the same value change nothing.
ParametersJSON Schema
NameRequiredDescriptionDefault
project_idYes
scenario_capNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only give a generic safety profile (readOnly false, destructive false, openWorld false), and the description goes well beyond it: it explains refusal semantics, the coupling of screenshot cap to scenario cap, and the mode switch on the retained-evidence catalog (capture-run cap vs catalog sizing, with catalog and run_limits reported separately). The only friction is 'repeated calls with the same value change nothing', which sits in tension with idempotentHint=false; read as a conservative default hint rather than a misstatement, the disclosure is still unusually rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and parameter range are front-loaded, and the paragraphs are ordered purpose → parameter → refusal rules → mode caveat. It is longer than necessary for a 2-parameter tool, and the stray double-backtick line endings plus the unexplained 'keelen pushes' reference add noise that a sentence or two of trimming would remove.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with an output schema, the description covers what an agent needs before calling: prerequisite read, valid range, refusal conditions, mode-dependent semantics, and a pointer to the response fields (evidence_lifecycle, catalog, run_limits) it should interpret. Return-value detail is correctly left to the output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% (no per-parameter descriptions), so the description has to carry the load and does: scenario_cap is an integer 6–32 or null for the default of 12, it derives the screenshot cap, and its meaning changes on evidence_lifecycle projects. project_id is never described in either place, which is the only gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states a specific verb and resource: setting (or resetting) this web_app project's ui-review scenario budget. Nothing in the sibling list covers the same resource, so an agent can route to it unambiguously from the purpose alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the prerequisite ('Read project_status.ui_review first to see the live occupancy'), the reset condition ('null to reset to the platform default of 12'), when raising is allowed, and the exact conditions under which lowering is refused (over-cap default branch, or unreadable manifest). That is explicit when-to-use and when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

signupA
Destructive
Inspect

Create a Keelen account (or start agent login) — emails a 6-digit code.

UNAUTHENTICATED — the only tool besides verify_email that works before a
bearer key is configured. `email` is where the code is sent. Flow:
signup(email) -> the user reads the 6-digit code from their inbox ->
verify_email(email, code) returns a reveal-once API key -> save it as this
server's `Authorization: Bearer <api_key>` header in your MCP client config
-> reconnect -> get_onboarding_status() to continue setup. The code expires
in 15 minutes; call signup again to resend. Response is uniform whether or
not the email already has an account (enumeration-safe), so signup doubles
as agent LOGIN. Rate-limited per IP and per email.

ASK THE USER for `email` in chat and WAIT for their answer before calling
this. Do NOT infer it from your client profile, the logged-in account, git
config, or any other ambient source; if you already hold a candidate, echo
it back and get an explicit yes first. Because this call doubles as LOGIN, a
guessed address signs the user in to whatever workspace owns it, and the
rest of setup then mints an API key on, and creates a project in, an account
they did not choose.
ParametersJSON Schema
NameRequiredDescriptionDefault
emailYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already indicate destructiveHint=true and readOnlyHint=false, but the description adds substantial context beyond these flags: it emails a code, expires in 15 minutes, is enumeration-safe, rate-limited per IP/email, and doubles as login with potential to sign the user into an unintended account. It also discloses that calling signup again resends the code. This enriches the agent's understanding of side effects and security behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence earns its place. The core action and flow are front-loaded, followed by expiration, enumeration-safety, rate-limiting, and a crucial user-consent admonition. It is structured with a clear sequence and prioritized warnings; no filler exists. This is a model of efficient technical writing for an AI agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's sensitivity (auth and account creation) and the absence of rich schema descriptions, the description covers all essential aspects: prerequisites (unauthenticated), flow, expected inputs, security implications, rate limits, and behavior. The output schema is provided separately, so return-value details are not required. An agent has everything needed to call it correctly and safely.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With schema description coverage at 0% and only one parameter (email), the description must add meaning. It clarifies that 'email' is where the 6-digit code is sent and is the address used for account/login. It also frames email as security-sensitive, instructing the agent to confirm it with the user. This goes beyond the bare schema definition, though it does not specify email format validation or edge cases, so a 4 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action: 'Create a Keelen account (or start agent login) — emails a 6-digit code.' It also clarifies the tool's dual role as signup and login, and distinguishes it from verify_email and other post-auth tools by noting it works 'before a bearer key is configured'. This is a precise, unambiguous purpose that an agent can act on.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit when-to-use guidance: 'UNAUTHENTICATED — the only tool besides verify_email that works before a bearer key is configured.' It also gives a step-by-step flow (signup -> verify_email -> save key -> reconnect) and a critical instruction: 'ASK THE USER for email in chat and WAIT for their answer.' It warns against inferring the email from ambient sources and explains the consequences of guessing. This leaves no ambiguity about correct usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

submit_requestAInspect

Ask for a change in plain English: a feature, a bug fix, or a new direction.

Submit ONE feature or intent per call; split a multi-feature ask into
separate requests. `text` must be under 16000 characters. Returns
{thread_id, status, next_action, poll_after_seconds, next_step}; follow
next_step (re-check get_request_status after poll_after_seconds).
ParametersJSON Schema
NameRequiredDescriptionDefault
textYes
project_idYes
workflow_profileNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (not read-only, not idempotent, not destructive), the description discloses asynchronous behavior: it returns a thread_id, status, next_action, poll_after_seconds, and next_step, and instructs the agent to poll get_request_status. This adds meaningful operational context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: purpose, cardinality constraint, text limit, and return behavior are each stated in a few sentences with no filler. The use of backticks for parameter names and a clear next-step instruction keeps it scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a submission tool with a provided output schema, the description covers the key usage loop, polling behavior, and limits, making it actionable. The main gap is the unexplained workflow_profile parameter, which keeps it from being fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage, the description compensates for the text parameter by defining its format, scope, and length limit, but it does not explain project_id or workflow_profile. The required project_id is only self-evident from its name, and workflow_profile is left undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Ask for a change') and the resource/domain (feature, bug fix, new direction) with clear intent. It does not explicitly differentiate from sibling refine_request or answer_request, but the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides concrete usage constraints: one feature/intent per call, splitting multi-feature asks, a 16000-character text limit, and explicit follow-up steps via get_request_status. It lacks explicit 'do not use when...' exclusions, but the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

verify_emailA
Destructive
Inspect

Redeem the emailed 6-digit code for a reveal-once workspace API key.

UNAUTHENTICATED. `email` + `code` must match a code issued by signup(email)
within the last 15 minutes (5 attempts max). The returned `api_key` is shown
exactly ONCE — store it ONLY in the MCP client config
("Authorization: Bearer <api_key>"), NEVER in a repo or a file you might
commit. Then reconnect this server with the header set and call
get_onboarding_status(). An invalid/expired/consumed code returns a uniform
error — call signup(email) for a fresh one.
ParametersJSON Schema
NameRequiredDescriptionDefault
codeYes
emailYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond annotations, the description reveals critical behaviors: the api_key is shown exactly once, it must only be stored in MCP client config, invalid/expired/consumed codes return a uniform error, and the operation consumes the code. This exceeds what readOnlyHint/destructiveHint alone convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence earns its place: purpose, constraints, security handling, next steps, and failure behavior. The essential one-time key warning is front-loaded and clearly emphasized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's authentication/consumption complexity, the description covers trigger, constraints, error handling, and post-redeem workflow. The output schema exists, so return-value details are not needed in the description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description carries the burden. It fully explains 'email' must correspond to a signup(email) address and 'code' must be the 6-digit code issued within 15 minutes, with attempt limits. Both required parameters are semantically covered.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Redeem the emailed 6-digit code for a reveal-once workspace API key.' This clearly states what the tool does and differentiates it from signup and get_onboarding_status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit usage context: unauthenticated, code must match signup(email), 15-minute validity, 5 attempts max, and instructs to call signup(email) for a fresh code on failure. It also tells the agent exactly what to do after redeem, and names get_onboarding_status as the next step.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool update
    • Changedsubmit_request1 field changed
      • addedInput schema / properties / workflow_profile
        Added value: +{
        +  "anyOf": [
        +    {
        +      "type": "string"
        +    },
        +    {
        +      "type": "null"
        +    }
        +  ],
        +  "default": null,
        +  "title": "Workflow Profile"
        +}
  2. 1 tool update
    • Addedset_ui_review_scenario_cap
  3. 3 tool updates
    • Addedrearm_roadmap_item
    • Addedreplan_task
    • Addedresolve_platform_policy_conflict
  4. 3 tool updates
    • Addedclose_task
    • Addedget_product_goal
    • Addedget_product_vision
  5. 1 tool update
    • Addedretry_blocked_task
  6. 33 tool updates
    • First observedanswer_request
    • First observedarchive_project
    • First observedcancel_roadmap_item
    • First observedclear_horizon_pin
    • First observedconnect_github
    • First observedcontrol_scheduler
    • First observedcreate_project
    • First observeddelete_project
    • First observedget_billing
    • First observedget_control_gap_findings
    • First observedget_legal_exposure_findings
    • First observedget_onboarding_status
    • First observedget_provisioning_status
    • First observedget_request_status
    • First observedimport_project
    • First observedlist_escalations
    • First observedlist_github_repos
    • First observedlist_projects
    • First observedlist_roadmap
    • First observedopen_dashboard
    • First observedproject_status
    • First observedrefine_request
    • First observedreorder_roadmap
    • First observedresolve_escalation
    • First observedrollback_roblox_place
    • First observedrun_control_gap_review
    • First observedrun_legal_exposure_review
    • First observedrun_security_review
    • First observedset_product_goal
    • First observedset_product_vision
    • First observedsignup
    • First observedsubmit_request
    • First observedverify_email

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    Automates DevOps workflows like vulnerability resolution, code review, test generation, and DORA metrics through Claude Code slash commands, using a state machine for reliable execution.
    14 npm
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Autonomous coding pipeline exposed as an MCP server: plan, dispatch, review, and merge software stories through worktree-isolated agents. A frontier model (Claude) handles judgment — planning, review, risk adjudication — while a local model does the implementation, gated by TDD and a merge-time test rerun on the rebased branch.
    25
    5
    Apache 2.0
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.