Skip to main content
Glama

TestChimp

Server Details

QA platform for agents: coverage signals, in-repo test plans, verified tests and release governance.

Ownership verified
Status
Healthy
Last Tested
Transport
Streamable HTTP · MCP 2025-11-25
URL
Repository
TestChimp/testchimp-mcp-client
GitHub Stars
0

TDQS

B3.2/5.0

Scored across 91 tools

Disambiguation3/5

Many tools cluster around overlapping areas such as releases (get-release vs get-release-details), execution history (get-execution-history, get-suite-execution-stats, fetch-execution-report, get-workflow-execution, get-last-run-workflow-detail), and distinct-marking (mark-entity-distinct vs mark-semantic-tests-distinct). The detailed descriptions usually clarify intended use, but the sheer breadth makes selection non-trivial without careful reading.

Naming Consistency5/5

All tool names follow a consistent kebab-case verb-first pattern (create-, get-, list-, update-, upsert-, report-, mark-, etc.) with descriptive object nouns. Minor acronym casing (truecoverage, eaas) does not break the overall convention.

Tool Count1/5

91 tools is an extreme overshoot for a single MCP server, far above the recommended 3-15 range. The surface spans many distinct domains (releases, executions, perf, security scans, meetings, TrueCoverage, BunnyShell, bot management, semantic analysis), creating enormous cognitive load for an agent.

Completeness4/5

The surface covers a very broad QA lifecycle: stories, scenarios, issues, executions, coverage, performance, security scans, meetings, environments, policies, and bot events. Minor gaps exist (little direct deletion/editability beyond status updates for some entities), but core agent workflows appear well supported.

Available Tools

91 tools
ack-bot-eventsAInspect

Acknowledge QA-bot webhook events by eventId (1-100 per call) so TestChimp stops redelivering them. Ack every event you handled or decided to ignore (including expired / irrelevant ones); re-acking is safe. Pass the delivery's ackUrl when present (defaults to the ingress /bot/events/ack). Returns per-id status: BOT_ACK_ACCEPTED / BOT_ACK_ACKED / BOT_ACK_ALREADY_ACKED / BOT_ACK_EXPIRED_RECORDED are fine; BOT_ACK_UNKNOWN_EVENT / BOT_ACK_NOT_A_TARGET / BOT_ACK_MISSING_BOT_ID indicate a wrong id or bot identity.

ParametersJSON Schema
NameRequiredDescriptionDefault
ackUrlNo
eventIdsYes

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: idempotency ('re-acking is safe'), batch limits (1-100 per call), default endpoint behavior, the redelivery-thwarting effect, and the meaning of each status code including which indicate a wrong id or bot identity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the action and effect, then usage rule, then parameter guidance, then the status-code decoding. Dense but every clause is actionable; the status list substitutes for the absent output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-param, no-annotation, no-output-schema tool, the description supplies everything needed: purpose, when to call, batch bound, URL handling, and return-code interpretation. Nothing an agent needs to invoke or react correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it does: eventIds is scoped (1-100 per call), and ackUrl is explained as the delivery's ack URL with a stated default ingress path. Both parameters gain meaning not present in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (acknowledge) and resource (QA-bot webhook events) with the operational consequence (stops TestChimp from redelivering). It is unmistakably the write-side counterpart to the get-bot-*/register-bot-profile siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs to ack every event handled OR deliberately ignored, including expired/irrelevant ones, and notes that re-acking is safe. It also names the fallback condition for ackUrl ('pass when present'). No ambiguity about when to call it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compare-perf-to-baselineBInspect

Compare a run (runId) or filtered target (testchimpId plus optional dimensions) to its promoted baseline. envClass is required. Optional thresholds override max p95 regression percent and maximum fail-rate increase.

ParametersJSON Schema
NameRequiredDescriptionDefault
runIdNo
datasetNo
llmModeNo
profileNo
envClassYes
environmentNo
testchimpIdNo
maxFailRateIncreaseNo
maxP95RegressionPercentNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose that envClass is mandatory and that the optional thresholds override default regression gates, which is real behavioral context. However, it says nothing about permissions, side effects, or what the comparison returns (pass/fail vs. metric deltas), leaving the safety and outcome profile thin.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences with the two selection modes and the required parameter front-loaded; nothing is padding. Slightly compressed phrasing ('max p95 regression percent') costs a little readability but not much space.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter tool with no annotations, no output schema, and no schema descriptions, the description covers the core comparison semantics but omits four parameters and the result shape. An agent can call it, but cannot confidently interpret the outcome or know which optional params matter.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 9 parameters, so the description must compensate. It explains runId, testchimpId, envClass, maxP95RegressionPercent, and maxFailRateIncrease, but leaves dataset, llmMode, profile, and environment entirely undefined, and never states that runId and testchimpId are mutually exclusive alternatives.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (compare) and resource (a run or filtered target vs. its promoted baseline), and names the key identifiers (runId, testchimpId). It is distinguishable from siblings like get-perf-run or promote-perf-baseline, though it never explicitly contrasts itself with them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implies two invocation modes (runId, or testchimpId plus optional dimensions) and flags envClass as required, which is useful context. But it never says when to prefer this tool over get-perf-run, list-perf-baselines, or fetch-execution-report, and gives no exclusions or prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create-issueAInspect

Create a TestChimp issue in the current project. title is required. Use simple fields for common creates, or pass the full curated contract via --json-input (description, issueType, category, severity, status, reportedReleaseId, dueDateMillis, assignee, linkTargets, labels, source, environment, attachments, artifactReference). For /testchimp implement TASK_ISSUE creates: set labels=["TestChimp Implement"] (not source), severity from task priority, category (e.g. FUNCTIONAL), and linkTargets for STORY and/or SCENARIO ordinals. Optional agentTraceability (or flat workflowId/workflowExecutionId/policyFile/…) records CREATED Activity inline — prefer this over a separate report-agent-action for issue creates. For Activity/timeline attachment both workflowId and workflowExecutionId (stable Plan ULID) are required; do not omit workflowExecutionId or mint a new ULID per issue. Authenticated via project API key; project is resolved from the key.

ParametersJSON Schema
NameRequiredDescriptionDefault
titleYes
gitShaNo
labelsNo
sourceNo
statusNo
userIdNo
assigneeNo
categoryNo
severityNo
actorTypeNo
issueTypeNo
agentModelNo
branchNameNo
cliVersionNo
policyFileNo
workflowIdNo
attachmentsNo
descriptionNo
environmentNo
linkTargetsNo
skillVersionNo
dueDateMillisNo
policyVersionNo
agentTraceabilityNo
artifactReferenceNo
reportedReleaseIdNo
workflowExecutionIdNo

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose meaningful traits: authenticated via project API key with project resolved from the key, and the hard constraint that Activity/timeline attachment requires both workflowId and a stable workflowExecutionId (don't mint a new ULID per issue). It still omits idempotency, error behavior, and what a create returns.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose is front-loaded in the first sentence, followed by progressively more specific guidance, and for a 27-param tool with zero schema coverage the length is largely justified. It is dense single-paragraph prose with some run-on content that could be better structured, but few sentences are pure filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given a high-complexity create tool with nested objects, no annotations, and no output schema, the description supplies enough to call it correctly for the main scenarios (common create, TASK_ISSUE, traceability). Remaining gaps are the undocumented secondary traceability params and lack of any return-value or failure expectations.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 27 params, so the description must compensate and largely does: it enumerates the curated contract fields and adds real semantics for labels ("TestChimp Implement"), severity mapping from task priority, category, linkTargets ordinals, and agentTraceability. It leaves several traceability fields (gitSha, actorType, branchName, cliVersion, policyVersion) without meaning, so it is strong but not complete.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ("Create a TestChimp issue in the current project") and names the primary scenario. It also distinguishes itself from siblings by explicitly preferring this over report-agent-action for issue creates, so an agent can route correctly without opening another schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context: simple fields for common creates vs. --json-input for the full contract, plus the specific /testchimp TASK_ISSUE recipe. It names an alternative (report-agent-action) and when to prefer this tool, but never states an explicit when-not-to-use or failure condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create-test-scenarioAInspect

Create a test scenario linked to a user story and allocate a real TS-. Response includes content: canonical stub markdown already containing id: TS- and story: US-. BLOCKING workflow: call this FIRST → Write the returned content to the repo plans/scenarios path (edit body as needed but keep id: and story:) → call update-test-scenario with the full markdown. Never write scenario markdown that omits id. update-test-scenario rejects missing id/story with a clear error. platformFilePath must be under plans/scenarios/ and end with .md. userStoryOrdinalId is the numeric part of the parent US- id. Optional agentTraceability records Activity inline (no separate report-agent-action for this create).

ParametersJSON Schema
NameRequiredDescriptionDefault
titleYes
gitShaNo
userIdNo
actorTypeNo
agentModelNo
branchNameNo
cliVersionNo
policyFileNo
workflowIdNo
skillVersionNo
policyVersionNo
platformFilePathYes
agentTraceabilityNo
userStoryOrdinalIdYes
workflowExecutionIdNo

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden and does well: it discloses the BLOCKING call order, that update-test-scenario rejects missing id/story with a clear error, and that agentTraceability records Activity inline (superseding report-agent-action here). It omits auth/permission needs and duplicate/idempotency behavior, but for a create operation this is strong disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose in the first sentence, then lays out the workflow with arrow notation that is easy to follow. It is dense but each clause adds a constraint or behavior; only the parenthetical about update-test-scenario's error slightly repeats the preceding instruction.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 15-parameter, nested-object, no-output-schema create tool with no annotations, the description covers the critical path (required inputs, path constraints, call sequence, returned content). It stops short of explaining the bulk of optional metadata params, which is the main remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it adds real meaning for three parameters: platformFilePath must be under plans/scenarios/ and end in .md, userStoryOrdinalId is the numeric part of the parent US-<n> id, and agentTraceability records Activity inline. However, the other ~12 metadata parameters (gitSha, userId, actorType, etc.) receive no explanation anywhere, so the coverage gap is only partially filled.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Create a test scenario') plus the side effect of allocating a real TS-<ordinalId>, and links it to a parent user story. An agent can immediately distinguish it from siblings like update-test-scenario and get-test-scenarios.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit ordered workflow: call this FIRST, write the returned content to the repo path, then call update-test-scenario with the full markdown. It names the alternative tool and the exact condition that triggers it, leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

create-user-storyAInspect

Create a user story on the TestChimp project and allocate a real US-. Response includes content: canonical stub markdown already containing id: US-. BLOCKING workflow: call this FIRST → Write the returned content to the repo plans/stories path (edit body as needed but keep id:) → call update-user-story with the full markdown. Never write story markdown that omits id. platformFilePath must be under plans/stories/ and end with .md. Optional agentTraceability (or flat workflowId/workflowExecutionId/policyFile/policyVersion/gitSha/…) records AGENT_WORKFLOW_ACTIVITY inline — no separate report-agent-action needed for this create.

ParametersJSON Schema
NameRequiredDescriptionDefault
titleYes
gitShaNo
userIdNo
actorTypeNo
agentModelNo
branchNameNo
cliVersionNo
policyFileNo
workflowIdNo
skillVersionNo
policyVersionNo
platformFilePathYes
agentTraceabilityNo
workflowExecutionIdNo

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does substantial work: it discloses the return content (canonical stub markdown with id), the mandatory write-then-update sequence, the inline AGENT_WORKFLOW_ACTIVITY recording side effect, and a hard validation rule ('Never write story markdown that omits id'). It stops short of stating permission requirements, idempotency, or failure behavior, so not a full 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose then the workflow and constraints; every sentence carries information. The compressed parenthetical listing flat traceability fields is dense but functional, not filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description covers the return content, the required call sequence, path validation, and traceability recording. For a 14-param mutation tool this is close to complete, though most individual params remain unexplained.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% across 14 params, so the description must compensate. It adds real meaning for platformFilePath (must be under plans/stories/ and end with .md) and explains the nested agentTraceability vs flat field form, but the remaining ~11 traceability params (gitSha, actorType, agentModel, branchName, etc.) get no explanation. Partial compensation only.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Create a user story') and adds the distinctive behavior of allocating a real US-<ordinalId>. It explicitly distinguishes itself from update-user-story and report-agent-action, so an agent can route correctly without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use it ('call this FIRST') within a described blocking workflow (create → write to repo → update-user-story), and states when an alternative is not needed ('no separate report-agent-action needed for this create'). Nothing about ordering is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

destroy-ephemeral-environmentCInspect

Delete a BunnyShell environment created for this project.

ParametersJSON Schema
NameRequiredDescriptionDefault
bnsEnvironmentIdYes

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It signals a destructive delete but says nothing about irreversibility, whether the environment data is recoverable, what permissions are required, or whether dependent resources are removed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no wasted words, front-loaded with the action verb. It is tight but too sparse to be exemplary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a destructive, annotation-free tool with no output schema, the description is too thin: it omits irreversibility, permission requirements, and any notion of what the caller receives or should do afterward.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single parameter bnsEnvironmentId has only type/minLength. The description mentions 'environment' generically but does not clarify the id format, where it comes from, or whether it is the same id returned by provision-ephemeral-environment, so it fails to compensate for the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Delete') and resource ('BunnyShell environment created for this project'), which is enough to distinguish it from the sibling provision-ephemeral-environment. It stops short of naming siblings explicitly, so it lands at 4 rather than 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to call this versus alternatives, no prerequisites, and no statement about confirming or being sure the environment is no longer needed. Teardown usage is only implied by the verb.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

fetch-execution-reportAInspect

Fetch a detailed execution report for failing SmartTests, given a batchInvocationId (batch run) or jobId (single run). Returns only failing tests and includes error details and a trace viewer URL when available.

ParametersJSON Schema
NameRequiredDescriptionDefault
jobIdNo
batchInvocationIdNo

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries full behavioral burden. It usefully discloses the filtering behavior ('Returns only failing tests') and output contents (error details, trace viewer URL 'when available'), but says nothing about permissions, param exclusivity, or behavior on empty/unknown IDs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with zero filler. The primary verb and resource are front-loaded, with the scoping and return details following compactly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema or annotations, the description does cover the key return shape (failing tests, error details, trace URL) and the dual ID input model. It is fairly complete for a read tool, missing only the required/exclusive semantics of the two IDs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does explain the meaning of both params (batchId = batch run, jobId = single run), which is genuinely additive, but it does not state whether one is required or whether passing both is allowed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Uses a specific verb ('Fetch') plus a specific resource ('detailed execution report for failing SmartTests') and narrows the scope to failing tests. It is clear about what it returns, though it does not explicitly differentiate itself from siblings like get-execution-history or get-last-run-workflow-detail.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for when to reach for it (failing SmartTests) and disambiguates the two input paths: batchInvocationId for a batch run vs jobId for a single run. It stops short of naming alternatives or exclusion conditions against the many sibling list/get tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-api-operation-detailAInspect

Fetch detailed API operation coverage (request/query/response fields, response codes, covering tests) — the operation also carries obsMappingState and the latest runtimeObservation daily summary when available. Use request volume/error rate to rank coverage risk and p95/p99 latency to prioritize performance-test work; do not treat absent observations as zero or production latency as a test threshold. Prefer TestChimp operation id (--id ULID); or rootFilePath + oasOperationId; or rootFilePath + httpMethod + pathTemplate.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNo
httpMethodNo
serviceKeyNo
pathTemplateNo
rootFilePathNo
includeManualNo
includeRemovedNo
oasOperationIdNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses data-availability semantics (obsMappingState, runtimeObservation 'when available', do not treat absent observations as zero, production latency is not a test threshold), but says nothing about permissions, rate limits, or behavior when the operation cannot be resolved.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then interpretation caveats, then parameter-selection order; each sentence carries information. The em-dash run-on makes it dense to parse, but nothing is padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 8-parameter, zero-annotation, no-output-schema tool, the description covers return content and selector logic well, but leaves three parameters (serviceKey, includeManual, includeRemoved) undocumented and does not explain resolution/failure behavior, so an agent still has gaps before calling it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% across 8 parameters, so the description must compensate. It documents the three selector patterns covering id, rootFilePath, oasOperationId, httpMethod and pathTemplate, but leaves serviceKey, includeManual and includeRemoved completely undefined in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Fetch) and resource (detailed API operation coverage) and enumerates the payload (request/query/response fields, response codes, covering tests). It is clearly a single-operation detail tool, distinguishable by name from list-api-operations, but it never names or contrasts those siblings explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides concrete usage guidance: rank coverage risk by request volume/error rate, prioritize perf work by p95/p99 latency, and prefer the id selector over rootFilePath combinations. It does not, however, say when to call this versus list-api-operations or list-api-operation-interactions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-batch-view-urlBInspect

Resolve the TestChimp batch execution viewer deeplink for the authenticated project. Pass --batch-invocation-id. Returns batchViewUrl for messaging back to the user.

ParametersJSON Schema
NameRequiredDescriptionDefault
batchInvocationIdYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full disclosure burden. It does convey that the operation is authenticated/scoped to the project and that the result is a batchViewUrl intended to be relayed to a user, which implies a read-only lookup. It does not state behavior when the ID is unknown or invalid, nor whether results are cached or live.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the action and scope, then the input and the return field. Nothing is padded, though the '--' flag phrasing is a minor stylistic slip.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with no output schema, the description covers the action, the input name, and the returned field, which is most of what an agent needs. It omits error/failure behavior and any prerequisite for obtaining a valid batch invocation ID.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and there is one required parameter, so the description must supply meaning. It only restates the parameter name in CLI-flag form ('Pass --batch-invocation-id'), adding no detail about what a batch invocation ID is, its format, or where to obtain it, and the flag notation is slightly at odds with an MCP object parameter.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (resolve) and resource (TestChimp batch execution viewer deeplink) scoped to the authenticated project, which is far more informative than a name restatement. It does not explicitly distinguish itself from near-neighbors like fetch-execution-report or get-last-run-workflow-detail, leaving some selection ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for messaging back to the user' implies the use case (produce a shareable link), which is a useful contextual hint. However, no alternative tool is named and no when-not-to-use condition is given, so the agent must infer that report-fetching siblings return data rather than a viewer URL.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-bot-compatAInspect

Return the minimum testchimp skill version, minimum @testchimp/cli version, and the bot event schema version this TestChimp deployment requires from QA bots. Call on bot startup / daily; if the local skill or CLI is older, ask the user to approve an upgrade.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose meaningful behavior: the tool is a read-style version query, what it returns, and the upgrade-approval workflow it feeds. It omits auth/permission requirements and failure behavior, but for a simple version-check read that is a minor gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with zero waste, front-loading what the tool returns before the usage guidance. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must convey return values, and it lists all three version values it provides, plus the call cadence and follow-up action. Adequate for this simple, parameterless read tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters, so per the baseline rule a 4 applies. There are no parameters whose semantics need explaining, and the description correctly adds none.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ("Return") and three concrete resources (minimum testchimp skill version, minimum @testchimp/cli version, bot event schema version). This clearly distinguishes it from siblings like get-bot-profile, register-bot-profile, and ack-bot-events; an agent can tell exactly what it yields without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says when to call it ("on bot startup / daily") and what to do with the result (ask user to approve an upgrade if local skill or CLI is older). It gives clear operational context but names no exclusions or alternative tools, so it stops short of a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-bot-profileAInspect

Fetch this QA bot's registration: botId, project, the team member it represents (userId), platform, status, role, responsibilities, capabilities, event subscriptions, paused flag, and webhook health. botId defaults to the bot-id header (TESTCHIMP_BOT_ID) or the OAuth token's bot. Use to confirm identity before onboarding and to check paused / capabilities before running routines.

ParametersJSON Schema
NameRequiredDescriptionDefault
botIdNo

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose useful behavior: botId falls back to the TESTCHIMP_BOT_ID header and then to the OAuth token's bot, and it names the returned fields including webhook health and the paused flag. It says nothing about auth/permission requirements, failure behavior when the bot is unknown, or whether this is strictly read-only.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the resource and its returned fields, then the usage guidance. The field enumeration is long but every item is a distinct piece of information an agent may need, so it earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, listing the returned fields (botId, project, userId, platform, status, role, responsibilities, capabilities, event subscriptions, paused, webhook health) is exactly the compensating detail needed, and the identity-resolution default is included. Only error/permission behavior is missing, which is minor for a read-only profile lookup.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single botId parameter has 0% schema description coverage, so the description must compensate and largely does, explaining the full default-resolution chain (header then OAuth token). It stops short of documenting format constraints or what happens on an unresolvable botId, but the key semantic is covered.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Fetch this QA bot's registration') and enumerates exactly what the registration contains, so an agent knows the payload without opening a schema. It is implicitly distinct from the sibling register-bot-profile, but it never names that alternative, which is the only thing keeping it from a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives two concrete use cases: 'confirm identity before onboarding' and 'check paused / capabilities before running routines'. That is clear context for invoking it, but there is no when-not guidance and no explicit pointer to alternatives such as register-bot-profile or get-bot-compat.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-branch-specific-endpoint-configCInspect

Resolve BASE_URL for a Git branch from TestChimp Branch Management (URL template and per-branch overrides).

ParametersJSON Schema
NameRequiredDescriptionDefault
branchNameNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It hints at the resolution mechanism ('URL template and per-branch overrides'), which is useful behavioral context, but it omits auth requirements, what happens when a branch has no override, and whether a default/template fallback is returned.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the core verb and resource lead. It is efficient, though the parenthetical could have been spent on parameter or fallback detail rather than restating the override concept.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter read tool with no output schema, the description conveys what is resolved and the override/template concept. It still leaves gaps on return format and no-override behavior, but is minimally adequate given the tool's simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

One parameter (branchName) with 0% schema description coverage, so the description must compensate. The phrase 'for a Git branch' only loosely maps to branchName and gives no format guidance (branch vs ref, full name, etc.), leaving the single parameter weakly documented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Resolve') and a precise resource ('BASE_URL for a Git branch') plus the source system (TestChimp Branch Management), so the agent knows exactly what it retrieves. It does not explicitly distinguish itself from the many get-* siblings, but the resource is distinctive enough to be identifiable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is for resolving a branch's endpoint URL but gives no explicit when-to-use guidance, no prerequisites, and no mention of alternatives (e.g. get-eaas-config or get-git-folder-mapping). The agent must infer the trigger condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-bunnyshell-workflow-job-logsCInspect

Troubleshooting: fetch logs for a BunnyShell workflow job.

ParametersJSON Schema
NameRequiredDescriptionDefault
workflowJobIdYes
bnsEnvironmentIdYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations and no output schema, the description carries the full burden, yet it only says it fetches logs. It discloses nothing about log format, pagination or truncation, retention window, permissions, or whether the environment must still exist. This is thin for an unannotated read tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short, front-loaded sentence with no wasted words. It is efficient, though its brevity is more under-specification than disciplined concision.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-required-parameter tool with zero annotation coverage, no output schema, and 0% schema descriptions, the definition is incomplete. An agent cannot tell how the ids relate, what a log payload looks like, or what failure modes to expect.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and both required parameters (workflowJobId, bnsEnvironmentId) are completely undocumented in the description. The description adds no meaning about which job id to use or why the environment id is required.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('fetch logs for a BunnyShell workflow job'), which cleanly separates it from list-bunnyshell-workflow-jobs and get-workflow-execution. It does not, however, name or distinguish itself from the most adjacent siblings, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'Troubleshooting:' prefix hints at the intended scenario (diagnosing a failed/hung job) but gives no explicit when-to-use vs when-not, and never names an alternative sibling such as get-last-run-workflow-detail. Usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-eaas-configAInspect

Return the project's BunnyShell (Environment-as-a-Service) settings. Secrets are never returned.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It adds a useful security trait by stating 'Secrets are never returned,' but does not disclose read-only nature, authentication requirements, or return format.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no wasted words. The core purpose is front-loaded, and the security constraint follows efficiently.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read tool with no annotations or output schema, the description is largely sufficient: it identifies the resource and notes that secrets are excluded. It could be more complete by clarifying what settings are included or whether it applies to the current project, but the essential context is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the schema provides no parameter semantics to compensate for. The description appropriately adds no parameter details, matching the baseline for parameterless tools.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Return') and resource ('project's BunnyShell (Environment-as-a-Service) settings'), making the tool's purpose clear. It does not explicitly differentiate from sibling config or getter tools, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to use this tool versus alternatives, nor any exclusions or prerequisites. The description only implies usage by stating what it returns.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-ephemeral-environment-statusCInspect

Poll BunnyShell for environment status and component_urls_json.

ParametersJSON Schema
NameRequiredDescriptionDefault
bnsEnvironmentIdYes

TDQS

C2.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does not state whether this is a read-only operation, how long an environment typically remains in a transient state, whether component_urls_json is only populated once provisioning completes, or how often the caller should poll. Only the vague verb 'Poll' hints at repeated invocation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with no filler, front-loaded with the act (poll) and the target (BunnyShell environment status). It is efficient but under-specified rather than well-structured.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

One required parameter is undocumented, no output schema exists, and no annotations are present, so the description is the only specification an agent has. It omits the parameter semantics, safety profile, return content (beyond a cryptic JSON name), and polling guidance, leaving the tool under-specified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the single required parameter 'bnsEnvironmentId' has no description, so the description must compensate but does not. It never mentions the parameter, its format, where the ID comes from, or whether it is the environment ID returned by provision-ephemeral-environment.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description restates the tool name: 'get-ephemeral-environment-status' → 'Poll BunnyShell for environment status'. The addition of 'component_urls_json' is the only real information, and it's an ambiguous internal artifact name rather than a description of what data is returned. It does not distinguish this from siblings like list-bunnyshell-environment-events.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance is given. The word 'Poll' weakly implies a repeated-check pattern, but the description does not say when to call this versus provision-ephemeral-environment-and-wait (which presumably also yields readiness) or list-bunnyshell-environment-events. No exclusions are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-execution-historyAInspect

Fetch SmartTest execution history for a testId (top 5 recent runs), an optional platform-rooted folder/file scope, or a scenario when scenarioId is set. Prefer testId when you have it from fetch-execution-report. Typically omit environment to avoid env scoping. Use branchName and scope.filePaths as for coverage. Optional platform (web|ios|android) and dimensionFilters narrow results.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
scopeNo
offsetNo
testIdNo
releaseNo
platformNo
branchNameNo
scenarioIdNo
environmentNo
dimensionFiltersNo

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses that results default to the top 5 recent runs and warns against env scoping, but omits read-only confirmation, permission requirements, pagination behavior (limit/offset exist in schema), and the shape of returned history entries.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, front-loaded with the primary input modes before the optional narrowing parameters. No filler, though the coverage-comparison clause ('as for coverage') is slightly cryptic for an agent without external context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter tool with nested objects, no annotations, and no output schema, the description covers the selection-relevant parameters but leaves limit/offset/release unexplained and says nothing about what a history record contains. Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 10 params, so the description must compensate, and it does for most: testId, scenarioId, environment, branchName, scope.filePaths, platform, and dimensionFilters. It leaves limit, offset, and release entirely unaddressed, which keeps it from a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (Fetch) and resource (SmartTest execution history) and lays out three distinct input modes: testId, folder/file scope, or scenarioId. It references the sibling fetch-execution-report as the source of testId, giving partial differentiation, but does not contrast its own output with that sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit routing advice: 'Prefer testId when you have it from fetch-execution-report' and 'Typically omit environment to avoid env scoping.' It also points to branchName/scope.filePaths usage 'as for coverage.' No explicit when-not-to-use-this-tool statement, so not a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-git-folder-mappingBInspect

Return git provider, repository, mapped plans/tests folder paths, and plans branch (empty = repository default) for the project.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full disclosure burden. It implies a safe read ("Return") and adds a genuinely useful behavioral detail: the plans branch is empty when the repository default applies. It does not state permission requirements, scoping (per project), or whether the mapping can be absent, so disclosure is partial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler, leading with the verb and then the enumerated return values. The parenthetical about empty branch values is the only nuance and it earns its place, though the dense enumeration makes it slightly list-like.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must convey return content, and it does so by naming each field returned. Combined with zero input parameters, this is nearly complete; only the shape/nesting of the response and the absence case are left unspecified.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there are no parameter semantics to document; the baseline for a 0-param tool is 4. The description correctly avoids inventing parameter detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb+resource ("Return git provider, repository, mapped plans/tests folder paths, and plans branch") and even enumerates the returned fields, so the agent knows exactly what this yields. It doesn't differentiate itself from the sibling update-git-folder-mapping or explain the read vs. write relationship. Still, the purpose itself is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, prerequisites, or alternative tools are mentioned. An agent is not told to call this before update-git-folder-mapping, or in what workflow context the mapping is needed. Only the implied read semantics suggest usage.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-issue-detailsAInspect

Fetch a TestChimp issue (bug) by ordinal id. Accepts flexible issueId formats: #B-123, B-123, #B123, B123, or plain 123. Returns title, description, status, linked entities, artifact references, and short-lived signed URLs for GCS attachments/screenshots.

ParametersJSON Schema
NameRequiredDescriptionDefault
issueIdYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the full burden and does meaningful work: it enumerates the returned fields and discloses the non-obvious behavioral trait that attachment/screenshot URLs are short-lived and signed. It stops short of stating permission requirements or not-found/error behavior, so it is strong but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences: identity and key first, accepted input formats second, return contents third. No filler and nothing that restates the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description compensates well by describing both the input formats and the return payload, including the short-lived signed URL caveat. Minor gaps remain around failure modes (invalid or nonexistent id) and the fact that this is a non-mutating read.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the sole parameter has no description in the schema, yet the description fully compensates by spelling out five accepted issueId formats (#B-123, B-123, #B123, B123, plain 123) and clarifying that the id is an ordinal. This is exactly the compensating detail a low-coverage schema requires.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Fetch') and resource ('TestChimp issue (bug)') with the keying mechanism ('ordinal id'), which cleanly separates it from siblings like create-issue, update-issue-status, and the various get-meeting-*/get-release-* tools. An agent can identify the target entity without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the fetch verb and the entity name, but there is no explicit when-to-use guidance, no mention of when the issue does not exist, and no note about interactions with the similarly-scoped create-issue / update-issue-status tools. Adequate but leaves routing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-last-run-workflow-detailBInspect

Fetch the last workflow execution for a workflow-id on a branch (optional userId for per-user last run). Used for since-last-run scoping.

ParametersJSON Schema
NameRequiredDescriptionDefault
userIdNo
branchNameNo
workflowIdYes

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. 'Fetch' implies a read, but it says nothing about what happens when no prior run exists (critical for a 'last run' tool), auth requirements, or pagination. The only behavioral hint is the 'since-last-run scoping' use case.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the core action front-loaded and no filler. Efficient, though the parenthetical is slightly dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read tool with no output schema and no annotations, the description covers the core action and parameter roles but omits edge-case behavior (no-run-found) and any return-shape hint, leaving real gaps given 0% schema coverage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does map all three params: workflow-id, branch, and the optional userId 'for per-user last run'. It clarifies meaning but adds no format, syntax, or default details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Fetch the last workflow execution') and qualifies it by workflow-id and branch, which is enough to separate it from list-workflow-executions and get-workflow-execution. However it never explicitly names the sibling it differs from or explains the 'last' scoping relative to those tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The trailing clause 'Used for since-last-run scoping' gives a usage context, but there is no when-to-use-vs-alternatives guidance, no exclusions, and no mention of when to prefer list-workflow-executions or get-workflow-execution instead.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-manual-session-detailsAInspect

Fetch a manual test session by id. Returns project id, title, environment, steps (playwright commands, signed screenshot URLs, notes), and linked scenario ordinal ids. Use when authoring a SmartTest from a recorded manual session.

ParametersJSON Schema
NameRequiredDescriptionDefault
manualSessionIdYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does well by disclosing the response shape (project id, title, environment, steps with playwright commands and signed screenshot URLs, notes, linked scenario ordinal ids). It omits operational traits such as permission requirements or whether signed screenshot URLs expire, which keeps it below 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, with the core action front-loaded and the return payload following; every clause carries information and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter read tool with no output schema and no annotations, the description covers both invocation (by id) and return contents, which is the key gap-filling. Minor omissions (id format, auth, URL expiry) keep it from being fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single manualSessionId parameter, and the description only adds 'by id' without stating the expected id format or origin. The parameter name is largely self-describing, but the description does not fully compensate for the missing schema documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource ('Fetch a manual test session by id') and even enumerates the returned payload, so the tool's function is unambiguous. It does not explicitly distinguish itself from adjacent siblings such as get-test-scenarios or create-test-scenario, so it falls short of the top tier.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It supplies a clear usage context: 'Use when authoring a SmartTest from a recorded manual session,' which effectively routes the agent toward this tool in that workflow. There is no explicit when-not condition or named alternative, so it stops short of 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-meeting-setAInspect

Fetch a meeting-set context by id (ULID from /testchimp using meeting-set context <id>, created on the Meetings page via Start Chat). Returns the filters / search text applied and the meetings in scope (id, title, start). Fetch each meeting with get-meeting-transcript (summaryOnly first). Sets expire after 7 days.

ParametersJSON Schema
NameRequiredDescriptionDefault
meetingSetIdYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose real behavioral traits: the return payload (filters/search text plus in-scope meetings with id, title, start) and a lifecycle constraint ('Sets expire after 7 days'). It omits permission/auth requirements and error behavior, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences, front-loaded with purpose and then built outward to provenance, return shape, next step, and expiry. Dense and largely waste-free, though the return-contents clause and expiry note could be sequenced more crisply.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema and no annotations exist, so the description must cover both, and it does describe the return fields and the 7-day expiry. It leaves auth requirements and failure modes unaddressed, which keeps it just under fully complete for a fetch tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the schema only says 'string, minLength 1', so the description does the work by explaining that meetingSetId is a ULID sourced from the `/testchimp using meeting-set context <id>` flow or the Meetings page. That is meaningful added semantics, though format/validation detail is still thin.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Fetch a meeting-set context by id') and then defines what a meeting-set is and how the id is produced, so an agent understands the object without opening a schema. It also distinguishes itself from the sibling get-meeting-transcript by naming it as the follow-up step.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear context for use: the id is a ULID obtained via the meeting-set context flow, and the recommended next action is 'Fetch each meeting with get-meeting-transcript (summaryOnly first)'. It stops short of explicit when-not-to-use or alternative retrieval paths, so it is not a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-meeting-transcriptAInspect

Fetch a cloud-synced Meeting Bots transcript by meeting id (calendar event id, or URL hash for ad-hoc meetings). Returns title, start time, post-meeting summary, and transcript. Use summaryOnly to fetch just the summary (much smaller) and pull the full transcript only when the summary is not enough. Prefer local ~/.testchimp/data/meetings//transcript.md when present on Studio / the recording machine.

ParametersJSON Schema
NameRequiredDescriptionDefault
meetingIdYes
summaryOnlyNo

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden, and it does disclose the return shape (title, start time, post-meeting summary, transcript) and that the source is cloud-synced versus a local file. It omits auth/permission requirements, but read semantics and output contents are covered well for a tool with zero annotation support.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences, front-loaded with the verb, resource, and id format before the optional-mode and local-file guidance. Every sentence adds distinct information with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description supplies the return values, the two id forms, and the summaryOnly tradeoff, which is enough to call the tool correctly. It could have noted permissions or the absence/presence of transcript for summary-only meetings, leaving a small gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate and it largely does: meetingId is explained as a calendar event id or URL hash for ad-hoc meetings, and summaryOnly is explained as returning just the (much smaller) summary. Both of the two parameters get meaningful semantics beyond their bare types.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Fetch) and resource (cloud-synced Meeting Bots transcript), scoped to retrieval by meeting id. Distinguishes the resource from sibling get-meeting-set and list-meetings. An agent knows exactly what this returns without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit routing guidance: use summaryOnly when the summary suffices and pull the full transcript only when it doesn't, plus a preference for the local transcript.md file when present. These are concrete when-to-use conditions, though no sibling tool is named as an alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-my-tasksAInspect

List a team member's open QA work in this project: manual test scenarios assigned to them in test runs (with testRunUrl), issues assigned to them (status, severity, due date, issueUrl), and SmartTests they authored that still await verification. With an OAuth token the user is the token's user; with an API key pass userId. Use for daily reminders and 'what should I work on' questions.

ParametersJSON Schema
NameRequiredDescriptionDefault
userIdNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full disclosure burden and does meaningful work: it explains that auth mode changes behavior (OAuth token resolves to the token's user; API key requires userId) and enumerates the fields returned per category (testRunUrl, status, severity, due date, issueUrl). It stays silent on pagination, result limits, and what happens when the user has no open work, which keeps it out of the 5 range.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with purpose, then content, then auth mechanics, then usage. Every clause (including the parenthetical field lists) adds information an agent needs, and there is no restatement of the name or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter list tool with no annotations and no output schema, the description is nearly self-sufficient: it covers project scope, the three data categories, field-level return detail, and conditional auth. Remaining gaps are pagination and any time-window constraint on "open" work, so it is complete but not exhaustive.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, yet the description supplies the only semantic guidance for userId: it is required only when authenticating with an API key, and is unnecessary with an OAuth token where the user is derived from the token. That is exactly the contextual meaning the bare string schema lacks, though the expected identifier format is never described.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a concrete verb+resource ("List a team member's open QA work") and enumerates the three result categories — assigned manual test scenarios, assigned issues, and authored SmartTests awaiting verification — so an agent knows precisely what it returns. It implicitly scopes to the calling user, which separates it from generic siblings like get-test-scenarios or list-tests-awaiting-verification, but it never names an alternative explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states a clear use context — "Use for daily reminders and 'what should I work on' questions" — plus the auth-mode condition that selects the argument handling. It offers no explicit when-not guidance or named alternatives, so it stops short of the top score.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-org-capabilitiesAInspect

Fetch the organization's enabled capabilities (e.g. TRUE_COVERAGE, API_CONTRACT_COVERAGE) and freeTrialActive flag. Call before relying on TrueCoverage / API contract coverage features so playbooks can soft-skip gated insights instead of failing. Authenticated via project API key.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses the authentication requirement ('Authenticated via project API key') and the intended behavior of soft-skipping gated insights rather than failing. It implies a read-only fetch, though it does not describe return shape or failure modes in detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tightly written sentences, each earning its place: purpose, usage guidance, and auth requirement. The core purpose is front-loaded and no extraneous detail is included.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read tool with no output schema, the description covers what it returns, when to call it, and authentication. It could be slightly richer about the exact response structure, but it is sufficient for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is no parameter-level semantics to document. The description appropriately names the returned capability examples and freeTrialActive flag, providing useful context even without an input schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Fetch the organization's enabled capabilities ... and freeTrialActive flag.' It names concrete capability examples (TRUE_COVERAGE, API_CONTRACT_COVERAGE), which distinguishes it from sibling coverage tools like get-truecoverage-events and get-truecoverage-event-details.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear when-to-use guidance: 'Call before relying on TrueCoverage / API contract coverage features so playbooks can soft-skip gated insights instead of failing.' It does not name alternative tools or explicit when-not conditions, but the triggering context is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-perf-runAInspect

Fetch one performance run by runId; set includeRaw to include its raw payload.

ParametersJSON Schema
NameRequiredDescriptionDefault
runIdYes
includeRawNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. 'Fetch' implies a read operation and it does disclose the includeRaw toggle's effect, but it says nothing about permissions, behavior when runId is unknown, or response shape.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler; the core action and the one behavioral option are both stated without waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter read tool this covers the minimum, but with no output schema and no annotations it never describes what a performance run contains or what the raw payload is, which an agent would benefit from knowing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It adds real meaning for includeRaw ('include its raw payload') but 'by runId' merely restates the parameter name, leaving half the parameters effectively undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Fetch one performance run') and names the key identifier (runId), which clearly separates it from the list-perf-runs sibling. It does not, however, explicitly call out that distinction, leaving the agent to infer single-vs-list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied: the presence of runId suggests you already know which run you want. There is no explicit when-to-use vs list-perf-runs or compare-perf-to-baseline, and no prerequisites are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-plans-support-fileAInspect

Fetch a file under the mapped plans root from the platform by relative path (no git required). Primary use: load a workflow execution plan named in a Continue Locally / implement prompt before falling back to the repo copy. filePath is relative to the plans mapped root (leading plans/ is stripped). Under workflow_plans/, filenames are coerced to *.plan.md. Response: found (false if missing), supportFileId, filePath (canonical), filetype, content.

ParametersJSON Schema
NameRequiredDescriptionDefault
filePathYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose real behavior: no git dependency, path normalization (leading 'plans/' stripped), filename coercion to *.plan.md under workflow_plans/, and a missing-file signal via found=false. It omits auth/permission requirements and any caching or error behavior beyond the found flag.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose and scope, then path semantics, then the response shape — dense and mostly waste-free. The response field enumeration is justified given there is no output schema, though the sentence is packed enough to require a second read.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter fetch with no output schema, the description covers purpose, path resolution rules, missing-file behavior, and the returned fields, which is close to what an agent needs. The remaining gaps are auth requirements and non-'not found' failure behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the single parameter is undocumented in the schema, so the description must compensate — and it largely does, explaining that filePath is relative to the plans mapped root, that a leading 'plans/' is stripped, and that names under workflow_plans/ are coerced to *.plan.md. It still gives no concrete example path or note on the minLength constraint.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource ('Fetch a file under the mapped plans root from the platform by relative path') and a distinguishing qualifier ('no git required'). It does not name its obvious write counterpart upsert-plans-support-file or other retrieval siblings, so it stops short of full sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the primary use case explicitly ('load a workflow execution plan named in a Continue Locally / implement prompt before falling back to the repo copy'), which tells the agent when this tool is the right first choice. It does not name an alternative tool for the fallback or any exclusion conditions, so guidance is contextual rather than complete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-policyAInspect

Fetch a workflow policy file by name (e.g. run-qa.policy.md) from the platform POLICY_FILE store. Filename is coerced to *.policy.md (same as upsert-policy).

ParametersJSON Schema
NameRequiredDescriptionDefault
policyFileNameYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations the description carries the full burden, and it does disclose one genuinely useful trait beyond the schema: the filename is coerced to *.policy.md, matching upsert-policy. It omits what a fetch returns, error behavior when the file is absent, and any auth/permission context, so the behavioral picture is incomplete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the action and resource, with no filler. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple single-parameter read tool with no annotations and no output schema, the input semantics are adequately covered. However, it says nothing about the return value (file contents) or failure mode for a missing policy, leaving meaningful gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and the single property has no description, so the description must compensate. It does so well by explaining that policyFileName is coerced to *.policy.md and giving a concrete example, adding real semantics beyond the bare string type.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (Fetch) plus resource (workflow policy file) plus store (POLICY_FILE) and a concrete example filename. The purpose is unambiguous, but sibling differentiation (e.g. vs list-policies) is only implicit rather than stated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied: fetch when you know the policy file name. The reference to upsert-policy concerns the filename-coercion behavior, not when to choose this tool over list-policies or another sibling, so no explicit when-to-use routing is given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-project-init-statusBInspect

Return project-init progress for the current project (platform comms, folder mapping, test env, CI, optional imports).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. 'Return' implies a read-only, non-mutating call, and the enumerated components give some sense of what is inspected, but nothing is said about permissions, error behavior when a project is uninitialized, or rate limits.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler; the parenthetical is an efficient way to enumerate the tracked components without extra clauses.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description is the only documentation of the response, yet it lists only the tracked areas, not their shape (statuses, booleans, percentages, per-item errors). Adequate to identify the tool but thin for interpreting results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the baseline of 4 applies. The description correctly implies no input is needed, operating on the implicit current project context.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Return') and resource ('project-init progress') with clear scope ('for the current project'), and the parenthetical enumerates exactly which init components are reported. The read-only nature distinguishes it from the sibling update-project-init-status, though that contrast is only implicit in the verb choice.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no prerequisites (e.g., project must exist/be initialized), and no named alternatives. The agent must infer from the tool name alone that this is the read counterpart to update-project-init-status.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-qa-postureAInspect

Project-wide QA posture snapshot: releases (version, lifecycleStatus, due date, passed / failed / blocked / not-attempted counts), issues (active, inProgress, blocked, openBySeverity), activeTestRuns with result counts, and testsAwaitingVerificationCount. Use for weekly QA digests and release-health questions; summarise for the reader's role rather than dumping raw JSON.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations and no output schema, the description carries the full burden and largely meets it by enumerating the payload's top-level sections and nested counts. It also adds a behavioural instruction to summarise per the reader's role rather than dumping JSON. It does not state freshness, caching, or scope requirements (e.g., implicit project context), leaving a modest gap for a read tool with no annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One dense payload sentence followed by a short usage/behaviour sentence, with the core purpose front-loaded. The field list is long but each item is informative rather than filler, so it earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter aggregate with no output schema, the description compensates well by naming the returned sections and counts, so an agent knows what it will get. Remaining gaps are minor: no statement of data freshness or how 'project-wide' scope is determined.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is no parameter semantics to document beyond noting the snapshot is implicitly project-scoped. Baseline 4 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete resource (project-wide QA posture snapshot) and enumerates the specific contents: releases with lifecycle/status counts, issue breakdowns, active test runs, and testsAwaitingVerificationCount. This is clearly distinguishable from siblings like get-release, get-release-details, or get-suite-execution-stats, which return narrower slices.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit use cases ('weekly QA digests and release-health questions'), which tells the agent when this aggregate view is the right choice over per-entity getters. It stops short of naming an alternative or stating when-not to use it, so it is clear context without exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-releaseBInspect

Fetch release catalog details for a version/label in the current project (McpGetReleaseRequest/Response: cut git SHA, prior release + SHA, focus areas, payload). Pass version (CLI: --version). Authenticated via project API key.

ParametersJSON Schema
NameRequiredDescriptionDefault
versionYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose the auth requirement ('Authenticated via project API key') and that the operation reads a catalog. It omits whether the call is read-only, what happens for an unknown version, and any rate or caching behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with the resource and scope first, then the payload parenthetical and the parameter note. Dense but every clause carries information; the CLI-flag aside is slightly incidental.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter read with no annotations or output schema, the description is only partially complete: it names some return fields and the auth model, but leaves the sibling ambiguity with get-release-details and the parameter format unresolved.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and there is one required parameter, so the description must compensate. It clarifies that the argument is a version/label and maps it to the CLI flag --version, but never defines the accepted format (semver, label string, branch) or how to obtain a valid value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Fetch release catalog details for a version/label in the current project') and even sketches the payload contents. However, it never distinguishes itself from the near-identical sibling get-release-details, leaving the agent unable to tell which release tool to pick.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance, no exclusion criteria, and no mention of alternatives. The only contextual hint is the phrase 'in the current project', which implicitly scopes the call but does not tell the agent when this is the right tool over get-release-details.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-release-detailsAInspect

Fetch gate-oriented release details for a version/label: scope, per-environment priority×status test stats, open issue stats, scan summaries, and detailed in-scope scenario/issue records (McpGetReleaseDetailsRequest/Response). Pass version (CLI: --version). Use for CI/agent release gating. Authenticated via project API key.

ParametersJSON Schema
NameRequiredDescriptionDefault
versionYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the disclosure burden. It adds that authentication is via a project API key, which is useful. But it doesn't state whether this is read-only (implied by 'get'), pagination behavior, or result size; the mention of McpGetReleaseDetailsRequest/Response is a naming reference, not behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Single dense sentence that front-loads what is returned, followed by the parameter and usage notes. It is information-dense but not wasteful, though the McpGetReleaseDetailsRequest/Response parenthetical is low-value for an agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read tool with no annotations and no output schema, the description does reasonably cover return payload shape and usage context, but leaves gaps on parameter format and behavioral traits like rate limits or result size — an agent could invoke it but with unresolved uncertainty.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

One parameter with 0% schema description coverage, so the description must compensate. It provides a hint ('version/label') and a CLI flag mapping (--version), but doesn't clarify the accepted format (semantic version, tag, branch name) or whether 'label' is an accepted alias.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (fetch) and resource (gate-oriented release details) and enumerates the returned data scope (priority×status stats, issue stats, scan summaries, scenario/issue records). It is somewhat differentiated from a sibling get-release, but the distinction between the two isn't made explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states usage context: 'Use for CI/agent release gating.' This tells the agent when to reach for this tool. However, it does not mention when NOT to use it or name the alternative sibling (get-release) for non-gating needs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-requirement-coverageAInspect

Fetch requirement (scenario) coverage under an optional platform-rooted folder scope (tests/... or plans/...). Use scope.filePaths or scope.folderPath (platform tests/plans roots). Omit branchName for cross-branch coverage (aggregates branch copies; execution jobs deduped by stable hash of tests-root-relative path + test name). Pass branchName only when results must be limited to one Git branch. Optional platform (web|ios|android) filters rollup. For top-N gap recommendations: set scenarioLifecycleStatuses, considerScenarioPriority / considerSemanticCoverage, and limit; prefer response rankedScenarios (gaps only). Server excludes verification_strategy=manual by default (autoVerificationOnly).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
scopeNo
releaseNo
platformNo
branchNameNo
environmentNo
recordTypesNo
autoVerificationOnlyNo
considerScenarioPriorityNo
considerSemanticCoverageNo
scenarioLifecycleStatusesNo
includeNonCoveredUserStoriesNo
includeNonCoveredTestScenariosNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose meaningful behavior: the server default excluding verification_strategy=manual via autoVerificationOnly, branch aggregation/dedup semantics, and rollup filtering by platform. It omits permissions or pagination behavior, but the defaults and aggregation rules are genuinely useful non-obvious context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose then dense parameter/behavior detail in a single efficient paragraph; nearly every clause carries information. It is a touch run-on but not padded or repetitive.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 13-param, nested-object tool with no output schema, it covers scope, branch semantics, defaults, and hints at the response shape (rankedScenarios for gaps). Remaining gaps are the undocumented params and the absence of return/output detail, but the description is substantially complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate and does so for many of the 13 params: scope.filePaths/folderPath, branchName, platform, scenarioLifecycleStatuses, considerScenarioPriority, considerSemanticCoverage, limit, and autoVerificationOnly. It leaves release, environment, recordTypes, and the two includeNonCovered* flags undocumented, so it does not fully close the gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Fetch requirement (scenario) coverage,' with a clear scope qualifier (platform-rooted tests/plans folder). An agent can tell it retrieves coverage data rather than scenarios or events. It does not name any sibling for differentiation, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives real conditional guidance: 'Omit branchName for cross-branch coverage' vs 'Pass branchName only when results must be limited to one Git branch,' plus a prescribed recipe for top-N gap recommendations. It never names an alternative sibling tool, so it is strong context without explicit when-not/alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-requirement-quality-reportAInspect

Fetch the stored requirement quality report (metrics + findings with user states) for a user story or test scenario. Use before local DeFOSPAM to dedupe: do not re-report findings already IGNORED or APPLIED (match by fingerprint). Pass subjectType STORY|SCENARIO plus subjectEntityId or ordinalId (numeric part of US- / TS-). When no prior report exists, response still includes report.subject with resolved subjectEntityId.

ParametersJSON Schema
NameRequiredDescriptionDefault
ordinalIdNo
subjectTypeYes
subjectEntityIdNo

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose non-obvious behavior: the fingerprint-based dedupe contract and the empty-state behavior ('when no prior report exists, response still includes report.subject with resolved subjectEntityId'). It omits auth/permission requirements and error behavior, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, front-loaded with the fetch purpose, then the dedupe rule, then the parameter and edge-case notes. No filler, though 'DeFOSPAM' is domain jargon that will not help an unfamiliar agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers purpose, dedupe semantics, parameter addressing and the empty-result case despite having no annotations, no output schema and 0% schema coverage. Auth requirements and the shape of the returned findings remain unstated, so it is not fully complete for a 3-param tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does: it explains subjectType is STORY|SCENARIO, that subjectEntityId or ordinalId is accepted, and that ordinalId is the numeric part of US-<n>/TS-<n>. It does not define the expected format of subjectEntityId itself.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Fetch the stored requirement quality report (metrics + findings with user states)' scoped to a user story or test scenario. This clearly distinguishes a read of stored findings from the write-side sibling report-requirement-quality-findings, though it never names that sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit when ('Use before local DeFOSPAM to dedupe') and a when-not ('do not re-report findings already IGNORED or APPLIED, match by fingerprint'). Strong operational context, but it does not name the alternative tool an agent should call when it does want to report.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-security-scan-configBInspect

Fetch security scan config by scan id (detail.dastCheckConfig / detail.sastCheckConfig / detail.depsCheckConfig / detail.leaksCheckConfig, release label, status). Pass id (CLI: --id). Used by /testchimp run security scan. For DAST honour allowActiveScan / useEphemeralSandbox / scope on dastCheckConfig.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It adds meaningful context by disclosing the shape of the returned config and the DAST-specific fields to honour (allowActiveScan / useEphemeralSandbox / scope), which is real behavioral value. However, it says nothing about read-only nature, auth requirements, or failure behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is front-loaded with the core verb+resource, but the two parenthetical blocks (detail.* field list and DAST fields) make it a dense run-on that is harder to parse. It is compact but the nested lists reduce readability.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter read tool with no output schema, the description partially compensates by describing the returned config structure. It still leaves gaps around the id's origin, error cases, and permissions, so it is adequate but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% for the single 'id' parameter, so the description must compensate. It clarifies that the id is a scan id and gives the CLI flag (--id), but omits the id format or where to obtain it, leaving the parameter only partially documented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource ('Fetch security scan config by scan id') and enumerates what the config contains (dast/sast/deps/leaks check configs, release label, status). It clearly distinguishes itself from the report-* and update-* siblings, though it does not name any alternative tool explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides some context ('Used by /testchimp run security scan') and how to pass the id (CLI: --id), but gives no explicit when-to-use vs. when-not, no prerequisites, and no routing to sibling tools. Usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-spec-lifecycle-detailsAInspect

Fetch lifecycle_fields for user stories and/or test scenarios by ordinal id (DB only; no markdown). Pass scenarioIds / storyIds as lists of bare ordinals (canonical) or prefixed forms (TS-107, #US-12). Use after identifying scenarios in scope for create-tests to read verification_strategy (auto|manual) and skip manual ones.

ParametersJSON Schema
NameRequiredDescriptionDefault
storyIdsNo
scenarioIdsNo

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses the data source ('DB only; no markdown') and what is returned (lifecycle_fields including verification_strategy), but says nothing about permissions, pagination, empty/not-found behavior, or cost. Adequate but with real gaps for a tool with zero annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the verb+resource, then ID-format syntax, then usage workflow. No filler and every sentence adds operational information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description steps in by naming the returned field (lifecycle_fields) and the key sub-value (verification_strategy). Combined with ID-format guidance and usage context, an agent can call it correctly; only permissions, pagination, and missing-ID behavior are unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it does: it names both parameters (scenarioIds/storyIds), clarifies they are lists, and specifies accepted ID forms — bare ordinals as canonical, or prefixed forms like TS-107 and #US-12. It leaves optionality/combination rules to the 'and/or' phrasing rather than stating them, so not quite a 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Fetch) and resource (lifecycle_fields for user stories and/or test scenarios) plus the lookup key (ordinal id). The '(DB only; no markdown)' qualifier distinguishes it from markdown-backed siblings such as get-user-stories or get-test-scenarios.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit workflow context: 'Use after identifying scenarios in scope for create-tests to read verification_strategy (auto|manual) and skip manual ones.' This tells the agent both the preceding step and the downstream purpose. It does not state when-not to use it or name alternative retrieval tools, so it falls short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-suite-execution-statsAInspect

Aggregate suite timing from list_execution_history testStats (same filters as get-execution-history). Returns testCount, timedTestCount, sumSuccessMeanSecs (UI ExecutionTimingSummary total), sumFailMeanSecs, maxSuccessMeanSecs, slowTestCount (success mean > 30s). Sum of means is aggregate test CPU-time, not CI wall-clock under parallelism. Compare sumSuccessMeanSecs against any suite budget in the agent — this tool only returns stats.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
scopeNo
offsetNo
testIdNo
releaseNo
platformNo
branchNameNo
scenarioIdNo
environmentNo
dimensionFiltersNo

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it enumerates returned fields, defines slowTestCount (success mean > 30s), and warns that the sum of means is aggregate CPU-time rather than CI wall-clock under parallelism. It omits auth/permission and rate-limit context, keeping it just short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and packed with useful specifics; each sentence (returns list, CPU-vs-wall-clock caveat, budget comparison) earns its place. Dense but not padded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Since no output schema exists, the description appropriately enumerates the returned fields, and the caveat is valuable. But with 10 parameters at 0% schema coverage and no annotation coverage, the definition is incomplete on the input side for a tool of this complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 10 parameters, including a nested scope and dimensionFilters object, and the description explains none of them individually. Pointing to 'same filters as get-execution-history' is only partial compensation for such a large undocumented surface.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (aggregate suite timing stats) and distinguishes itself from siblings by noting it derives from list_execution_history testStats and shares filters with get-execution-history. An agent can tell this is the aggregated-stats variant rather than a raw history fetch.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implies usage by stating 'this tool only returns stats' and referencing get-execution-history for filters, and advises comparing sumSuccessMeanSecs against a suite budget. However it never explicitly states when to pick this over get-execution-history or what the filters actually are, leaving routing partly to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-test-scenariosAInspect

Fetch test scenarios from the TestChimp platform by ordinal id (numeric part of TS-) and/or external TMS ids (e.g. C12345, PROJ-101 — server strips prefixes and matches numerical part). Returns full plan markdown content, title, platform file path, linked user story ordinal ids, and external_source / external_system_id when present. Use when plan files are not yet synced to the repo, or when linking imported tests to scenarios by TMS id.

ParametersJSON Schema
NameRequiredDescriptionDefault
externalIdsNo
scenarioOrdinalIdsNo

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the burden and does disclose non-obvious behavior: the server strips ID prefixes and matches only the numerical part, and both id families can be combined ('and/or'). It omits auth/permission needs, empty-result behavior, and pagination, so it is good but not exhaustive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly packed sentences, front-loaded with the action and identification mechanism, then the return payload, then usage. No filler sentences; every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description usefully enumerates the return payload (plan markdown, title, file path, linked user story ids, external source/system id) and the input-id semantics. It leaves out only secondary operational details like permission requirements and empty-result handling.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate and does: it explains scenarioOrdinalIds as the numeric part of TS-<n> and externalIds as TMS ids with prefix-stripping and concrete examples (C12345, PROJ-101). This fully documents both otherwise-undescribed parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Fetch) and resource (test scenarios) plus the keys used to identify them, so the tool's job is unambiguous. It does not explicitly name a sibling (e.g. list-test-scenarios-for-scope) to distinguish itself from, which keeps it short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear when-to-use context: 'when plan files are not yet synced to the repo, or when linking imported tests to scenarios by TMS id.' This implicitly routes the agent away from local-plan lookups, but it never names an alternative tool explicitly, so it stops short of full when/when-not/alternatives guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-truecoverage-child-event-treeAInspect

TrueCoverage next-event tree after an event (ListChildEventTreeRequest). Requires eventTitle, baseScope (environment + timeWindow); optional coverageScope for PRESENT/ABSENT. Note: metadataFilters on scopes are ignored for transition stats. Field names are baseScope/coverageScope (not baseExecutionScope).

ParametersJSON Schema
NameRequiredDescriptionDefault
baseScopeYes
eventTitleYes
coverageScopeNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden, and it does add one genuinely valuable behavioral caveat: metadataFilters on scopes are ignored for transition stats. It also pre-empts a common naming error. It still omits read-only nature, permission needs, and any rate/limit behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence with the purpose front-loaded and zero filler. The parenthetical request-type and naming caveat earn their place, though the sentence is packed enough to be slightly hard to parse on first read.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description should say what the 'tree' returns, and it doesn't. Given a deeply nested input schema, no annotations, and no output schema, the behavioral coverage is respectable but leaves the result shape opaque.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Despite 0% schema description coverage, the description explains baseScope's composition (environment + timeWindow), the role of coverageScope, and warns that the fields are baseScope/coverageScope rather than baseExecutionScope. That is meaningfully beyond the raw schema, though it doesn't cover the nested window formats or metadataFilters shape.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete resource and scope: a TrueCoverage next-event (child) tree after a given event, with the underlying request type named. It doesn't differentiate from the closely-named sibling get-truecoverage-event-transition, so the agent can't route between the two on the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a usage condition for the optional parameter ('optional coverageScope for PRESENT/ABSENT') and names the two required inputs, which is useful trip-planning. It never says when to prefer this tool over the sibling transition/event-tree tools, so the when-vs-alternatives guidance is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-truecoverage-event-detailsBInspect

TrueCoverage drill-down for one event (GetEventDetailsRequest). Requires eventTitle plus baseExecutionScope (environment + timeWindow). Optional comparisonExecutionScope for coverage columns (automationEmitsOnly on comparison). Same timeWindow / platform rules as get-truecoverage-events.

ParametersJSON Schema
NameRequiredDescriptionDefault
eventTitleYes
baseExecutionScopeYes
comparisonExecutionScopeNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It usefully discloses that supplying comparisonExecutionScope yields coverage columns and that automationEmitsOnly applies only to comparison/coverage scopes, but it says nothing about auth requirements, return format, or scaling behavior for a deeply nested request.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tightly packed sentences, front-loaded with the core purpose and then the required/optional scope structure. Very little waste, though it leans on the sibling reference to carry the timeWindow/platform rules.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-param tool with deeply nested objects, no annotations, and no output schema, the description covers the required skeleton and the comparison-coverage behavior but omits the nested metadata filter semantics and platform/release meaning. It is adequate but not complete enough for a complex nested request.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is reported at 0%, so the description must compensate. It adds real meaning for the three top-level params (eventTitle, baseExecutionScope = environment + timeWindow, comparisonExecutionScope's purpose), but leaves nested inputs like metadataFilters, release, branchName, and the platform enum values unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ("drill-down for one event") and scopes it to a single event, which distinguishes it from the plural get-truecoverage-events sibling. It does not explicitly name the sibling tradeoff, but the singular scope plus the GetEventDetailsRequest reference make the intent unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides prerequisites (eventTitle + baseExecutionScope required; comparisonExecutionScope optional for coverage columns) and routes the agent to get-truecoverage-events for timeWindow/platform rules. However, it never states when to prefer this tool over the sibling list tool or when the comparison scope is worth supplying, so usage is only implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-truecoverage-event-metadata-keysCInspect

List metadata keys for a given event title (ListEventMetadataKeysRequest). Pass eventTitle (CLI: --event-title or json-input).

ParametersJSON Schema
NameRequiredDescriptionDefault
eventTitleYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are supplied, so the description carries the full burden. The verb 'List' weakly implies a read-only operation, but nothing is said about whether keys are distinct, ordered, or empty when the event has no metadata. The parenthetical class name adds no behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with the operation stated first and the required parameter second. The trailing '(ListEventMetadataKeysRequest)' is mild internal-implementation noise rather than agent-facing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description is the only source of context, and it stops short: it does not describe what a 'metadata key' is, how the result is shaped, or the matching semantics of eventTitle. Adequate to call the tool, but not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It names eventTitle and its CLI/json-input forms, which is genuinely useful, but it never explains what value is expected (exact event title vs. ID, case sensitivity), leaving the semantic gap only partly filled.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (metadata keys) scoped to 'a given event title', which lets an agent distinguish it from the sibling get-truecoverage-session-metadata-keys (session-level). It does not explicitly name or route against that sibling, keeping it short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to use this tool versus the closely related get-truecoverage-session-metadata-keys, get-truecoverage-event-details, or get-truecoverage-event-transition. Usage is only inferable from the phrase 'for a given event title'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-truecoverage-eventsAInspect

TrueCoverage event funnel summaries (ListEventsRequest). baseExecutionScope is the real-user / primary environment; optional comparisonExecutionScope for coverage (set automationEmitsOnly:true on comparison for test-tagged emits only). Each scope needs environment + timeWindow: { relativeWindow: "604800s" } or { fixedWindow: { startTime, endTime } } (RFC 3339). Optional platform: web|ios|android or WEB_/IOS_/ANDROID_EXECUTION_PLATFORM. Prefer list-rum-environments first.

ParametersJSON Schema
NameRequiredDescriptionDefault
baseExecutionScopeYes
comparisonExecutionScopeNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses that baseExecutionScope is the real-user/primary environment and comparisonExecutionScope is for coverage, plus that automationEmitsOnly restricts to test-tagged emits — real behavioral context beyond the schema. It omits read-only/side-effect status, permissions, and output characteristics, so it is adequate but not complete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded and every clause carries information — scope semantics, timeWindow formats, platform enums, and the prerequisite. The single dense paragraph reads like packed notes rather than clean prose, but there is little waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with nested objects, no output schema, and no annotations, the description covers how to invoke it well but says nothing about what the funnel summary returns or how results are shaped. The calling contract is fairly complete; the return-side contract is not.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Top-level schema coverage is 0%, so the description must compensate, and it does: it explains the two-scope model, the required environment+timeWindow pairing, both timeWindow shapes with a concrete example ('604800s') and RFC 3339 fixed bounds, and the accepted platform value sets. It leaves metadataFilters, release, and branchName unaddressed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource and output shape ('TrueCoverage event funnel summaries'), so an agent knows this returns aggregated funnel data rather than raw events. It does not, however, distinguish itself from the several sibling tools (get-truecoverage-event-details, -time-series, -transition, -child-event-tree), leaving the agent to infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives one real usage cue — 'Prefer list-rum-environments first' — which tells the agent to resolve environments before calling. Beyond that there is no when-to-use/when-not guidance and no routing against the other truecoverage event tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-truecoverage-event-time-seriesBInspect

TrueCoverage daily time series for one metric (EventTimeSeriesRequest). Requires baseExecutionScope (environment + timeWindow). Optional eventTitle and metricType: SESSION_COUNT | RELATIVE_FREQUENCY | PERCENTAGE_TERMINAL_EVENT | SESSION_POSITION | TIME_TO_NEXT_EVENT | REVERSE_INDEX | TIME_FROM_START | TIME_TO_END | TIME_SINCE_PREVIOUS_EVENT.

ParametersJSON Schema
NameRequiredDescriptionDefault
eventTitleNo
metricTypeNo
baseExecutionScopeYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the burden; it discloses the aggregation granularity (daily) and that the result covers a single metric, which is genuinely useful behavioral context. It omits auth/permission needs, error behavior, and pagination, so it is not complete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then prerequisites and optional params in a compact second sentence. The long inline enum is somewhat bulky but maps to the parameter, so little is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a deeply nested schema with no output schema and 0% description coverage, the description covers the top-level requirement but leaves most nested fields (timeWindow shape, platform, metadataFilters) unexplained. Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It names the required baseExecutionScope and its required subfields and lists the optional metricType values, but it repeats the schema enum and gives no meaning for the enum members, timeWindow formats, platform, metadataFilters, or automationEmitsOnly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'TrueCoverage daily time series for one metric', which is more precise than the sibling names alone. It does not explicitly contrast with near-neighbors like get-truecoverage-event-transition or get-truecoverage-events, so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states the prerequisite ('Requires baseExecutionScope (environment + timeWindow)') and that only one metric is returned, which implies usage. But it gives no when-to-use-vs-alternative guidance against the many TrueCoverage siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-truecoverage-event-transitionBInspect

TrueCoverage detailed transition stats between two events (GetDetailedEventTransitionSummaryRequest). Requires eventTitle, nextEventTitle, baseScope (environment + timeWindow); optional coverageScope. metadataFilters on scopes are ignored. Uses baseScope/coverageScope field names.

ParametersJSON Schema
NameRequiredDescriptionDefault
baseScopeYes
eventTitleYes
coverageScopeNo
nextEventTitleYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses one non-obvious behavior (metadataFilters on scopes are ignored) and the field-naming convention, but says nothing about authentication, permissions, result shape, or what 'transition stats' contains. Partial coverage only.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Compact and front-loaded: the purpose leads, then requirements, then caveats. Sentences carry distinct facts, though the parenthetical request-type reference adds little for an agent.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a nested-object, four-parameter tool with no annotations and no output schema, the description covers the essentials (required vs optional, key caveat, naming). It remains thin on what the transition statistics represent and how the two scopes differ, so it is adequate but not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% at the top level, so the description must compensate. It maps the required/optional params and clarifies that baseScope includes environment + timeWindow and that the tool uses baseScope/coverageScope field names. It still leaves the semantics of eventTitle/nextEventTitle and the metadata-filter ignore rule's scope ambiguous.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'detailed transition stats between two events' for TrueCoverage. It is distinguishable from siblings like get-truecoverage-child-event-tree and get-truecoverage-event-details, though it does not explicitly name an alternative. Purpose is clear without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It discloses required (eventTitle, nextEventTitle, baseScope) and optional (coverageScope) parameters and warns that metadataFilters on scopes are ignored, which is genuine usage context. However, it offers no explicit when-to-use guidance vs. the other TrueCoverage event tools, leaving the choice implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-truecoverage-session-metadata-keysAInspect

List session-level metadata keys observed in RUM for TrueCoverage filters.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses that keys are 'observed' (empirically discovered from RUM traffic rather than a fixed enumerated schema) and that they are session-scoped, but says nothing about whether results are cached, how fresh they are, or whether the listed keys are environment/time-range dependent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with the scope qualifier front-loaded before the domain and use case. No filler, no redundancy, and the key discriminating word ('session-level') appears early.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only enumeration tool with no output schema, the description covers what the tool returns (session-level metadata key names) and its domain. The only material omission is any indication of how the returned keys relate to constructing filters or how they differ from event-level keys.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate; the baseline of 4 applies. The description does not need to explain any argument semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and a precisely scoped resource (session-level metadata keys observed in RUM for TrueCoverage filters). The 'session-level' qualifier implicitly separates it from the sibling get-truecoverage-event-metadata-keys, but the sibling is never named, so differentiation is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for TrueCoverage filters' hints at why the keys matter, but there is no explicit when-to-use, no prerequisites, and no direction to the event-level counterpart when a caller actually wants event metadata keys rather than session keys.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-user-storiesAInspect

Fetch user stories from the TestChimp platform by ordinal id (numeric part of US-). Returns full plan markdown content, title, and platform file path for each found story. Use when plan files are not yet synced to the repo.

ParametersJSON Schema
NameRequiredDescriptionDefault
userStoryOrdinalIdsYes

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose the return shape (full plan markdown content, title, platform file path), which is genuine behavioral value, but says nothing about auth requirements, rate limits, or behavior when an ordinal id is not found.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, both load-bearing: the first defines the operation and identifier format, the second defines the return payload and the usage condition. Nothing is padded and the identifier detail is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description supplies the essentials an agent needs: identifier format, payload contents, and when to prefer it. Missing only edge-case behavior such as not-found handling or partial results across multiple ids.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single parameter, so the description must compensate — and it does, explaining that userStoryOrdinalIds are the numeric part of US-<n> and that multiple stories may be returned. It omits that the input is an array requiring at least one element, but the core semantics are supplied.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Fetch) and resource (user stories) on a named platform, and uniquely clarifies the identifier format as 'ordinal id (numeric part of US-<n>)'. This cleanly separates it from create-user-story and update-user-story among the siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit triggering condition: 'Use when plan files are not yet synced to the repo.' That is a real when-to-use rule, but it does not name the alternative (e.g. get-plans-support-file or reading repo plan files) by name, so the routing is implied rather than spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get-workflow-executionBInspect

Get a workflow execution by id; pass includeActions=true for the action timeline.

ParametersJSON Schema
NameRequiredDescriptionDefault
includeActionsNo
workflowExecutionIdYes

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. Beyond the implicit read semantics of 'Get' and the includeActions toggle, it says nothing about error behavior for a missing id, permissions, or the shape of the result.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that identifies the resource and immediately notes the key optional parameter. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple get-by-id tool with no output schema, the definition covers the essentials but omits result shape, error handling, and any differentiation from the many other retrieval siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It explains includeActions ('for the action timeline') but says nothing about workflowExecutionId's format or constraints, leaving one of two parameters to be inferred from its name alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: retrieve a single workflow execution by id. This is clearly distinguishable from the sibling list-workflow-executions (singular fetch vs list), though it does not explicitly name any sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It implies usage (fetch one execution when you have its id) and gives a hint for the includeActions toggle, but offers no guidance on when to prefer this over siblings like get-execution-history or get-last-run-workflow-detail, and no exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-api-operation-interactionsAInspect

List recorded API operation interactions. Requires testId and/or operationId. Defaults to REAL interactions; limit is capped at 100.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
testIdNo
operationIdNo
interactionTypeNoREAL

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It usefully discloses the default filter (REAL), the limit cap, and the required filtering inputs, but says nothing about return shape, pagination, or ordering beyond the cap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with purpose, then requirements, then defaults/limits. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers the essentials for invocation (filters, default, cap), but with no output schema and no annotations, the description could do more to explain the result set, pagination behavior, or how REAL vs MOCKED interactions differ.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it does add real meaning: the 'testId and/or operationId' requirement is not expressed as required in the schema, and the REAL default clarifies the enum's behavior. The limit cap merely restates the schema's maximum.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'List recorded API operation interactions.' An agent can tell what it does, though it doesn't explicitly differentiate itself from siblings like list-api-operations or get-api-operation-detail.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides a key prerequisite ('Requires testId and/or operationId'), which is actionable context, but gives no when-to-use vs alternatives guidance or exclusions relative to the many sibling list/detail tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-api-operationsBInspect

List API operations for a service resource with covering-test previews, coverageSummary scores, observability mapping state, and runtimeObservation from the latest finalized production daily summary (with latest-hour fallback). runtimeObservation may include window bounds, requestCount/rpm, errorCount/errorRate, p50/p95/p99LatencyMs, status-class counts, and syncStatus. Missing observability is unknown/not configured, not zero. Use these signals to prioritize uncovered high-volume/high-error operations; prefer --root-file-path.

ParametersJSON Schema
NameRequiredDescriptionDefault
serviceKeyNo
rootFilePathNo
includeManualNo
includeRemovedNo

TDQS

B3.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations the description carries the full burden and does well: it discloses the production-daily-summary source with a latest-hour fallback, lists the runtimeObservation fields, and adds the important interpretation rule 'Missing observability is unknown/not configured, not zero.' It does not mention pagination, result limits, or permission/auth needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded, but the bulk is a single very long sentence stacking the runtimeObservation field inventory, which would normally live in an output schema. Given no output schema exists the detail is defensible, yet the nesting and enumeration make it harder to scan than it needs to be.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Rich on the return payload and observability semantics, but with 0% parameter coverage and no annotations the definition is incomplete: three of four inputs are undocumented and there is no guidance on limits, ordering, or when to prefer siblings. Adequate but with visible gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for 4 parameters, so the description must compensate but largely does not. It hints that rootFilePath is preferable but never defines it, and serviceKey, includeManual, and includeRemoved are not explained at all, leaving three of four inputs opaque.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List API operations for a service resource') and enumerates the payload it returns (coverage previews, coverageSummary, observability mapping, runtimeObservation). It does not differentiate itself from close siblings like list-api-operation-services, list-api-operation-interactions, or get-api-operation-detail, so an agent must infer the boundary from the service-resource scoping.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides some actionable direction: 'Use these signals to prioritize uncovered high-volume/high-error operations' and 'prefer --root-file-path'. However it never states when to choose this tool over get-api-operation-detail or list-api-operation-services, nor any when-not condition, so the routing guidance is partial.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-api-operation-servicesAInspect

List API operation service resources for the project (configured OpenAPI root file paths + operation counts). Use rootFilePath as the service resource id for list-api-operations / get-api-operation-detail.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose the returned content (root file paths and operation counts) and that the listing is scoped 'for the project', which is useful. However it says nothing about read-only semantics, pagination, result size, or how the project is resolved given there are no parameters.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences with no filler. The identification of what this tool returns comes first, and the routing instruction follows, so the most decision-relevant information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless list tool with no output schema, the description does the necessary compensating work by naming the fields returned and the role they play downstream. It is nearly complete; only pagination/result-size expectations and how the project scope is determined are left unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero input parameters, which sets the baseline at 4. The description still adds value by explaining that the returned rootFilePath field serves as the service resource id in other tools, linking output semantics to sibling calls.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (API operation service resources), and immediately disambiguates the jargon with a parenthetical defining what a service resource actually is: configured OpenAPI root file paths plus operation counts. It implicitly separates itself from list-api-operations and get-api-operation-detail, though that separation is framed as usage routing rather than as part of the purpose statement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit downstream routing: 'Use rootFilePath as the service resource id for list-api-operations / get-api-operation-detail.' This tells the agent this tool is the entry point that yields the identifier the sibling tools consume. It does not state when not to use it, but for a zero-param listing tool the workflow guidance is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-bunnyshell-environment-eventsCInspect

Troubleshooting: list BunnyShell platform events for an environment.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageNo
eventTypeNo
eventStatusNo
bnsEnvironmentIdYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. Beyond the word 'list', it says nothing about pagination (a page param exists), ordering, retention window, or required permissions for reading environment events. For a zero-annotation tool this is a substantial gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single short sentence with the troubleshooting framing front-loaded and no wasted words. It is efficient, though arguably under-specified rather than optimally concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and four parameters at 0% description coverage, the definition leaves an agent guessing about filtering, pagination, and the shape of returned events. It is too thin for a 4-parameter tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description only obliquely accounts for bnsEnvironmentId via 'for an environment'. The other three parameters (page, eventType, eventStatus) — including what event types/statuses are valid — are undocumented in both the schema and the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (list) and resource (BunnyShell platform events) scoped to an environment, which is enough to distinguish it from generic siblings like list-workflow-executions. It does not explicitly contrast itself with the nearest neighbors (list-bunnyshell-workflow-jobs, get-bunnyshell-workflow-job-logs), so it falls short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'Troubleshooting:' prefix implies a usage context — reach for this when diagnosing an environment — but gives no when-not guidance and names no alternative tool for related questions. Usage is only loosely implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-bunnyshell-workflow-jobsCInspect

Troubleshooting: list BunnyShell workflow jobs for an environment.

ParametersJSON Schema
NameRequiredDescriptionDefault
pageNo
bnsEnvironmentIdYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It implies a read-only listing via 'list', but it does not disclose pagination behavior despite a 'page' parameter, authentication/permission requirements, rate limits, or what the returned job records contain.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler. It is appropriately sized for a simple listing tool and contains no unnecessary repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a listing tool with no annotations, no output schema, and zero schema descriptions, the definition is too thin. It omits pagination behavior, return contents, and differentiation from the sibling log retrieval tool, leaving an agent with significant gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for two parameters. The phrase 'for an environment' loosely maps to the required bnsEnvironmentId, but the parameter is not named or explained, and the optional 'page' parameter is not mentioned at all, so the description does little to compensate for the schema gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'list BunnyShell workflow jobs', scoped 'for an environment'. This is clear and not tautological, but it does not explicitly distinguish itself from closely related siblings such as get-bunnyshell-workflow-job-logs or list-workflow-executions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'Troubleshooting:' prefix gives only a vague contextual cue. There is no guidance on when to use this tool rather than get-bunnyshell-workflow-job-logs, list-workflow-executions, or get-workflow-execution, and no prerequisites or exclusions are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-meeting-filter-optionsAInspect

List the labels, participants (key = user id or email, with display name), and participant email domains that appear on team-wide Meeting Bots meetings. Use the exact values as list-meetings filters (labels, participantKeys, participantDomains).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden. It adds useful scope ('team-wide Meeting Bots meetings') and clarifies the participant key is a user id or email with a display name, but says nothing about permissions, rate limits, pagination, or freshness of the enumerated values.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler. The return contents come first and the actionable directive (use as filters) is front-loaded immediately after.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly explains what is returned (labels, participants, domains) and how to consume it. It is nearly complete for a zero-param read tool, missing only the team-wide scoping caveats an agent might need.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the schema baseline of 4 applies. The description's only parameter-adjacent content is naming the downstream list-meetings filter fields, which is helpful framing but not parameter documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and three concrete resources (labels, participants with key format, participant email domains) scoped to team-wide Meeting Bots meetings. An agent can distinguish it from get-meeting-set and list-meetings without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly instructs the agent to use the returned values as list-meetings filters, naming the exact filter fields (labels, participantKeys, participantDomains). This gives clear usage context, though it does not state when NOT to use the tool or what happens if no team-wide meetings exist.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-meetingsAInspect

List / search cloud-synced Meeting Bots meetings, newest first, like the Meetings page filters + search. Team-wide meetings only (visibility: all team members). Filters: from / to (YYYY-MM-DD, ISO datetime, or epoch millis; inclusive), labels, participantKeys (user id or email), participantDomains, searchText (full-text over title + transcript; hits include a searchSnippet with matches wrapped in ⟦ ⟧). Returns meetingId, title, startMillis, labels, participants, summaryStatus; page with nextPageToken. Use list-meeting-filter-options for exact label / domain / participant values, then get-meeting-transcript (summaryOnly first) per hit.

ParametersJSON Schema
NameRequiredDescriptionDefault
toNo
fromNo
labelsNo
pageSizeNo
pageTokenNo
searchTextNo
endDateMillisNo
participantKeysNo
startDateMillisNo
participantDomainsNo

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden but does disclose key traits: read-only implication of 'list/search', sort order ('newest first'), scope restriction ('team-wide only'), inclusive date bounds, and the searchSnippet wrapping convention. It stops short of an explicit read-only/no-side-effects statement or any permission/rate-limit note.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense but front-loaded: purpose and scope come first, then filters, then return shape, then sibling routing. Every sentence carries information; it is packed rather than padded, though the single long middle sentence is heavy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema or annotations exist, yet the description supplies the return fields (meetingId, title, startMillis, labels, participants, summaryStatus), pagination via nextPageToken, and full filter inventory. Gaps remain around the duplicate startDateMillis/endDateMillis params and pageSize limits.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and there are 10 params, so the description must compensate. It adds real semantics for from/to formats ('YYYY-MM-DD, ISO datetime, or epoch millis; inclusive'), participantKeys ('user id or email'), participantDomains, and searchText ('full-text over title + transcript'). It leaves startDateMillis/endDateMillis, pageSize, and pageToken semantics unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('List / search cloud-synced Meeting Bots meetings'), its sort order ('newest first'), and its scope ('Team-wide meetings only, visibility: all team members'). It clearly distinguishes itself from siblings like list-meeting-filter-options and get-meeting-transcript.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent: 'Use list-meeting-filter-options for exact label / domain / participant values, then get-meeting-transcript (summaryOnly first) per hit.' This gives both a prerequisite and a follow-up alternative, naming the conditions that select them.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-perf-baselinesCInspect

List promoted performance baselines, optionally filtered by TestChimp id.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
offsetNo
testchimpIdNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it says almost nothing beyond 'List'. It does not disclose pagination semantics, default ordering, permissions, or what 'promoted' concretely entails, all of which matter for a baseline-listing tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with zero waste, and the resource is front-loaded before the optional filter. There is nothing extraneous to trim.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a paged list tool with three undocumented parameters, no annotations, and no output schema, the description is too thin. It omits pagination behavior and result shape, which the agent must otherwise infer.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, but it only addresses one of three parameters (testchimpId). The limit and offset parameters are left entirely unexplained, so an agent gets no help on result-sizing or pagination behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb+resource ('List ... performance baselines') and qualifies scope with 'promoted', so the agent knows these are the promoted subset. It does not differentiate from siblings like list-perf-runs or compare-perf-to-baseline, leaving the boundary between baseline listing and run listing implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The only guidance is 'optionally filtered by TestChimp id', which is a filtering hint rather than a when-to-use rule. There is no statement of when this tool should be chosen over list-perf-runs or compare-perf-to-baseline, and no prerequisites or exclusions are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-perf-runsCInspect

List performance runs, optionally filtered by TestChimp id, JOURNEY/COMPOSITE kind, branch, profile, dataset, LLM mode, or environment.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNo
limitNo
offsetNo
datasetNo
llmModeNo
profileNo
branchNameNo
environmentNo
testchimpIdNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, yet it discloses nothing about pagination behavior (limit/offset exist in the schema but the description never mentions paging, defaults, or maximum result count), sorting order, or required permissions. It only restates the filterable dimensions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that wastes no words and puts the resource before the filter list. It is appropriately sized for the amount of information conveyed.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter list tool with no annotations and no output schema, the description is too thin: it never explains pagination (limit/offset), result ordering, or what a returned run looks like, all of which an agent needs to page through results correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does map most parameter names to human-readable concepts (TestChimp id, kind, branch, profile, dataset, LLM mode, environment). However it omits limit/offset entirely and never explains the JOURNEY/COMPOSITE enum values or expected string formats, leaving two parameters and the only enum undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List performance runs') plus the filtering capability, which is enough to distinguish it from the singular sibling get-perf-run. It does not, however, explicitly differentiate itself from the other listing siblings such as list-perf-baselines or list-related-perf-tests.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The word 'optionally filtered' implies the tool supports both a full listing and a filtered query, but there is no guidance on when to reach for this tool versus get-perf-run (single run), list-perf-baselines, or compare-perf-to-baseline. No prerequisites, exclusions, or alternative routing are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-policiesAInspect

List policy files for an optional workflow-id. Marks isDefault when filename is .policy.md.

ParametersJSON Schema
NameRequiredDescriptionDefault
workflowIdNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it does disclose one meaningful trait: the isDefault marker is derived from the filename matching '<workflow-id>.policy.md'. However, it says nothing about permissions, pagination, ordering, or the shape of the returned list.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core action and followed by the one non-obvious behavioral rule. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-only list tool with one optional parameter and no output schema, the description covers the essential action and the isDefault derivation rule. Remaining gaps (ordering, scope when workflowId is omitted) are minor given the tool's low complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for the single workflowId parameter, so the description must compensate. It does establish that the parameter is optional and ties it to the file naming convention, but adds no format, matching, or default-behavior detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (policy files) plus scope (optional workflow-id), which distinguishes it from get-policy and upsert-policy siblings. It stops short of explicitly naming those siblings or contrasting behavior, so it is clear but not fully differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'for an optional workflow-id' implies you may call it with or without the filter, but it never states when to use this versus get-policy or upsert-policy, nor what omitting the parameter returns. Usage is inferable rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-rum-environmentsAInspect

List distinct RUM environment tags for this project. Call this first to choose environment values for TrueCoverage ExecutionScope.environment (e.g. QA, production).

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses that results are distinct tags and are meant to feed ExecutionScope.environment, but says nothing about scoping to a project, permissions, result ordering, or whether the list is exhaustive. Adequate for a trivial read, thin on behavioral detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero waste, and the primary purpose is front-loaded ahead of the usage hint. Nothing is repeated from the schema or title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter read tool with no output schema, the description covers what it does and why an agent would call it. The only real gap is the shape of the return value (a flat list of tag strings), which is implied but not stated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing to document and the baseline is 4. The description's note that 'this project' is the implicit scope is a small useful addition beyond the empty schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'List distinct RUM environment tags for this project.' An agent knows exactly what the tool returns, and it is clearly distinct from any sibling in the list (many of which retrieve TrueCoverage events/sessions rather than environment tag values).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Call this first to choose environment values for TrueCoverage ExecutionScope.environment,' giving clear sequencing and downstream purpose. It does not name an alternative or a when-not condition, but no plausible sibling alternative exists for this narrow lookup.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-screen-statesAInspect

Fetch the project's screen/state vocabulary (relational atlas) for SmartTests and traces. Optional environment field is accepted for forward compatibility; v1 may be project-global.

ParametersJSON Schema
NameRequiredDescriptionDefault
environmentNo

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does add real behavioral context by disclosing that the environment field is accepted 'for forward compatibility' and that v1 may be project-global, warning the agent that the filter may be silently ignored. It stops short of describing the return shape, ordering, or whether the call is a pure read.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the primary purpose front-loaded and the parameter caveat second. Nothing is padded, though the parenthetical jargon is slightly more decorative than necessary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description is the sole source of truth, and it omits what the returned vocabulary looks like (screens, states, their relations) and how the result is typically consumed. For a simple zero-required-parameter list call this is adequate but leaves the agent guessing about the response.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and there is only one parameter, so the description must explain it — and it does, stating the field is optional and may have no effect in v1. That is exactly the caveat an agent needs before passing a value and expecting filtering. It could still say what values are valid, but the semantic intent is covered.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Fetch') and resource ('the project's screen/state vocabulary'), with a parenthetical gloss ('relational atlas') that clarifies what the vocabulary actually is. It is distinguishable from the sibling upsert-screen-states by the read verb, though it never names that sibling or any other alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'for SmartTests and traces' hints at the downstream consumers and therefore implies when the data is needed, but there is no explicit when-to-use, when-not-to-use, or named alternative. The agent must infer the trigger condition from the audience description alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-semantic-nearbyBInspect

List semantically nearby entities across types (Story/Scenario/Test/Issue/Event). TEST uses TestLocator; STORY/SCENARIO/ISSUE use sourceOrdinalId; EVENT uses sourceEventTitle. Response TEST hits include TestLocator (never platform test_id).

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
sourceTestNo
sourceOrdinalIdNo
sourceEntityTypeYes
sourceEventTitleNo
targetEntityTypesNo

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses an output trait ('Response TEST hits include TestLocator (never platform test_id)'), which is real behavioral value. However, it says nothing about permissions, rate limits, pagination, or how many default results are returned, leaving meaningful behavioral gaps for a read tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, front-loaded with the core purpose, then the parameter routing and output caveat. No filler, though the parenthetical entity list partly duplicates the enum already present in the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema and 0% param coverage mean the description should carry more. It covers entity-type-to-parameter routing well but omits the meaning of 'limit' and 'targetEntityTypes' and any notion of the response shape beyond the TEST-locator note, so it is adequate but not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it does the most important part: mapping each sourceEntityType to its required input (TEST→TestLocator, STORY/SCENARIO/ISSUE→sourceOrdinalId, EVENT→sourceEventTitle). It still leaves 'limit' and 'targetEntityTypes' semantics unexplained, but this is strong compensation for the identifier parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('List semantically nearby entities') and scopes it to five entity types. The 'across types' framing implicitly distinguishes it from the similarly named sibling list-semantic-similar-tests, though that contrast is never made explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description never says when to reach for this tool versus list-semantic-similar-tests or any other list/get sibling, nor does it state any precondition. It only tells the agent how to call it, not when.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-semantic-similar-testsBInspect

List semantically similar SmartTest pairs in scope using TestLocators (no test_id). Pairs are deduped (A→B only when A.testId < B.testId). Distinct-marked pairs are excluded.

ParametersJSON Schema
NameRequiredDescriptionDefault
scopeNo

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose genuinely useful behavior: dedup rule (A→B only when A.testId < B.testId), exclusion of distinct-marked pairs, and scope-wide rather than per-test operation. It says nothing about permissions, result size, pagination, or the shape of a returned pair, so the behavioral picture is partial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short, dense sentences with the core verb and scope up front and the dedup/exclusion rules following. Every sentence carries information, though the parenthetical 'using TestLocators (no test_id)' is cryptic and slightly interrupts the flow.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations, no output schema, and a nested parameter whose fields are undocumented, the description should say more about the returned pair structure and how scope resolves. It covers the dedup and exclusion semantics well but leaves the parameter and output side thin.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single scope parameter is a nested object with filePaths and folderPath fields. The description only says 'in scope' and never explains what those fields mean or which one wins, so it does not compensate for the documentation gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (semantically similar SmartTest pairs) with the distinguishing detail that it operates across a scope rather than on a single test_id. That phrasing implicitly separates it from list-semantic-nearby and mark-semantic-tests-distinct, though it never names a sibling outright.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use or when-not-to-use guidance, and no alternative tool is named. The closest thing to guidance is the implicit 'in scope' framing and the note that distinct-marked pairs are excluded, but the agent is left to infer when this list is preferable to list-semantic-nearby or list-related-perf-tests.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-tests-awaiting-verificationAInspect

List SmartTests whose latest executions still need human verification before they earn a verified badge (testId, testName, filePath, scenarioId / scenarioOrdinalId / scenarioTitle, verificationStatus NOT_VERIFIED or VERIFIED_STALE, reportedAtMillis, workflowExecutionId). Optional userId narrows to tests authored for that user (defaults to the OAuth token's user); optional limit. Use after writing E2E tests to prompt the user to verify the executions.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
userIdNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose real behavior: the scoping rule ('latest executions'), the two status values that qualify, and the userId default (OAuth token's user). It does not state read-only status explicitly or mention sorting/pagination limits, but 'List' plus the enumerated semantics give solid behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and triggering condition are front-loaded, and the tool is otherwise compact. The long parenthetical field list is dense but earns its place as a return-shape substitute given there is no output schema; the sentence is a bit run-on but not wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations and no output schema, the description usefully enumerates returned fields and the filter semantics, so an agent knows what it will get and when. It lacks any pagination or ordering note for a capped list tool, which is the main remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It fully explains userId (narrows to tests authored for that user, defaults to the OAuth token's user), but limit is given only as 'optional limit' with no mention of its 1-500 range or behavior. Half the parameters are enriched, half are not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (SmartTests) plus a precise scope: those whose latest executions still need human verification. The parenthetical enumerates the exact verificationStatus values (NOT_VERIFIED or VERIFIED_STALE), which disambiguates it from sibling list-* and mark-tests-for-review tools without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly tells the agent when to call it: 'Use after writing E2E tests to prompt the user to verify the executions.' That is a genuine triggering condition, though it names no alternatives (e.g. mark-tests-for-review) and states no exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-test-scenarios-for-scopeAInspect

List in-scope test scenarios (ordinalId + title only) for exactly one locator: a named test run, a release label, or a platform plans folder/file path. Does not return markdown — use get-test-scenarios for detail by known ordinal. Use when executing tests for a plans path, release, or named test run.

ParametersJSON Schema
NameRequiredDescriptionDefault
releaseNo
plansPathNo
namedTestRunIdNo

TDQS

A4.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It usefully discloses the return shape (ordinalId + title only, no markdown), which is a real behavioral trait, but says nothing about pagination, invalid/ambiguous locator handling, or authentication requirements for a tool where that context matters.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the verb and scope, then the return shape, then routing and usage. It is efficient overall, though the final sentence largely repeats the locator types already listed in sentence one, adding mild redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly takes on the return-value explanation (ordinalId + title only, not markdown) and covers all three otherwise-undocumented parameters. It stops short of describing error/pagination behavior, but nothing essential for a correct call is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and none of the three parameters are described in the schema, so the description must compensate. It does: it maps the enumerated locators (named test run, release label, plans folder/file path) to the release/plansPath/namedTestRunId semantics and adds the mutual-exclusion rule ('exactly one locator').

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (List) and resource (in-scope test scenarios) plus the exact return shape (ordinalId + title only) and the scoping constraint (exactly one locator). It also names the sibling get-test-scenarios as the not-this path, so an agent can distinguish the two without opening either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use it ('when executing tests for a plans path, release, or named test run') and routes the agent to the alternative ('use get-test-scenarios for detail by known ordinal'). The 'exactly one locator' constraint also implies the not-when (multiple locators or detail lookups).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-workflow-catalogBInspect

List supported TestChimp workflows with Active / Disabled / Missing Config status for the project.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full load. It usefully discloses the status vocabulary it returns (Active / Disabled / Missing Config), which is real behavioral context, but says nothing about auth/project scoping, ordering, or whether results are paginated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; the resource and the returned status dimension come first. Nothing is wasted, though it is terse to the point of leaving the scope unexplained.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-param, read-only catalog list with no annotations and no output schema, the description gives the gist (workflows + status) but omits return shape, project scoping, and whether the project is inferred from context. Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so the schema-based baseline of 4 applies and there is no parameter semantics for the description to clarify.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and resource ('supported TestChimp workflows') plus the status dimension it reports (Active/Disabled/Missing Config). It is distinguishable from the many get-* workflow tools, though it never names list-workflow-executions or explains the catalog-vs-execution boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no prerequisites, and no mention of the closest sibling (list-workflow-executions). The agent must infer that this is a one-time discovery call rather than a status query over runs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

list-workflow-executionsBInspect

List recent workflow executions for the project, optionally filtered by workflowId.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNo
offsetNo
workflowIdNo

TDQS

B3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. 'Recent' is undefined (how many, over what window), the default page size is never stated, and nothing explains that limit/offset control pagination or what the ordering guarantees are.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with zero filler. The scoping ('for the project') and the filter are placed where they matter, and nothing is repeated from the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and 0% schema coverage, the description should compensate but does not. It omits default limits, pagination semantics, ordering, and the shape of returned executions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across three parameters, so the schema contributes no semantics. The description explains only workflowId; limit and offset (with their bounds) are left entirely undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('List') and resource ('workflow executions') with scope ('for the project') and the optional filter. It reads clearly against the singular sibling get-workflow-execution, though it never explicitly contrasts with it.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'optionally filtered by workflowId' implies when the filter applies, but there is no guidance on when to choose this list tool over get-workflow-execution, get-execution-history, or get-last-run-workflow-detail.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mark-entity-distinctBInspect

Mark two same-type entities as distinct. TEST uses TestLocators; STORY/SCENARIO/ISSUE use ordinals; EVENT uses titles.

ParametersJSON Schema
NameRequiredDescriptionDefault
focusTestNo
otherTestNo
entityTypeYes
focusOrdinalIdNo
otherOrdinalIdNo
focusEventTitleNo
otherEventTitleNo

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden, yet it discloses nothing about effects: what 'distinct' actually changes, whether it is reversible, what permissions are required, or what happens to previously marked pairs. For a mutation tool with zero annotation coverage this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the core action and then the per-type parameter routing. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a nested, 7-parameter, annotation-free mutation tool the description covers the essential parameter routing but omits behavioral context: reversibility, effect of the operation, and error conditions. Adequate for invocation, incomplete for understanding consequences.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema description coverage on 7 parameters, the description does the heavy lifting by explaining the identifier scheme per entityType (TestLocators for TEST, ordinals for STORY/SCENARIO/ISSUE, titles for EVENT), which maps directly onto the focus/other param families. It doesn't clarify narrower details like the string-vs-number ordinal union, but this is substantial compensation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a clear verb and resource ('Mark two same-type entities as distinct'), and the word 'same-type' is a meaningful scoping constraint. However, it doesn't differentiate itself from closely related siblings such as unmark-entity-distinct or mark-semantic-tests-distinct, so an agent must infer the boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by mapping each entityType to its identifier scheme ('TEST uses TestLocators; STORY/SCENARIO/ISSUE use ordinals; EVENT uses titles'), which is real guidance. But it never states when to reach for this tool versus unmark-entity-distinct or mark-semantic-tests-distinct, and gives no prerequisites.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mark-plan-items-implementation-doneBInspect

Mark user stories and/or test scenarios implementation-complete in platform lifecycle (DB only; does not rewrite plan markdown). Use scenarioOrdinalIds / userStoryOrdinalIds (numeric parts of TS- / US-).

ParametersJSON Schema
NameRequiredDescriptionDefault
scenarioOrdinalIdsNo
userStoryOrdinalIdsNo

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden; it usefully discloses that the change is DB-only and does NOT rewrite plan markdown, which is real side-effect information. However it omits permissions, idempotency/reversibility, and what happens when ordinals don't exist.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the core action and the DB-only caveat front-loaded, then the id-format hint. No filler, though the parenthetical could be slightly tighter.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a small two-param mutation tool with no annotations and no output schema, the description covers what changes, the DB-only scope, and id semantics. It still leaves gaps around auth requirements, validation of ordinals, and the resulting state, so it is adequate but not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it does partially by explaining that scenarioOrdinalIds/userStoryOrdinalIds are the numeric parts of TS-<n>/US-<n>. It still omits that both are arrays, optional, and that at least one is presumably needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (mark user stories/test scenarios implementation-complete) and even names the boundary 'platform lifecycle (DB only)'. It is clear about what is affected, though it never names the closest sibling update-plan-items-lifecycle-status to differentiate the two.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when this applies (implementation-complete in the platform lifecycle) and gives id-format guidance, but offers no explicit when-to-use/when-not versus update-plan-items-lifecycle-status or update-test-scenario. Usage is inferable rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mark-semantic-tests-distinctCInspect

Mark two SmartTests as legitimately distinct (symmetric) using TestLocators. Agent/API calls use marked_by_user_id = 0.

ParametersJSON Schema
NameRequiredDescriptionDefault
focusTestYes
distinctTestYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does add two behavioral facts: the marking is symmetric, and Agent/API calls are attributed with marked_by_user_id = 0. However, it omits whether this is a mutation with side effects, whether it is reversible (the unmark-entity-distinct sibling implies so), idempotency, and permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the primary action front-loaded and no filler. The trailing sentence about marked_by_user_id = 0 is cryptically phrased but relevant rather than redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A mutation tool with zero annotations, no output schema, nested-object parameters, and no schema description coverage needs more than two short sentences. Missing are the TestLocator structure, the effect of marking, error/edge cases, and permission context, so an agent lacks what it needs to call this confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and both parameters are nested objects with four sub-fields (fileName, testName, testSuite, folderPath) that are entirely undocumented. The description gestures at the input shape with 'TestLocators' and distinguishes the two roles ('two... tests', 'distinct'), but does not explain the sub-field semantics or the focus vs distinct distinction, leaving most parameter meaning unstated.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (mark), the exact resource (two SmartTests), and the resulting state ('legitimately distinct'), plus the input mechanism (TestLocators). An agent can distinguish this from unmark-entity-distinct and mark-entity-distinct, though the description never explicitly differentiates from those siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use, when-not-to-use, or alternative-selection guidance. The word 'symmetric' hints at behavior but does not explain when an agent should call this versus mark-entity-distinct, nor what preconditions exist (e.g., tests previously flagged as near-duplicates).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mark-tests-for-reviewAInspect

Report existing SmartTests that an agent patched so humans can re-verify. Always send per-test confidence 0–100 (higher = less need for human review). Do not read project config. Call only from fix-test-execution after test-incorrect patches (never from run-qa / create-tests; never for product-broken cases). Optional agentTraceability / workflowExecutionId / gitCommitSha / branchName.

ParametersJSON Schema
NameRequiredDescriptionDefault
testsYes
gitShaNo
userIdNo
actorTypeNo
agentModelNo
branchNameNo
cliVersionNo
policyFileNo
workflowIdNo
gitCommitShaNo
skillVersionNo
policyVersionNo
agentTraceabilityNo
workflowExecutionIdNo

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses calling constraints and the effect ('so humans can re-verify') plus the confidence scale semantics, but says nothing about permissions, reversibility, what the review record looks like, or failure behavior. Useful context, but short of the full disclosure a mutation tool with zero annotation coverage requires.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose, then the hard calling constraints, then optional fields. Every sentence carries information, though 'Do not read project config' is cryptic and slightly interrupts the flow.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 14-parameter tool with a nested required object, no output schema and no annotations, the description covers purpose and routing well but is incomplete on the many optional traceability parameters and on what the call produces or returns.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 14 parameters, so the description must compensate. It adds real meaning for the required 'confidence' field (0-100, higher = less need for human review) and lists four optional traceability fields, but leaves roughly ten parameters (gitSha, userId, actorType, agentModel, workflowId, etc.) unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: reporting existing SmartTests that an agent patched so humans can re-verify. This is distinct from sibling readers like list-tests-awaiting-verification and from mark-semantic-tests-distinct/mark-entity-distinct, so an agent can route correctly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly names the only valid caller context (fix-test-execution after test-incorrect patches) and explicitly excludes others (never from run-qa / create-tests; never for product-broken cases). This is a textbook when/when-not/alternatives statement.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

promote-perf-baselineCInspect

Promote a performance run as the baseline for an environment class. Optional agent traceability records the mutation.

ParametersJSON Schema
NameRequiredDescriptionDefault
runIdYes
gitShaNo
userIdNo
envClassYes
actorTypeNo
agentModelNo
branchNameNo
cliVersionNo
policyFileNo
workflowIdNo
skillVersionNo
policyVersionNo
agentTraceabilityNo
workflowExecutionIdNo

TDQS

C2.7/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses that the operation is a mutation that can record optional agent traceability, but says nothing about whether an existing baseline is overwritten, whether the action is reversible, or what permissions are required.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core action and effect, with no wasted words. It is well sized for what it chooses to say.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 14-parameter mutation with a nested traceability object, no annotations, and no output schema, this description is far too thin. It omits baseline-replacement semantics and essentially all parameter meaning, so an agent could invoke it but not confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 14 parameters and 0% schema description coverage, the description must compensate and does not. It only alludes to the traceability fields ('agent traceability'), leaving runId, envClass, and the ten-plus traceability parameters undocumented beyond their names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (promote) plus resource (a performance run) and its effect (becomes the baseline for an environment class), which is enough to separate it from siblings like compare-perf-to-baseline, get-perf-run, and list-perf-baselines. No sibling is named explicitly, so it stops short of 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use or when-not-to-use guidance and no mention of alternatives such as compare-perf-to-baseline. The purpose implies its use case but nothing tells the agent under what conditions selecting a baseline is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

provision-ephemeral-environmentAInspect

Create a BunnyShell ephemeral environment (create + deploy trigger only). Prefer provision-ephemeral-environment-and-wait for normal flows.

ParametersJSON Schema
NameRequiredDescriptionDefault
branchNameNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does disclose one meaningful trait – that this only triggers creation + deploy and does not wait for readiness – but omits auth/permission requirements, side effects, and what state is left behind after the call.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two front-loaded sentences with no filler; the scope qualifier and the sibling routing are both stated in minimal space. The parenthetical is terse but effective rather than wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations and no output schema, the description is thin: no auth needs, no description of the resulting environment or how to observe its status, and the lone parameter is undocumented. An agent has enough to route but not enough to call it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single parameter branchName has 0% schema description coverage, so the description must compensate and does not. It never states the format (branch ref? name?), whether it defaults when omitted (it is not required), or how it maps to the created environment.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Create) and resource (BunnyShell ephemeral environment), and the parenthetical '(create + deploy trigger only)' scopes exactly what the operation does. It explicitly distinguishes itself from the sibling provision-ephemeral-environment-and-wait, so an agent can route correctly without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Names the alternative tool and the condition that selects it ('Prefer provision-ephemeral-environment-and-wait for normal flows'), implying this one is for flows where waiting is not wanted. Clear routing context, though it never states explicitly when this tool IS the right choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

provision-ephemeral-environment-and-waitBInspect

Provision a BunnyShell ephemeral environment, then poll until deployed and component URLs are available. Progress is logged to stderr (CLI) or MCP logging when supported.

ParametersJSON Schema
NameRequiredDescriptionDefault
branchNameNo
maxWaitMinutesNo
pollIntervalSecondsNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden; it does disclose two real behaviors an agent needs: the call blocks and polls, and progress goes to stderr/MCP logging. It omits timeout behavior, failure handling, default wait/poll values, and whether partial deploys are left behind — significant gaps for a long-running provisioning operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, verb-first, zero padding; the core action and its blocking nature are front-loaded. The logging sentence is secondary but still informative rather than filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

It does tell the agent what the call yields (deployed environment plus component URLs), which is useful given there is no output schema. But with no annotations and no parameter documentation, the definition is thin on timeout semantics, defaults, and error behavior for an operation that can block for minutes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and all three parameters (branchName, maxWaitMinutes, pollIntervalSeconds) are bare types. The description gestures at polling but never explains the wait/poll knobs, their defaults, or what branchName selects, so the schema gap is not compensated.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Provision'), the target resource ('BunnyShell ephemeral environment'), and adds the blocking behavior ('poll until deployed'). This is clearer than most, but it never contrasts itself with the sibling 'provision-ephemeral-environment', which the '-and-wait' suffix implicitly duplicates.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The polling/waiting behavior implies the usage context (you need the environment ready before the next step), but there is no explicit when-to-use statement, no mention of the non-waiting sibling as an alternative, and no guidance on combining it with 'get-ephemeral-environment-status' or 'destroy-ephemeral-environment'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

register-bot-profileAInspect

Register or replace this QA bot's profile and event subscriptions atomically (mutating — confirm with the user first). role: QA_LEAD | PM | QA_ENGINEER | DEVELOPER. responsibilities: the user's own words. capabilities: REQUIREMENTS_UPDATE, E2E_AUTHORING, ISSUE_FIX, MANUAL_TEST_COORDINATION, TEST_BATCH_FIX, QA_POSTURE. subscriptions: [{eventType, filters:[{field, op:'eq', value}]}] where value may be 'me' (e.g. git-push author=me, issue-assigned assignee=me, e2e-batch-completed). Derive subscriptions from capabilities per the testchimp skill's bot onboarding guide. botId defaults to the bot-id header / OAuth token.

ParametersJSON Schema
NameRequiredDescriptionDefault
roleYes
botIdNo
capabilitiesNo
subscriptionsNo
responsibilitiesNo

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses mutating nature, atomic semantics, and the 'replace' implication, plus the botId default source (header/OAuth token). It does not cover permissions/auth requirements, error behavior, or the consequences of replacing existing subscriptions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core action and its mutating/atomic warning, then uses compact colon-separated field lists. Dense but virtually every clause carries information; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 5-parameter mutation tool with no annotations and no output schema, the description covers the essentials: atomic replace, the confirm-first requirement, per-parameter semantics, and the subscription derivation source. Missing only return/error behavior and explicit permission requirements.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it largely does: it enumerates role and capability enum values, defines responsibilities ('the user's own words'), documents the nested subscriptions shape with op='eq', the 'me' value semantics, and botId defaulting. This adds substantial meaning beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource: 'Register or replace this QA bot's profile and event subscriptions atomically'. The scope (profile + event subscriptions, atomic replace) is clear and distinguishable from read siblings like get-bot-profile and ack-bot-events, though no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides procedural guidance ('confirm with the user first', 'derive subscriptions from capabilities per the testchimp skill's bot onboarding guide'), which implies this is a setup/onboarding step. However, it never states when to use this vs get-bot-profile or when not to call it, so routing must be inferred.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

report-agent-actionAInspect

Report a mutating agent action under a stable workflow-execution-id (ULID). First call for an id creates the workflow_executions row; later calls append Activity timeline rows (AGENT_WORKFLOW_ACTIVITY). Actions land on the entity's Activity timeline (plans, issues, SmartTest file). entityType: USER_STORY | SCENARIO | SMART_TEST | POLICY | ISSUE | TEST_EXECUTION | TEST_INVOCATION_BATCH | EXPLORATION | EVENT | WORKFLOW. actionType: CREATED | UPDATED | DELETED | ANALYZED | IMPLEMENTED | ACTION_COMPLETED | ACTION_FAILED. Identity: SMART_TEST uses test (TestLocator: folderPath/fileName/testSuite/testName); other artifact types use entityIdentity (ordinal / filename / opaque id). Do not use platform UUIDs. Completion (ACTION_COMPLETED / ACTION_FAILED): entityType WORKFLOW and entityIdentity = catalog workflow_id.

ParametersJSON Schema
NameRequiredDescriptionDefault
testNo
gitShaNo
userIdNo
actorTypeNo
actionTypeYes
agentModelNo
branchNameNo
cliVersionNo
entityTypeYes
policyFileNo
workflowIdYes
skillVersionNo
traceabilityNo
policyVersionNo
entityIdentityNo
workflowExecutionIdYes

TDQS

A3.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does well: it states that the tool is mutating, that the first call for an id creates the workflow_executions row, that later calls append Activity timeline rows, and that actions land on the entity's Activity timeline. It does not describe permissions, error behavior, or rate limits, but the core side effects are disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core action and then moves into identity and completion details. It is dense but generally earns its length given the 16-parameter schema, though it repeats enum values already present in the schema and could trim some of that duplication.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex, mutating tool with no annotations, no output schema, and 0% schema description coverage, the description supplies essential context about identity and lifecycle behavior. It remains incomplete regarding optional parameters, permissions, and return/error behavior, so an agent can invoke it but not fully understand all invocation scenarios.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It explains key semantics for entityType, actionType, SMART_TEST identity via the nested test object, other entity identities via entityIdentity, and completion using WORKFLOW plus catalog workflow_id. It leaves many optional traceability parameters such as gitSha, userId, actorType, agentModel, branchName, cliVersion, policyFile, and skillVersion undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb and resource: 'Report a mutating agent action under a stable workflow-execution-id (ULID).' It also explains the create-vs-append behavior, which makes the tool's scope clear. However, it does not explicitly distinguish this tool from the many sibling reporting tools such as report-sast-findings or report-deps-findings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied through 'Report a mutating agent action' and the note that the first call creates the row while later calls append timeline rows. The description does not state when to choose this tool over alternatives, nor does it give explicit exclusion conditions. The 'Do not use platform UUIDs' line is an identity constraint, not a usage guideline.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

report-dast-findingsAInspect

Upload a ZAP Traditional JSON report for a security scan. Pass --id and --report-file . Backend parses alerts, dedupes by bug hash, and inserts new SECURITY bugs linked to the scan. Does not mark the scan COMPLETED — the DAST playbook calls update-scan-progress COMPLETED after this.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
reportFileYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations the description carries the full burden, and it delivers: it discloses parsing, dedupe by bug hash, insertion of new SECURITY bugs linked to the scan, and the critical caveat that the scan is not marked COMPLETED. It omits auth/permission needs and error or overwrite behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences, front-loaded with the action and the input, then behavior, then the sequencing caveat. Nothing is redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers purpose, inputs, side effects, and the handoff to update-scan-progress for a mutation tool with no annotations and no output schema. The only real gap is that with no output schema it never hints at what the call returns (e.g. count of inserted bugs).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate, and it partially does by implying id is the scan and reportFile is a path to a ZAP Traditional JSON report. However it uses CLI flag names (--id, --report-file) that do not match the schema property names (id, reportFile), and gives no format/constraint detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Names a specific verb (upload), resource (ZAP Traditional JSON report) and the domain (security scan), which cleanly separates it from the sibling report-sast-findings/report-secrets-findings/report-deps-findings. It also states the downstream effect on SECURITY bugs and scan status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states it does NOT complete the scan and names update-scan-progress COMPLETED as the follow-up call, which is strong sequencing guidance. It does not explicitly say when to pick this over the SAST/secrets/deps siblings, but the ZAP/DAST framing makes the selection obvious.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

report-deps-findingsAInspect

Upload a full Trivy JSON report for a dependency security scan. Pass --id and --report-file . Backend stores the report, filters by security profile / ignore-unfixed, dedupes, and inserts SECURITY bugs. Does not mark the scan COMPLETED — the deps playbook calls update-scan-progress COMPLETED after this.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
reportFileYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations to lean on, the description carries the burden well: it discloses the backend pipeline (stores, filters by security profile / ignore-unfixed, dedupes, inserts SECURITY bugs) and an important side-effect boundary (does not mark the scan COMPLETED). It omits auth/permission requirements and whether re-uploading a report is idempotent despite dedupe.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four tight sentences, front-loaded with the core action and invocation, followed by pipeline behavior and the critical 'does not mark COMPLETED' caveat. No filler or restatement of the name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a write tool with no annotations and no output schema, the description covers purpose, mutation effects, and the downstream workflow step. It leaves the return value and the meaning of 'id' unstated, which is a minor gap since no output schema exists to cover responses.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and both parameters (id, reportFile) are required, so the description must compensate. It maps the params to '--id' and '--report-file <path>' and implies the file is a filesystem path, but does not explain what the id identifies (scan id?) or any format/collision constraints. Baseline-level adequacy given the 0% coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Upload a full Trivy JSON report for a dependency security scan') that distinguishes it cleanly from the sibling report-sast-findings, report-dast-findings, and report-secrets-findings tools. An agent can tell which scanner's findings this handles without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete invocation guidance ('Pass --id and --report-file <path>') and a clear workflow boundary: it does not mark the scan COMPLETED, and update-scan-progress must be called afterward by the deps playbook. It stops short of naming sibling alternatives or stating when a different reporting tool is appropriate, so it is short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

report-requirement-quality-findingsAInspect

Upload a DeFOSPAM / requirement quality analysis report for a user story or test scenario (local-agent path). Pass full RequirementQualityReport JSON via --report-file or --json-input {"report":{...}}. report.subject.subjectEntityId is required; use --subject-type + --ordinal-id to resolve via get-requirement-quality-report, or set subjectEntityId explicitly. Backend merges IGNORED/APPLIED findings carry-forward on re-report.

ParametersJSON Schema
NameRequiredDescriptionDefault
reportNo
ordinalIdNo
reportFileNo
subjectTypeNo
subjectEntityIdNo

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full behavioral burden. It usefully discloses that the backend merges IGNORED/APPLIED findings carry-forward on re-report and that subjectEntityId is required, but it says nothing about permissions, side effects, or response behavior, leaving significant gaps for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose is front-loaded, then invocation mechanics follow. It is dense but every sentence adds information; the json-input snippet and CLI-style flag references make it slightly run-on but not wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with five params, deeply nested objects, no annotations, and no output schema, the description is only partially complete. It covers identity resolution and re-report merge behavior but omits auth requirements, side effects, and any sense of the return/confirmation, which an agent would need.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It clarifies that subjectEntityId is required and explains the reportFile / json-input alternatives and the subjectType+ordinalId resolution path, which covers several of the five params. However, the large nested report object is only delegated to an external 'RequirementQualityReport' type with no field-level guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Upload) and resource (DeFOSPAM / requirement quality analysis report), plus scope (for a user story or test scenario, local-agent path). It also implicitly distinguishes itself from the read-side sibling get-requirement-quality-report, so an agent can route correctly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context: this is the local-agent path for uploading, and it names the alternative resolution route via get-requirement-quality-report when only subject-type + ordinal-id are known. It does not state when NOT to use this tool (e.g., the cloud path), so it falls short of a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

report-sast-findingsAInspect

Upload a full Semgrep CLI JSON report for a SAST security scan. Pass --id and --report-file . Backend stores the raw report, parses results, dedupes by bug hash, and inserts new SECURITY bugs. Does not mark the scan COMPLETED — the SAST playbook calls update-scan-progress COMPLETED after this.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
reportFileYes

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it discloses server-side storage of the raw report, parsing, dedupe-by-bug-hash, insertion of new SECURITY bugs, and the critical state boundary that it does NOT mark the scan COMPLETED. This side-effect and idempotency detail is exactly what an agent needs for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the primary action, then parameters, then backend behavior and the completion caveat. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-param mutation tool with no output schema, the description covers behavior, dedupe semantics, and the workflow boundary thoroughly. It could still note auth/permission requirements or what the call returns, but nothing essential to correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It maps both params to CLI flags (--id and --report-file <path>) and clarifies reportFile is a path, but the meaning of id (the scan identifier) is left to inference from context rather than stated.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (upload), resource (Semgrep CLI JSON report), and scan type (SAST security scan). The 'SAST' qualifier cleanly distinguishes it from the sibling report-dast-findings, report-deps-findings, and report-secrets-findings tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides strong workflow context by naming the downstream step: the SAST playbook calls update-scan-progress COMPLETED after this, so the agent knows where this fits. It does not, however, explicitly state when to choose it over the other report-*-findings siblings beyond the implicit SAST scope.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

report-secrets-findingsAInspect

Upload a full Gitleaks JSON report for a secrets security scan. Pass --id and --report-file . Backend redacts secret payloads, stores the report, dedupes by bug hash, and inserts SECURITY bugs. Does not mark the scan COMPLETED — the secrets playbook calls update-scan-progress COMPLETED after this.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
reportFileYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does much of it: it discloses payload redaction, report storage, dedup by bug hash, insertion of SECURITY bugs, and explicitly states the postcondition it does NOT perform (marking the scan COMPLETED). Missing only auth/permission requirements and failure behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences that front-load the purpose, then invocation, then side effects and postcondition. Every clause adds information; nothing is repeated from the name or schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema and no annotations, yet the description covers purpose, parameter passing, backend side effects, and the downstream workflow step an agent must take. For a two-param upload tool this is essentially complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for two required params, so the description must compensate. It ties the params to CLI flags ('--id', '--report-file') and specifies the reportFile format (Gitleaks JSON), but never defines what 'id' identifies (scan id, project id?) or gives the expected file path form.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Upload a full Gitleaks JSON report for a secrets security scan.' The 'Gitleaks'/'secrets' framing cleanly separates it from the sibling report-sast-findings, report-dast-findings, and report-deps-findings tools without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains invocation ('Pass --id and --report-file <path>') and, crucially, the workflow ordering: it does not mark the scan COMPLETED and the playbook calls update-scan-progress COMPLETED afterward. That routing guidance is strong, though it stops short of stating when this tool should be used instead of the other report-*-findings siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

unmark-entity-distinctCInspect

Remove a distinct mark between two same-type entities (same identity rules as mark-entity-distinct).

ParametersJSON Schema
NameRequiredDescriptionDefault
focusTestNo
otherTestNo
entityTypeYes
focusOrdinalIdNo
otherOrdinalIdNo
focusEventTitleNo
otherEventTitleNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It does not disclose idempotency (what happens if the mark is absent), permissions, reversibility, or any side effects, which is a meaningful gap for a mutation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler; the verb and object come first. It is efficient, though extremely terse for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter tool with nested objects, an enum, no annotations, and no output schema, one sentence is not enough. An agent lacks enough to correctly construct the focus/other entity pairs or know the result of the operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 7 parameters, including nested focusTest/otherTest objects and ordinal-id/title pairs, and the description explains none of them. The only crumb is the cross-reference to mark-entity-distinct's identity rules, which doesn't compensate for the undocumented parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Remove) and resource (a distinct mark between two same-type entities), and names the inverse sibling mark-entity-distinct, so an agent can tell it is the undo counterpart. 'Distinct mark' is domain jargon that isn't explained, but the operation is otherwise unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The reference to 'same identity rules as mark-entity-distinct' implies this is the inverse of that sibling, which is a weak form of usage guidance. There is no explicit when-to-use/when-not statement and no guidance on what happens if no mark exists.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update-git-folder-mappingCInspect

Update mapped plans/tests folder paths (and optional repository / plans branch) on the platform. Agent scaffolds folders in a PR; this records platform mapping.

ParametersJSON Schema
NameRequiredDescriptionDefault
plansBranchNoBranch plans sync against. Empty string resets to the repository default branch.
plansFolderPathNo
testsFolderPathNo
plans_folder_pathNo
tests_folder_pathNo
repositoryFullNameNo
repository_full_nameNo

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It implies a mutation ('Update', 'records platform mapping') but discloses no permissions required, no reversibility, no effect on unmapped/omitted fields, and does not explain the duplicated snake_case/camelCase parameter pairs in the schema. Behavioral context is thin for a write tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the mutation and its resource front-loaded, and the second sentence adds useful workflow context. Slightly cryptic phrasing ('Agent scaffolds folders in a PR') but no wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A mutation tool with no annotations, no output schema, 7 parameters all optional, and 14% schema coverage needs substantially more description than provided. Key gaps: why every parameter is optional, what happens on a partial update, and the purpose of the duplicated parameter names.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 14% (just plansBranch), so the description must compensate, and it only loosely maps to the parameter groups ('plans/tests folder paths', 'repository / plans branch'). It does not explain the duplicated plansFolderPath/plans_folder_path and repositoryFullName/repository_full_name pairs, the empty-string reset behavior, or the fact that all 7 parameters are optional.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Update) and resource (mapped plans/tests folder paths on the platform), which clearly differentiates it from the sibling get-git-folder-mapping reader. It does not explicitly name the sibling, but the write/read distinction is inferable from the name and description.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The second sentence clarifies the workflow context — the agent scaffolds folders in a PR and this tool only records the resulting platform mapping — which implies when to use it. However, it names no alternative and gives no explicit when-not guidance or prerequisites (e.g., when to call get-git-folder-mapping first).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update-issue-statusAInspect

Update a TestChimp issue status by ordinal id (same flexible issueId formats as get-issue-details). status must be one of: ACTIVE, IGNORED, FIXED, DUPLICATE, IN_PROGRESS_BUG, ARCHIVED_BUG, BLOCKED. For /testchimp fix issue: set IN_PROGRESS_BUG after applying a code fix; set FIXED only after user confirmation / commits pushed. Optional ignoreReason when status is IGNORED: INTENDED_BEHAVIOUR | INACCURATE_ASSESSMENT | NOT_IMPORTANT. Optional agentTraceability records UPDATED Activity inline (requires both workflowId and workflowExecutionId for Activity attachment).

ParametersJSON Schema
NameRequiredDescriptionDefault
gitShaNo
statusYes
userIdNo
issueIdYes
actorTypeNo
agentModelNo
branchNameNo
cliVersionNo
policyFileNo
workflowIdNo
ignoreReasonNo
skillVersionNo
policyVersionNo
agentTraceabilityNo
workflowExecutionIdNo

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full behavioral burden. It discloses the conditional requirement that setting IGNORED allows an ignoreReason, and that agentTraceability records UPDATED Activity inline but requires both workflowId and workflowExecutionId for attachment — genuinely useful side-effect and prerequisite disclosure. However, it doesn't state whether this requires specific permissions, is reversible, or what happens if workflowExecutionId is missing (does it fail silently? partial update?). One step short of full transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core mutation, then enumerates statuses, then conditional guidance. Dense but each clause carries load. Slightly crowded with enum breadcrumbs and nested parentheticals, but no filler sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 15 params, 0% schema coverage, nested object schema, and no output schema, the description supplies essential context: the status enum meaning, the ignoreReason condition, and the workflowId/workflowExecutionId coupling for Activity attachment. It's incomplete on actorType/agentModel/cliVersion/skillVersion/policyVersion semantics but covers the high-stakes ones. Missing logical guidance for gitSha (likely paired with FIXED) and the purpose of the many agent-traceability params.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It documents the status enum semantics (WHEN to use each value) and the ignoreReason enum values and condition. It does not explain most of the 15 parameters — e.g., actorType, cliVersion, skillVersion, policyVersion, policyFile, userId, gitSha, branchName — but does explain the key conditional relationships (workflowId + workflowExecutionId pairing). Well above baseline given 15 undocumented params and 0% coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (Update) and resource (TestChimp issue status) and clarifies the identifier format ('by ordinal id, same flexible issueId formats as get-issue-details'). It's readily distinguishable from siblings like 'get-issue-details' and 'create-issue' because it's the only status-mutation tool named. An agent knows exactly what this does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly documents when to use which status: set IN_PROGRESS_BUG after applying a code fix; set FIXED only after user confirmation / commits are pushed. It also routes to when ignoreReason applies (only when status is IGNORED). This is condition-specific, workflow-embedded guidance that goes well beyond the default.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update-plan-items-lifecycle-statusAInspect

Update lifecycle_fields.status for one user story or test scenario (DB only; does not rewrite plan markdown). entityType: story | scenario; ordinalId: numeric US-/TS- ordinal; status: draft | ready | in progress | blocked | done | archived. Used after /testchimp implement to set status to ready (unless policy overrides).

ParametersJSON Schema
NameRequiredDescriptionDefault
statusYes
ordinalIdYes
entityTypeYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it delivers key side-effect disclosure: it is DB-only and does not rewrite plan markdown. It also notes policy can override the intended ready status. It omits auth requirements, idempotency, and error/return behavior, but the mutation scope is unusually well disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Dense and front-loaded, leading with the effect and scope before params and usage. Every clause earns its place, though the single run-on structure for param definitions is slightly dense rather than strictly tight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-required-param mutation tool with no annotations and no output schema, the description covers purpose, side-effect boundaries, param semantics, and usage timing. It does not describe the response shape or failure modes, a minor gap for an otherwise complete definition.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% and there are no enums in the schema, yet the description supplies full semantics for all three params: entityType (story | scenario), ordinalId (numeric US-/TS- ordinal), and status (the complete draft/ready/in progress/blocked/done/archived value set). This more than compensates for the schema gap and provides enum values the schema lacks.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (Update) and precise resource (lifecycle_fields.status for one user story or test scenario), and adds the scope qualifier 'DB only; does not rewrite plan markdown.' This clearly distinguishes it from siblings like update-user-story and update-test-scenario, which edit content rather than lifecycle status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states the trigger context: 'Used after /testchimp implement to set status to ready (unless policy overrides).' This tells the agent when it applies and the policy caveat. It stops short of naming the alternative tools to use for other update needs, so it's clear context without full routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update-project-init-statusBInspect

Merge project-init progress fields for the current project. Server recomputes overall_complete when required items are DONE.

ParametersJSON Schema
NameRequiredDescriptionDefault
statusYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden. It does disclose two useful traits: 'Merge' signals a partial update rather than a replace, and the server recomputes overall_complete when required items are DONE, which tells the agent not to set overall_complete manually. It omits permission requirements, reversibility, and the meaning of the enum states.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the operation and followed by the notable server-side side effect. Nothing is padded or redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a mutation tool with no annotations, no output schema, and a high-complexity nested enum payload, the description covers the key behavioral quirk (auto-recomputation) but leaves the parameter vocabulary and permission/error context undocumented. It is adequate but noticeably thin for the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is one nested 'status' object with 0% schema description coverage and no descriptions on the many enum fields. The description hints that overall_complete is derived, but it does not name the item fields, explain the eight enum values, or clarify the duplicated snake_case/camelCase keys, leaving the agent to reverse-engineer the payload.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (merge) and resource (project-init progress fields) scoped to the current project, which clearly distinguishes it from the sibling get-project-init-status (a read). It does not name any sibling explicitly, but the read/write contrast is inferable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by 'merge ... progress fields for the current project,' so an agent can guess this is how you advance init progress. However, there is no explicit when-to-use guidance, no mention of prerequisites, and no routing against alternatives like get-project-init-status or mark-plan-items-implementation-done.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update-scan-progressAInspect

Update a scan's status. status must be one of: QUEUED, IN_PROGRESS, COMPLETED, EXCEPTION. Call IN_PROGRESS when starting. Each scan is a single checker type: the category playbook sets COMPLETED after a successful report-*-findings (or EXCEPTION on hard failure).

ParametersJSON Schema
NameRequiredDescriptionDefault
idYes
statusYes

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses the mutation nature and a state machine (playbook drives COMPLETED/EXCEPTION), which is useful context, but says nothing about permissions, reversibility, or what happens to a scan on an invalid transition.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three front-loaded sentences with no filler; the core purpose and enum come first, followed by operational context. The final sentence is dense but every clause carries the lifecycle information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter state-update tool with no annotations and no output schema, the description supplies the essential status semantics and workflow. It is nearly complete, with only minor gaps around the 'id' parameter and error/transition behavior.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the schema offers only raw types and an enum. The description re-lists the enum values (already in the schema) but does add real semantic guidance on which status to set and when, which is meaningfully beyond the bare enum. The 'id' parameter remains unexplained in both places.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Update a scan's status'), which cleanly distinguishes it from the many other update-* siblings that touch issues, plan items, and project-init status. It does not explicitly name an alternative, but the resource is unambiguous enough that an agent can route to it correctly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete lifecycle guidance: 'Call IN_PROGRESS when starting' and explains that the category playbook sets COMPLETED after a successful report-*-findings, or EXCEPTION on hard failure. This tells the agent when each status applies, though it never states explicit when-not-to-use conditions or names an alternative tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update-test-scenarioAInspect

Sync a test scenario markdown file to the platform after local edits. Requires frontmatter id: TS- and story: US-. Missing either returns an error telling you to call create-test-scenario first. Parses frontmatter and updates linking if story changes. Optional agentTraceability records UPDATED Activity inline.

ParametersJSON Schema
NameRequiredDescriptionDefault
gitShaNo
userIdNo
contentYes
actorTypeNo
agentModelNo
branchNameNo
cliVersionNo
policyFileNo
workflowIdNo
skillVersionNo
policyVersionNo
agentTraceabilityNo
workflowExecutionIdNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does disclose meaningful behavior: required frontmatter fields, the error path when they are absent, frontmatter parsing, and link updates when the story changes. It does not state permission/auth requirements or what a successful response looks like, so it stops short of full disclosure for a write-style sync.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, front-loaded with the core purpose and prerequisites, with no filler. Slightly dense but every sentence carries information; minor room to trim the traceability sentence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema and no annotations, so the description bears the full load. It covers purpose, prerequisites, and mutation behavior adequately, but with 13 parameters and a nested object largely unexplained, an agent still lacks guidance on most inputs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 13 parameters, so the description must compensate. It only explains one parameter meaningfully ('Optional agentTraceability records UPDATED Activity inline') and the implicit content format; the remaining eleven traceability/metadata fields (gitSha, userId, actorType, agentModel, branchName, cliVersion, policyFile, workflowId, skillVersion, policyVersion, workflowExecutionId) are undocumented.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Sync a test scenario markdown file to the platform') and adds the scoping condition ('after local edits'). It explicitly distinguishes itself from create-test-scenario, so an agent can route between them without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear when-to-use context ('after local edits') and names the alternative explicitly: missing frontmatter id/story routes the caller to create-test-scenario first. The condition that selects the alternative is spelled out rather than left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

update-user-storyAInspect

Sync a user story markdown file to the platform after local edits. Requires frontmatter id: US- (platform-issued). Missing id returns an error telling you to call create-user-story first. Parses frontmatter (id, title, priority) and updates the linked support file and entity. Optional agentTraceability records UPDATED Activity inline.

ParametersJSON Schema
NameRequiredDescriptionDefault
gitShaNo
userIdNo
contentYes
actorTypeNo
agentModelNo
branchNameNo
cliVersionNo
policyFileNo
workflowIdNo
skillVersionNo
policyVersionNo
agentTraceabilityNo
workflowExecutionIdNo

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses the platform-issued id precondition, explicit error behavior ("Missing id returns an error telling you to call create-user-story first"), side effects (updates the linked support file AND entity), and that agentTraceability records an UPDATED Activity inline. What is missing is auth/permission requirements and what happens on partial failures or conflicting versions.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Five tight sentences, front-loaded with the core action and precondition, then side effects and the optional traceability behavior. No filler, though the frontmatter-field enumeration could be trimmed or pushed to the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 13-parameter, nested-object, no-output-schema tool with no annotations, the description covers purpose, precondition, error handling, side effects, and the optional nested parameter's effect. It leaves auth requirements and the many tracing parameters undocumented, which is the main remaining gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 13 parameters, so the description must compensate and only partly does: it defines the content payload's expected structure (frontmatter with id, title, priority) and the purpose of the optional agentTraceability object. The remaining ten telemetry-ish parameters (gitSha, userId, actorType, agentModel, branchName, cliVersion, policyFile, workflowId, skillVersion, policyVersion, workflowExecutionId) are left entirely to the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource ("Sync a user story markdown file to the platform after local edits") with the operative mode of operation (sync, not create) made explicit. It also disambiguates from the create-user-story sibling by describing the frontmatter id requirement and the error path. An agent can distinguish this from create-user-story and update-test-scenario without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

States the triggering condition ("after local edits") and names the alternative (create-user-story) with the condition that selects it (missing frontmatter id). It does not describe when this tool should NOT be used, nor other update siblings, so it stops short of full routing guidance but the core when/when-else is present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upload-attachmentAInspect

Upload a file (e.g. agent screenshot evidence) to explore-snaps and return a stable view URL. Pass --file ; optional --filename and --content-type. Response includes viewUrl to paste in chat.

ParametersJSON Schema
NameRequiredDescriptionDefault
fileYes
filenameNo
contentTypeNo

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It does disclose the return value (viewUrl) and the persistence implication ('stable view URL'), but says nothing about permissions, size/type limits, or whether re-uploading overwrites — meaningful gaps for a write operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, zero waste, front-loaded with the action and outcome before parameter details. Every sentence adds distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema or annotations exist, so the description must cover behavior and returns; it does explain the response (viewUrl) and all params. It remains incomplete on auth/limits, but is adequate for a simple upload tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it names all three parameters ('--file <path>', optional '--filename', '--content-type'). It clarifies file is a path and the other two are optional, though it doesn't explain the format or behavior of filename/contentType overrides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (upload), resource (a file/attachment), and destination (explore-snaps), plus the outcome (returns a stable view URL). No sibling does uploading, so an agent can identify this tool unambiguously.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for use ('e.g. agent screenshot evidence') and a downstream purpose ('viewUrl to paste in chat'). However, it never names an alternative or states when NOT to use it, so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upsert-plans-support-fileAInspect

Create or update any file under the mapped plans root on the platform by relative path (no git commit/push required). Primary use: upload workflow execution plans at knowledge/workflow_plans//.plan.md after the Plan phase. filePath is relative to the plans mapped root (leading plans/ is stripped). Under workflow_plans/, filenames are coerced to *.plan.md and stored as WORKFLOW_EXECUTION_PLAN. Response includes supportFileId, filePath (canonical), filetype, created. Blocking step before Execute for cloud agents.

ParametersJSON Schema
NameRequiredDescriptionDefault
contentYes
filePathYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: no git commit/push needed, filePath is relative to the plans root with leading plans/ stripped, filenames under workflow_plans/ are coerced to *.plan.md and stored as WORKFLOW_EXECUTION_PLAN, and the response fields are listed. This is rich disclosure of non-obvious mutation behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the core purpose in the first sentence and each subsequent sentence adds operational detail (path semantics, coercion, response fields, blocking role). It is dense but every sentence earns its place; slightly long but justified.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-param mutation tool with no annotations and no output schema, the description covers purpose, path semantics, storage/coercion behavior, response fields, and workflow role. Nothing essential for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It explains filePath thoroughly (relative to the plans root, leading plans/ stripped, filename coercion rules), though the content parameter is only implied as the file body. Strong compensation for one of two params.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (create or update a file under the mapped plans root) with clear scope, and the read counterpart get-plans-support-file is easily distinguished. An agent can identify exactly what this tool does without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete primary use case (upload workflow execution plans after the Plan phase) and context that it is a blocking step before Execute for cloud agents. It does not explicitly name alternatives or when-not to use it, but the intended context is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upsert-policyAInspect

Create or update a workflow policy file on the platform (plans/knowledge/policies/*.policy.md). policyFileName is coerced to *.policy.md (same as get-policy). Prefer after writing the file locally so the policy is available immediately (git sync also works later).

ParametersJSON Schema
NameRequiredDescriptionDefault
contentYes
policyFileNameYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses the filename coercion to *.policy.md and the on-disk location, but omits write semantics such as whether an existing policy is overwritten, auth/permission requirements, or any confirmation behavior expected for a mutation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tightly packed sentences that lead with the core action and path, then the coercion rule, then usage timing. No filler, and each sentence adds a distinct fact, though the phrasing is somewhat dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Covers the essentials for locating and invoking the tool, but for a write/upsert operation with no annotations and no output schema, it should say more about overwrite behavior and permissions. Adequate but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It fully explains policyFileName (coerced to *.policy.md, matching get-policy), but says nothing about the content parameter beyond its name, leaving one of two required parameters undocumented in both the schema and the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (create or update = upsert) plus the exact resource and target path (plans/knowledge/policies/*.policy.md). An agent can immediately distinguish this from the read-oriented sibling get-policy and from upsert-plans-support-file, which targets a different file.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear conditional guidance: prefer this after writing the file locally so the policy is available immediately, noting git sync also works later as an alternative path. It stops short of naming a specific alternate tool or explicit when-not-to-use conditions, so it is strong context rather than full routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

upsert-screen-statesBInspect

Merge screen names and state strings into the project's relational atlas (idempotent upsert). Body uses camelCase screenStates: [{ screen, states: string[] }, ...] per UpsertScreenStatesRequest.

ParametersJSON Schema
NameRequiredDescriptionDefault
screenStatesYes

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It does disclose a genuine trait beyond schema: the operation is an idempotent upsert with merge semantics, which is useful. It omits auth requirements, what happens to states not present in the payload (merge vs replace), and return/error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with the purpose front-loaded and the payload detail second. Little waste, though the trailing 'per UpsertScreenStatesRequest' references a type whose definition is not supplied, adding a dangling pointer.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No annotations and no output schema exist, so the description is the only source of behavior. It covers purpose, idempotency, and body shape, but a mutation tool this bare should also say what it returns and whether it needs project scope; the definition is adequate but leaves real gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate. It usefully restates the camelCase screenStates body shape and the [{ screen, states: string[] }] nesting. However it does not clarify that 'screen' is optional while 'states' is required, the minItems constraint, or ordering/dedup semantics, so it does not fully close the coverage gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (merge/upsert) and resource (screen names and state strings into the project's relational atlas), so the action is unambiguous. It implicitly contrasts with the read-only sibling list-screen-states but never names it, and 'relational atlas' is somewhat opaque jargon, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no when-to-use guidance, no exclusions, and never points to the obvious alternative list-screen-states for reads. An agent must infer that this is the write counterpart; nothing about prerequisites or scenarios is provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 91 tool updates
    • First observedack-bot-events
    • First observedcompare-perf-to-baseline
    • First observedcreate-issue
    • First observedcreate-test-scenario
    • First observedcreate-user-story
    • First observeddestroy-ephemeral-environment
    • First observedfetch-execution-report
    • First observedget-api-operation-detail
    • First observedget-batch-view-url
    • First observedget-bot-compat
    • First observedget-bot-profile
    • First observedget-branch-specific-endpoint-config
    • First observedget-bunnyshell-workflow-job-logs
    • First observedget-eaas-config
    • First observedget-ephemeral-environment-status
    • First observedget-execution-history
    • First observedget-git-folder-mapping
    • First observedget-issue-details
    • First observedget-last-run-workflow-detail
    • First observedget-manual-session-details
    • First observedget-meeting-set
    • First observedget-meeting-transcript
    • First observedget-my-tasks
    • First observedget-org-capabilities
    • First observedget-perf-run
    • First observedget-plans-support-file
    • First observedget-policy
    • First observedget-project-init-status
    • First observedget-qa-posture
    • First observedget-release
    • First observedget-release-details
    • First observedget-requirement-coverage
    • First observedget-requirement-quality-report
    • First observedget-security-scan-config
    • First observedget-spec-lifecycle-details
    • First observedget-suite-execution-stats
    • First observedget-test-scenarios
    • First observedget-truecoverage-child-event-tree
    • First observedget-truecoverage-event-details
    • First observedget-truecoverage-event-metadata-keys
    • First observedget-truecoverage-event-time-series
    • First observedget-truecoverage-event-transition
    • First observedget-truecoverage-events
    • First observedget-truecoverage-session-metadata-keys
    • First observedget-user-stories
    • First observedget-workflow-execution
    • First observedlist-api-operation-interactions
    • First observedlist-api-operation-services
    • First observedlist-api-operations
    • First observedlist-bunnyshell-environment-events
    • First observedlist-bunnyshell-workflow-jobs
    • First observedlist-meeting-filter-options
    • First observedlist-meetings
    • First observedlist-perf-baselines
    • First observedlist-perf-runs
    • First observedlist-policies
    • First observedlist-related-perf-tests
    • First observedlist-rum-environments
    • First observedlist-screen-states
    • First observedlist-semantic-nearby
    • First observedlist-semantic-similar-tests
    • First observedlist-test-scenarios-for-scope
    • First observedlist-tests-awaiting-verification
    • First observedlist-workflow-catalog
    • First observedlist-workflow-executions
    • First observedmark-entity-distinct
    • First observedmark-plan-items-implementation-done
    • First observedmark-semantic-tests-distinct
    • First observedmark-tests-for-review
    • First observedpromote-perf-baseline
    • First observedprovision-ephemeral-environment
    • First observedprovision-ephemeral-environment-and-wait
    • First observedregister-bot-profile
    • First observedreport-agent-action
    • First observedreport-dast-findings
    • First observedreport-deps-findings
    • First observedreport-requirement-quality-findings
    • First observedreport-sast-findings
    • First observedreport-secrets-findings
    • First observedunmark-entity-distinct
    • First observedupdate-git-folder-mapping
    • First observedupdate-issue-status
    • First observedupdate-plan-items-lifecycle-status
    • First observedupdate-project-init-status
    • First observedupdate-scan-progress
    • First observedupdate-test-scenario
    • First observedupdate-user-story
    • First observedupload-attachment
    • First observedupsert-plans-support-file
    • First observedupsert-policy
    • First observedupsert-screen-states

Publisher details

Operator
Not applicable
Operator website
https://testchimp.io
Vendor relationship
First-party
Trust center
Unknown
Restrictions
Unknown

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables evidence-first regression testing for AI agents by turning production traces into reviewed, replayable cases that gate releases. It supports reproducible, auditable agent evaluation with controlled tool execution, evidence-based judging, and versioned quality gates.
    MIT
  • A
    license
    A
    quality
    A
    maintenance
    Provides coding agents with visibility into test health through tools for flaky test detection, test quality linting, and LLM evaluation harness, enabling them to triage failures, review test quality, and check prompt changes for regressions.
    9
    Apache 2.0
  • A
    license
    A
    quality
    D
    maintenance
    Policy and quality engine for AI coding agents that enforces team coding standards and provides validation gates for agent-assisted software delivery.
    7
    44 npm
    4
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Connect AI agents to your test results, insights, and targets. Query test runs, failures, flaky tests, and regressions across frameworks including Playwright, Jest, Pytest, Cypress and more.
    36 npm
    MIT
Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.