jira-xray-cloud-mcp
Exposes Xray Cloud's GraphQL API, including a raw GraphQL toolset with query, mutation, and schema-lookup tools. The Xray GraphQL schema is bundled so the model can inspect it and run arbitrary queries or mutations for anything the dedicated tools do not cover.
Provides tools for Xray Test Management running on top of Jira Cloud: creates, updates, and deletes Jira test-related issues (Tests, Preconditions, Test Sets, Test Plans, Test Executions), accepts Jira issue keys (e.g. CALC-12) as well as numeric issue ids (resolving keys to ids automatically), reads requirement coverage for Jira issues, and retrieves change history and project settings.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@jira-xray-cloud-mcplist test runs for issue CALC-12"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
jira-xray-cloud-mcp
MCP server for Xray Test Management on Jira Cloud. It exposes the Xray Cloud REST API v2 and GraphQL API as MCP tools. Tools are grouped into toolsets, each split into a read and a write half that you can switch on and off. Single tools can be added or removed on top of that, and a read-only switch removes every write tool from the server.
84 Xray tools in 14 toolsets / 25 halves (31 read, 53 write), plus
xray_list_toolsets.Credentials from the environment (one Xray API key for the server) or from HTTP request headers (every user brings their own key).
Tools take Jira issue keys (
CALC-12) as well as numeric issue ids. Xray's GraphQL API itself only accepts ids; the server looks the keys up.Write tools carry MCP annotations (
readOnlyHint,destructiveHint,idempotentHint), so clients can ask for confirmation before destructive calls.The raw GraphQL toolset (
xray_graphql_query,xray_graphql_mutation,xray_graphql_schema) covers whatever the dedicated tools don't. The Xray schema is bundled, so the model can look it up.
Configuration
All settings are environment variables (see .env.example):
Variable | Default | Meaning |
|
|
|
| – | Client id of an Xray API key (Xray → Global Settings → API Keys). |
| – | Client secret of that API key. Every call runs as the key's Jira user. |
|
| Regional hosts: |
|
|
|
|
| Which toolset halves to enable (see below). |
| – | Single tools to add to or remove from that selection (see below). |
|
| HTTP timeout in seconds. |
|
| Largest file that attachment/backup/export tools return inline. |
Credentials
Xray Cloud authenticates with API keys only (client id + secret, exchanged for a 24h token that the server caches).
XRAY_AUTH_MODE=env(default): the server usesXRAY_CLIENT_ID/XRAY_CLIENT_SECRETfor every caller. Suited to stdio and to single-user deployments.XRAY_AUTH_MODE=headers: every request must carryX-Xray-Client-IdandX-Xray-Client-Secret. Calls then run as that key's Jira user, and tokens are cached per key.XRAY_CLIENT_ID/XRAY_CLIENT_SECRETare ignored. This needs the HTTP transport; serve it over HTTPS only (e.g. behind a TLS-terminating reverse proxy), since the headers carry the secret.
Selecting toolsets
Each toolset is split into a read half (<name>_read) and a write half (<name>_write). Toolsets without
write tools only have a read half. XRAY_TOOLSETS is a comma-separated list, applied left to right:
default: both halves of every toolset except the opt-in ones (currentlybackup)all: every half<name>: both halves of a toolset, e.g.tests<name>_read/<name>_write: one half, e.g.tests_read*wildcards, e.g.*_read(this also matches opt-in toolsets)-<token>: remove, e.g.-graphqlor-*_write
Examples: default,backup · all,-graphql · *_read,test_runs_write (read everything, write only test
results) · tests_read,test_executions,test_runs.
Selecting single tools
XRAY_TOOLS works on top of XRAY_TOOLSETS, with the same syntax: <tool> adds a tool (also from a
toolset that is not selected), -<tool> removes one, * wildcards are allowed, and tokens are applied
left to right. Examples: -xray_delete_* (no delete tools) · xray_get_test_run,xray_update_test_run
(on top of a read selection). XRAY_READ_ONLY=true still wins over any tool added here.
The server refuses to start if a toolset, tool or pattern matches nothing, so a typo cannot silently
change the tool list. xray_list_toolsets reports what is enabled at runtime.
Toolsets and tools
Toolset | Scope |
|
|
| Test issues, steps, definitions, versions, datasets, links |
|
|
| Precondition issues |
|
|
| Test Set issues |
|
|
| Test Plan issues |
|
|
| Test Execution issues, environments |
|
|
| Results: status, comments, defects, evidence, steps, iterations, progress |
|
|
| Folder tree of the Test Repository and of Test Plans |
|
|
| Requirement coverage |
| – |
| Change history of Xray issues |
| – |
| Statuses, link types, project settings, step libraries |
| – |
| REST: result imports (Xray JSON, Cucumber, Behave, JUnit, TestNG, NUnit, xUnit, Robot), bulk test import, Cucumber features |
|
|
| REST: attachment storage |
|
|
| REST: Xray backups (admin) |
|
|
| Raw GraphQL + schema lookup |
|
|
xray_get_test_progress counts the Tests of a Test Plan or Test Execution per status on the server (e.g.
"412 Tests: 380 PASSED, 20 FAILED, 12 TODO"), so the model does not need to page through hundreds of runs.
Related MCP server: MCP Zephyr
Running
# stdio (local MCP clients)
XRAY_CLIENT_ID=... XRAY_CLIENT_SECRET=... jira-xray-cloud-mcp
# HTTP, as in docker-compose.yml
FASTMCP_TRANSPORT=http FASTMCP_HOST=0.0.0.0 FASTMCP_PORT=8000 jira-xray-cloud-mcp
# or: fastmcp run xray_mcp/server.py:create_server --transport http --host 0.0.0.0 --port 8000Register it with Claude Code, here locally over stdio, read-only, with a limited set of toolsets:
claude mcp add xray \
-e XRAY_CLIENT_ID=... -e XRAY_CLIENT_SECRET=... \
-e XRAY_READ_ONLY=true -e XRAY_TOOLSETS=tests,test_executions,test_runs,coverage \
-- jira-xray-cloud-mcpOr connect to a shared HTTP deployment running with XRAY_AUTH_MODE=headers, using your own API key:
claude mcp add --transport http xray https://xray-mcp.example.com/mcp \
--header "X-Xray-Client-Id: ..." --header "X-Xray-Client-Secret: ..."Docker
Every release is published to ghcr.io/netcare-io/jira-xray-cloud-mcp as <version>,
<major>.<minor> and latest:
# HTTP, every user brings their own API key
docker run --rm -p 8000:8000 -e XRAY_AUTH_MODE=headers \
-e FASTMCP_TRANSPORT=http -e FASTMCP_HOST=0.0.0.0 ghcr.io/netcare-io/jira-xray-cloud-mcp:latest
# stdio, e.g. as the command of a local MCP client
docker run --rm -i -e XRAY_CLIENT_ID=... -e XRAY_CLIENT_SECRET=... ghcr.io/netcare-io/jira-xray-cloud-mcp:latestdocker compose up builds the production image locally and reads the variables from .env.
Development
pip install -e ".[dev]"
pytestThe tests run the server in memory against a fake Xray (httpx2.MockTransport), and over the in-process
HTTP stack for header credentials. Every GraphQL document a tool sends is validated, together with its
variables, against the bundled Xray schema. A test also fails when a tool is added without a test case.
main only accepts pull requests. CI runs pytest on Python 3.12 and 3.13 and builds the Docker image
for every PR. To release, run scripts/release.sh <version>: the first run opens a PR that bumps the
version, and a second run on main after the merge pushes the v<version> tag, which publishes the
image and creates the GitHub release.
Notes for coding agents are in AGENTS.md. Xray API docs:
REST API v2,
GraphQL API.
Available Tools
82 toolsxray_add_test_associationsXray Add Test AssociationsBIdempotent
Link one Test to Preconditions, Test Sets, Test Plans and/or Test Executions.
| Name | Required | Description | Default |
|---|---|---|---|
| test | Yes | Jira issue key (e.g. 'CALC-12') or numeric issue id. | |
| test_sets | No | Issue keys or ids. | |
| test_plans | No | Issue keys or ids. | |
| version_id | No | Test version, for preconditions / test executions. | |
| preconditions | No | Issue keys or ids. | |
| test_executions | No | Issue keys or ids. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=true and an open-world network call, so the safety profile is covered structurally. The description adds no behavioral context of its own - it does not say whether existing associations are replaced or appended, whether some target fields require version_id, or what authorization is needed for an additive link.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the action and its targets are stated immediately. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema and annotations present, the description need not cover returns or safety. However, for a 6-parameter mutation that is only meaningful if at least one target array is supplied, it never states that constraint nor how this tool relates to the many sibling add/remove association tools, leaving a real gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter is already documented in the schema, making 3 the baseline. The description's enumeration of Preconditions/Test Sets/Test Plans/Test Executions roughly mirrors the schema's array properties but adds no format, cardinality, or mutual-dependency detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ('Link') and resource ('Test') plus the four target association types, and the singular 'one Test' distinguishes it from the plural bulk siblings. It does not, however, explicitly contrast itself with xray_add_tests_to_test_set / xray_add_tests_to_test_plan, which perform the reverse-direction association, so a reader must infer the distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied by the purpose statement ('Link one Test to...'); there is no explicit 'use this when you have a single test and want to attach it to many containers' guidance and no mention of the sibling alternatives that do the same job from the other direction. No exclusions or prerequisites are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_add_test_environmentsXray Add Test EnvironmentsAIdempotent
Add Test Environments to a Test Execution; unknown environments are created.
| Name | Required | Description | Default |
|---|---|---|---|
| test_execution | Yes | Jira issue key (e.g. 'CALC-12') or numeric issue id. | |
| test_environments | Yes | Test Environment names, e.g. ['Chrome', 'staging']. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=true, destructiveHint=false and openWorldHint=true, so safety is covered. The description adds genuinely new behavioral context by disclosing the side effect that unknown environment names will be created rather than rejected.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler; the action comes first and the non-obvious side effect follows immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema, full parameter coverage, and annotations in place, the description only needs to supply the missing side-effect information, which it does. It could optionally note whether re-adding an existing environment is a no-op, but idempotentHint already implies this.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters (test_execution as a Jira key/id, test_environments as a name array) are documented in the schema. The description contributes nothing beyond what the schema already says, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Add) and resource (Test Environments to a Test Execution), which cleanly distinguishes it from the sibling xray_remove_test_environments. It stops short of any explicit sibling routing, but the action and target are unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use guidance, no prerequisites (e.g. required permissions on the Test Execution), and does not point to xray_remove_test_environments or xray_add_tests_to_test_execution as alternatives. Usage is only inferable from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_add_test_executions_to_test_planXray Add Test Executions To Test PlanCIdempotent
Associate Test Executions with a Test Plan.
| Name | Required | Description | Default |
|---|---|---|---|
| test_plan | Yes | Jira issue key (e.g. 'CALC-12') or numeric issue id. | |
| test_executions | Yes | Jira issue keys (e.g. 'CALC-12') or numeric issue ids. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=true, destructiveHint=false and openWorldHint=true, so the safety profile is fully covered structurally. The description adds no behavioral context of its own -- nothing about idempotency semantics, failure modes, or partial success when some execution keys are invalid.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero wasted words. It is efficient, though its brevity borders on under-specification rather than deliberate conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and annotations covering safety and idempotency, the description does not need to explain return values. What remains missing is the operational context an agent would want for a membership-mutating call: prerequisites, behavior on invalid keys, and relationship to the remove/add-tests siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both parameters documented in the schema itself (issue key or numeric id). The description adds no format, cardinality, or batch-limit detail beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description pairs a specific verb ('Associate') with two specific resources ('Test Executions' and 'Test Plan'). This is enough to distinguish it from xray_add_tests_to_test_plan (tests, not executions) and xray_remove_test_executions_from_test_plan (inverse operation), though it never names a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus alternatives, no prerequisites (e.g. that the plan and executions must already exist), and no mention of the inverse removal tool. Usage must be inferred entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_add_test_run_evidenceXray Add Test Run EvidenceB
Attach evidence files (base64 or uploaded attachment ids) to a test run.
| Name | Required | Description | Default |
|---|---|---|---|
| evidence | Yes | ||
| test_run_id | Yes | Test run id, from xray_get_test_execution or xray_get_test_runs. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare that this is a non-read-only, non-idempotent, open-world mutation, so the safety profile is covered externally. The description adds useful context that evidence may be supplied inline as base64 or by an uploaded attachment id, but it does not disclose duplicate behavior, permissions, or the consequence of calling it repeatedly for a non-idempotent operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no wasted words. It states the operation, the accepted input forms, and the target resource immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter mutation tool with annotations and an output schema, the description is minimally adequate. It does not explain non-idempotent duplicate behavior or the practical prerequisite of uploading a file first when using attachment_id, leaving meaningful operational context to be inferred from sibling tools or schema descriptions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 50%, so the description must help compensate. It does add the meaningful distinction that the 'evidence' payload can be base64 data or an uploaded attachment id, but it says nothing about test_run_id provenance or filename/mime_type requirements beyond what the schema already supplies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Attach') and resource ('evidence files') targeted at a 'test run,' which cleanly distinguishes it from broader attachment upload tools. It does not explicitly name or contrast with sibling tools such as xray_remove_test_run_evidence or xray_upload_attachment, so sibling differentiation is left implicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It mentions two supported input modes ('base64 or uploaded attachment ids'), which gives some usage context, but it never states when to use this tool versus alternatives like xray_upload_attachment or xray_remove_test_run_evidence. There are no prerequisites, exclusions, or workflow guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_add_test_stepXray Add Test StepC
Append a step to a manual Test.
| Name | Required | Description | Default |
|---|---|---|---|
| data | No | ||
| test | Yes | Jira issue key (e.g. 'CALC-12') or numeric issue id. | |
| action | Yes | ||
| result | No | ||
| call_test | No | Make this a 'Call Test' step calling this Test (key or id). | |
| version_id | No | Test version id; default is the current version. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=false, destructiveHint=false, and openWorldHint=true, covering the safety and repeat-call profile. The description reinforces accumulation semantics with 'append', but adds no auth, ordering, or side-effect detail beyond the structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with the verb and resource front-loaded and zero filler. It is efficient, though the brevity is also a symptom of the missing detail scored elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter mutation tool with half its parameters undocumented, the description is too thin. Output schema means return values need not be explained, but step ordering, the meaning of 'data' vs 'action', and the Call Test option are left entirely to the schema or unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 50%: 'test', 'call_test', and 'version_id' are documented, while 'action', 'result', and especially the ambiguous 'data' are not. The description supplies none of those semantics and gives no hint of the relationship between action, data, and result.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Append a step to a manual Test.' Against siblings like xray_update_test_step and xray_remove_test_step, the verb 'append' is reasonably distinguishing. It stops short of explicitly naming those alternatives, but the action is clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No statement of when to use this over xray_update_test_step (edit an existing step) or xray_add_test_step variants. The scope restriction 'manual Test' hints at applicability, but there is no explicit guidance or exclusion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_add_tests_to_preconditionXray Add Tests To PreconditionCIdempotent
Link Tests to a Precondition.
| Name | Required | Description | Default |
|---|---|---|---|
| tests | Yes | Jira issue keys (e.g. 'CALC-12') or numeric issue ids. | |
| precondition | Yes | Jira issue key (e.g. 'CALC-12') or numeric issue id. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety/mutability profile is covered. The description adds no behavioral context of its own — nothing about duplicate handling, permission needs, or what happens if a test is already linked.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler. It is arguably too terse for the information an agent needs, but as conciseness goes it wastes nothing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be explained, and annotations cover the mutation safety profile. Still, for a link operation whose counterpart exists as a sibling, the definition leaves the agent to guess about duplicate links and the unlink path — adequate but not complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both parameters documented as Jira issue keys or numeric ids, so the baseline is 3. The description adds no additional meaning beyond the schema (no cardinality note, no ordering semantics for the array).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Link') and both resources ('Tests', 'Precondition'), so an agent knows exactly what operation occurs. It does not, however, distinguish itself from the sibling xray_remove_tests_from_precondition or the parallel add_tests_to_test_set/test_plan tools, so sibling differentiation is absent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites (e.g. the precondition must already exist), and no mention of the inverse operation for unlinking. The agent must infer usage entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_add_tests_to_test_executionXray Add Tests To Test ExecutionAIdempotent
Add Tests to a Test Execution (creates a test run per Test).
| Name | Required | Description | Default |
|---|---|---|---|
| tests | Yes | Jira issue keys (e.g. 'CALC-12') or numeric issue ids. | |
| test_execution | Yes | Jira issue key (e.g. 'CALC-12') or numeric issue id. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnly=false, destructive=false, idempotent=true, openWorld=true), so the bar is lower. The description still adds real value by disclosing the side effect that each added Test spawns a test run, which the annotations do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the core action front-loaded and the notable side effect supplied in a parenthetical. Nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With a full schema, an output schema present, and annotations declaring the safety and idempotency profile, the essential information is covered. The only real gap is the absence of routing guidance relative to the numerous sibling add/remove tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (tests array of Jira keys/ids and test_execution key/id) are already fully documented in the schema. The description adds no further syntax or format detail, placing it at the baseline of 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Add) and resource (Tests to a Test Execution), and adds the non-obvious side effect that a test run is created per Test. It does not differentiate from close siblings like xray_add_tests_to_test_set or xray_add_tests_to_test_plan, which share the same verb pattern, so an agent must infer the target-resource distinction from the name alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to use this tool versus alternatives such as xray_remove_tests_from_test_execution or xray_create_test_execution, nor any prerequisite or exclusion. Usage is only implied by the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_add_tests_to_test_planXray Add Tests To Test PlanCIdempotent
Add Tests to a Test Plan.
| Name | Required | Description | Default |
|---|---|---|---|
| tests | Yes | Jira issue keys (e.g. 'CALC-12') or numeric issue ids. | |
| test_plan | Yes | Jira issue key (e.g. 'CALC-12') or numeric issue id. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=true, destructiveHint=false and openWorldHint=true, so the safety/idempotency profile is covered structurally. The description adds nothing beyond that — no note on duplicate handling (relevant given idempotentHint), no auth/permission requirements, no effect on existing plan membership.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is short and front-loaded, but the single sentence is a verbatim echo of the name rather than an efficient information payload, so brevity here reflects under-specification rather than disciplined concision.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has a clear mutation semantic and an output schema exists, so return values need not be described. However, for a write operation with no description-level guidance on behavior (duplicates, limits, ordering), the definition is too thin to let an agent invoke it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters (test_plan, tests) are fully documented in the schema with key/id format examples. The description contributes no additional meaning, which matches the baseline 3 for high-coverage schemas.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Add Tests to a Test Plan.' is a direct restatement of the tool name (xray_add_tests_to_test_plan) and title, adding no distinguishing information. It conveys the verb and resource only because the name already does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus siblings such as xray_remove_tests_from_test_plan or xray_add_test_executions_to_test_plan. No prerequisites, no mention of what happens when a test is already in the plan, no alternative routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_add_tests_to_test_setXray Add Tests To Test SetCIdempotent
Add Tests to a Test Set.
| Name | Required | Description | Default |
|---|---|---|---|
| tests | Yes | Jira issue keys (e.g. 'CALC-12') or numeric issue ids. | |
| test_set | Yes | Jira issue key (e.g. 'CALC-12') or numeric issue id. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=true, and destructiveHint=false, so the safety profile is covered structurally. The description adds nothing beyond that – it does not explain what happens on repeat calls, whether existing memberships are preserved, or what the operation returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
It is one short sentence, but this is under-specification rather than conciseness. There is no wasted wording because there is almost no wording at all; nothing is front-loaded because nothing is provided.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a mutation tool in a large family of parallel add/remove operations, the description is far too thin to guide correct selection. An output schema exists so return values need not be explained, but the absence of any behavioral or routing context leaves the definition inadequate.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both parameters documented (Jira issue keys or numeric ids). The description adds no meaning beyond the schema, so the baseline of 3 applies when the schema does all the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description "Add Tests to a Test Set" essentially restates the tool name and title verbatim. It names a verb and resource, but provides no scope or differentiation from siblings like xray_add_tests_to_test_plan or xray_add_tests_to_precondition, which share the identical 'add tests to X' pattern.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool, when not to, or how it differs from the parallel add/remove tools for test plans, preconditions, and test executions. Usage is only implied by the tool name, and nothing routes the agent among the near-identical siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_create_folderXray Create FolderA
Create a folder in a project's Test Repository or in a Test Plan, optionally with Tests.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Path of the new folder, e.g. '/Regression/API'. | |
| tests | No | Test keys or ids to move into the folder. | |
| project | No | Jira project key or id (for Test Repository folders). | |
| test_plan | No | Test Plan key or id (for Test Plan folders instead of the repository). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnlyHint=false, idempotentHint=false, destructiveHint=false, openWorldHint=true), so the description only needs to add context. It notes that tests can be moved into the new folder in the same call, which is useful, but says nothing about failure modes (existing path, invalid project/test plan) or permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler; the resource and the two placement modes come first, and the optional behavior comes last.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter create tool with an output schema and full annotation coverage, this is nearly sufficient: it names the two folder locations and the optional test attachment. It leaves out only edge-case behavior (path collisions, failure conditions), which keeps it short of a 5.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description reinforces the project-vs-test_plan mode distinction but adds no format, syntax, or constraint detail (e.g. path normalization, whether tests are moved or copied) beyond what the schema already documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Create) and resource (folder), and scopes it to two concrete parent locations ('a project's Test Repository or in a Test Plan') with an optional payload. This distinguishes it cleanly from sibling folder operations like xray_rename_folder, xray_move_folder, and xray_delete_folder.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied rather than stated: the description hints at the two modes via 'Test Repository or in a Test Plan' and the optional Tests, and the schema's project/test_plan params disambiguate them. But it never says when to prefer this over xray_move_to_folder for relocating existing tests, nor any prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_create_preconditionXray Create PreconditionB
Create a Precondition issue, optionally linked to Tests.
| Name | Required | Description | Default |
|---|---|---|---|
| tests | No | Test keys or ids to link. | |
| project | Yes | Jira project key or id. | |
| summary | Yes | ||
| definition | No | Precondition text / Gherkin background. | |
| folder_path | No | Test Repository folder, e.g. '/Setup'. | |
| extra_fields | No | Additional Jira fields for the new issue, in Jira REST format, e.g. {"description": "...", "labels": ["smoke"], "fixVersions": [{"name": "1.0"}]}. | |
| precondition_type | No | Precondition type name, e.g. 'Manual', 'Cucumber', 'Generic'. | Manual |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, openWorldHint=true, idempotentHint=false, and destructiveHint=false, so the mutation/safety profile is covered. The description adds one behavioral detail ('optionally linked to Tests'), but omits what is returned, whether linking requires pre-existing test keys, and any auth or folder constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler. It is efficient, though its brevity borders on under-specification for a 7-parameter create tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be explained, and the schema covers most parameters. However, for a create tool with optional linking, a defaulted type, and an open-ended extra_fields object, the description leaves real gaps around interaction with sibling linking tools and the meaning of the precondition types.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 86%, so the schema already documents project, tests, definition, folder_path, extra_fields, and precondition_type. The description adds nothing beyond what the schema provides, so a baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a Precondition issue') and adds the linking capability. It is clear but does not distinguish itself from sibling creation tools such as xray_create_test, relying on the reader to infer the differences from the name alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this versus xray_create_test or xray_add_tests_to_precondition, nor any prerequisites or warnings. The agent must infer usage entirely from the tool name and sibling list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_create_testXray Create TestB
Create an Xray Test issue (manual steps, Gherkin or unstructured definition).
| Name | Required | Description | Default |
|---|---|---|---|
| steps | No | Steps, for step-based (manual) tests. | |
| gherkin | No | Gherkin scenario, for Cucumber tests. | |
| project | Yes | Jira project key or id. | |
| summary | Yes | ||
| test_type | No | Xray test type name, e.g. 'Manual', 'Cucumber', 'Generic'. | Manual |
| folder_path | No | Test Repository folder, e.g. '/Regression'. | |
| extra_fields | No | Additional Jira fields for the new issue, in Jira REST format, e.g. {"description": "...", "labels": ["smoke"], "fixVersions": [{"name": "1.0"}]}. | |
| unstructured | No | Free-text definition, for Generic tests. | |
| preconditions | No | Precondition keys or ids to link. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the write profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false, openWorldHint=true), so the description's main added value is that the created issue is an Xray Test with one of three definition modes. It says nothing about required permissions, duplicate-summary behavior, or what happens on conflicting definition inputs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler or restated boilerplate. It is terse to the point of being a fragment, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With 9 parameters, two required, and an output schema present, the description does not need to explain return values. However, for a creation tool with three mutually exclusive definition payload styles, the lack of routing between test_type and steps/gherkin/unstructured leaves a real gap for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 89%, so the schema already documents nearly every parameter. The description echoes the definition modes (steps/Gherkin/unstructured) but adds no mapping detail such as 'Manual => steps, Cucumber => gherkin, Generic => unstructured' that would help the agent pick the right field.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Create an Xray Test issue') and enumerates the three definition flavors the tool supports, which distinguishes it from sibling create_* tools for preconditions, test sets and test plans. It does not explicitly name those siblings, but the resource noun is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance and no alternatives named. The agent must infer on its own whether to use this rather than xray_import_tests_bulk or xray_add_test_step, and nothing tells it which definition type to pick for a given input.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_create_test_executionXray Create Test ExecutionA
Create a Test Execution issue with Tests (one test run per Test). Link it to a plan via the test_plans tools.
| Name | Required | Description | Default |
|---|---|---|---|
| tests | No | Test keys or ids to run. | |
| project | Yes | Jira project key or id. | |
| summary | Yes | ||
| extra_fields | No | Additional Jira fields for the new issue, in Jira REST format, e.g. {"description": "...", "labels": ["smoke"], "fixVersions": [{"name": "1.0"}]}. | |
| test_environments | No | Test Environment names. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare a non-read-only, non-idempotent, open-world mutation. The description adds that creating an execution generates one test run per Test, which is useful behavioral context. It does not cover permissions, duplicate-call effects, or rate limits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the core action, no wasted words. The linking note is useful and brief.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With annotations covering safety, an output schema for returns, and 80% schema coverage, the description supplies the essential purpose and a key workflow hook (linking to plans). It could say more about idempotency or prerequisites, but it is largely sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 80%, so the schema documents most parameters (tests, project, extra_fields, test_environments). The description only gestures at tests ('with Tests') without adding format or constraint details beyond the schema. Baseline 3.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Create a Test Execution issue with Tests (one test run per Test) is a precise verb+resource and scope. It distinguishes from sibling creators (Test, Test Plan, Test Set, Precondition) by naming the exact issue type. The linking instruction further clarifies its place in the Xray workflow.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use versus alternatives (e.g., when to create a new execution vs. reusing an existing one, or when to use xray_add_test_executions_to_test_plan). The only guidance is a follow-up linking action via test_plans tools, which is not a usage condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_create_test_planXray Create Test PlanC
Create a Test Plan issue, optionally with Tests (e.g. fixVersions via extra_fields).
| Name | Required | Description | Default |
|---|---|---|---|
| tests | No | Test keys or ids to add. | |
| project | Yes | Jira project key or id. | |
| summary | Yes | ||
| extra_fields | No | Additional Jira fields for the new issue, in Jira REST format, e.g. {"description": "...", "labels": ["smoke"], "fixVersions": [{"name": "1.0"}]}. | |
| saved_filter | No | Jira filter id whose Tests are added to the plan. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare this is non-read-only, non-idempotent, non-destructive and open-world, so an agent knows calling it mutates state. The description adds nothing beyond that — it doesn't say a repeat call creates a second plan, what permissions are needed, or how the two test-sourcing paths (tests vs saved_filter) interact.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence is appropriately short, but the 'e.g. fixVersions via extra_fields' clause is a non sequitur attached to the tests clause and creates rather than removes ambiguity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With five parameters, two required, and two competing ways to attach tests (tests keys/ids vs saved_filter), the description should explain how those interact and which wins. Output schema exists so return values need no explanation, but the input-side ambiguity is left unaddressed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 80%, so the schema already documents tests, project, summary, extra_fields and saved_filter. The description restates the optional-Tests behavior and gives a fixVersions example that duplicates the extra_fields schema description rather than clarifying semantics. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Create a Test Plan issue') and notes the optional association with Tests, which distinguishes it from siblings like xray_create_test and xray_create_test_set. The trailing parenthetical is muddled, but the core purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus xray_add_tests_to_test_plan (for existing plans) or when to prefer saved_filter over an explicit tests list. The 'optionally with Tests' phrasing implies a use case but names no alternatives or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_create_test_setXray Create Test SetB
Create a Test Set issue, optionally with Tests.
| Name | Required | Description | Default |
|---|---|---|---|
| tests | No | Test keys or ids to add. | |
| project | Yes | Jira project key or id. | |
| summary | Yes | ||
| extra_fields | No | Additional Jira fields for the new issue, in Jira REST format, e.g. {"description": "...", "labels": ["smoke"], "fixVersions": [{"name": "1.0"}]}. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, openWorldHint=true, and idempotentHint=false, so the safety and mutability profile is largely covered. The description adds only that Tests may be included optionally, but it does not disclose auth requirements, side effects beyond creation, or rate limits. It is consistent with the annotations and adds limited value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no wasted words. It states the core action and the optional add-on directly. Every part of the sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. Annotations cover safety and mutability. However, the description omits usage guidance, does not mention required project and summary parameters, and does not mention extra_fields, leaving some gaps despite the structured schema support.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 75%, so the schema already documents most parameters including tests, project, and extra_fields. The description mentions the optional Tests parameter but does not clarify project, summary, or extra_fields behavior. Baseline 3 is appropriate because the schema does most of the semantic work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource: 'Create a Test Set issue'. It also notes the optional inclusion of Tests. It does not explicitly differentiate from nearby siblings such as xray_create_test or xray_add_tests_to_test_set, but the resource name 'Test Set' is distinct enough to be clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no when-to-use guidance, prerequisites, or alternatives. It is implied that this tool creates a new Test Set issue, but the agent receives no explicit direction on when to choose this over xray_add_tests_to_test_set or how to handle existing Test Sets.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_delete_folderXray Delete FolderADestructiveIdempotent
Delete a folder and its subfolders. The Test / Precondition issues in it are not deleted.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Folder path, e.g. '/Regression/API'. '/' is the root. | |
| project | No | Jira project key or id (for Test Repository folders). | |
| test_plan | No | Test Plan key or id (for Test Plan folders instead of the repository). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false, so the safety profile is covered. The description adds genuinely non-obvious behavior: subfolders are deleted recursively while the Test/Precondition issues inside are preserved, which an agent could otherwise get wrong. Permissions and irreversibility are not restated, hence not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero filler; the destructive scope is front-loaded and the preservation caveat follows immediately. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be explained, and annotations plus full schema coverage carry the rest. The description covers the two facts an agent most needs for a destructive folder operation (recursion and issue preservation); only permission requirements are unstated.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for all three parameters (path, project, test_plan), including the '/' root convention and the repository-vs-test-plan distinction. The description contributes nothing beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Delete a folder') plus the destructive scope ('and its subfolders'), which clearly separates it from siblings like xray_create_folder, xray_rename_folder, and xray_move_folder. It never names an alternative explicitly, so it stops short of full sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus alternatives such as xray_remove_from_folder (which detaches rather than deletes) or xray_move_folder. No prerequisites or preconditions are given, leaving the agent to infer the selection context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_delete_preconditionXray Delete PreconditionBDestructiveIdempotent
Delete a Precondition issue permanently.
| Name | Required | Description | Default |
|---|---|---|---|
| precondition | Yes | Jira issue key (e.g. 'CALC-12') or numeric issue id. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, so the safety profile is largely covered. The description adds the word 'permanently', reinforcing that this is not a recoverable/soft delete, which is useful context beyond the annotations. It does not disclose cascading effects (e.g., removal from linked tests) or auth requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short, front-loaded sentence with no wasted words. The destructive nature is conveyed immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and annotations cover the safety profile. However, for a permanent delete of an issue type that other entities (tests) reference, the description omits any note about side effects or linked-entity impact, leaving a meaningful gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Single parameter with 100% schema description coverage; the schema explains it accepts a Jira issue key or numeric id. The description adds nothing about the parameter, which is the expected baseline when the schema fully documents it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Delete) and resource (a Precondition issue), making the target object clear. It does not differentiate itself from siblings such as xray_delete_test or xray_delete_test_set beyond the resource noun, but the resource is distinct enough that selection is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus alternatives, no prerequisites, and no warning about consequences for tests that reference the precondition. The agent must infer the usage context entirely from the tool name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_delete_testXray Delete TestADestructiveIdempotent
Delete a Test issue permanently (Jira issue included).
| Name | Required | Description | Default |
|---|---|---|---|
| test | Yes | Jira issue key (e.g. 'CALC-12') or numeric issue id. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already flag destructiveHint, idempotentHint, and readOnlyHint=false, so the safety profile is covered. The description still adds real value by disclosing the non-obvious side effect that the Jira issue itself is deleted, expanding the blast radius beyond what the annotations state. It does not cover cascading effects on related test runs, steps, or executions, nor permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, zero filler, with the destructive nature and the surprising Jira-issue side effect front-loaded. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be explained, and annotations carry the disposition hints. For a destructive one-parameter tool the description is nearly sufficient, missing only cascade scope and permission/undo context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Single parameter with 100% schema description coverage, so the schema already documents the accepted key/id formats. The description adds nothing about the 'test' argument, making this a baseline 3 case where structured data does the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Delete), a precise resource (Test issue), and adds that the underlying Jira issue is removed too. The word 'Test' distinguishes it from sibling deletions like xray_delete_test_set, xray_delete_test_plan, and xray_delete_test_execution, though it never explicitly names those alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this versus the other delete tools or the alternative update/archive paths, and no stated prerequisites. 'Permanently' hints at caution but is not framed as a when-to-use condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_delete_test_executionXray Delete Test ExecutionADestructiveIdempotent
Delete a Test Execution issue permanently, including its test runs.
| Name | Required | Description | Default |
|---|---|---|---|
| test_execution | Yes | Jira issue key (e.g. 'CALC-12') or numeric issue id. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, but they do not say what is destroyed. The description adds that the deletion is permanent and removes associated test runs, which is meaningful cascade information beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the verb and resource, and the consequential detail (permanent, test runs included) is stated without padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the single parameter is fully documented in the schema. The cascade behavior is covered; only irreversibility warnings and auth expectations are absent, which is a minor gap for a one-parameter delete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The single parameter has 100% schema description coverage (Jira issue key or numeric id), so the schema already carries the semantics. The description adds nothing about the identifier format, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Delete) and resource (Test Execution issue), and adds scope: the deletion is permanent and cascades to the test runs. This distinguishes it from siblings like xray_delete_test, xray_delete_test_plan, and xray_delete_test_set by resource.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus alternatives such as xray_remove_tests_from_test_execution or xray_delete_test. No prerequisites, permissions, or confirmation requirements are stated, so the agent must infer usage entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_delete_test_planXray Delete Test PlanBDestructiveIdempotent
Delete a Test Plan issue permanently.
| Name | Required | Description | Default |
|---|---|---|---|
| test_plan | Yes | Jira issue key (e.g. 'CALC-12') or numeric issue id. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true and openWorldHint=true, so the safety profile is largely covered. The description's only addition is the word 'permanently', which usefully reinforces irreversibility but says nothing about cascading effects on associated test executions or required permissions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no waste. It is efficient, though arguably under-specified rather than optimally concise for a destructive operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With one fully documented parameter, an output schema, and annotations covering safety and idempotency, the description is sufficient to invoke the tool correctly. The only missing element is a note on side effects for related executions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and there is a single parameter, with the schema already documenting accepted formats (Jira issue key or numeric id). The description adds nothing beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Delete') and resource ('Test Plan issue') with the scope qualifier 'permanently'. It is distinguishable from siblings like xray_delete_test and xray_delete_test_set by naming the exact entity, though it does not explicitly contrast with them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus alternatives, nor on prerequisites or when deletion is inappropriate. For a destructive, irreversible operation, the absence of any 'when-not' context is a real gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_delete_test_setXray Delete Test SetADestructiveIdempotent
Delete a Test Set issue permanently (the Tests are kept).
| Name | Required | Description | Default |
|---|---|---|---|
| test_set | Yes | Jira issue key (e.g. 'CALC-12') or numeric issue id. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, so safety is covered. The description adds genuinely non-obvious behavioral context: the deletion is permanent and the member Tests survive the operation, which an agent could not infer from the schema or annotations alone.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence with the destructive action front-loaded and the critical side-effect in a parenthetical. Nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a one-parameter destructive tool with full annotation coverage and an output schema, the description supplies the essential behavior (permanent, non-cascading). It could mention any permission requirements or irreversible-confirmation guidance, but is nearly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter (test_set) with 100% schema description coverage, so the schema already documents the accepted key/id formats. The description adds no additional parameter meaning; baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (delete) and resource (Test Set issue) and adds the key scoping fact that contained Tests are kept, which meaningfully distinguishes it from xray_delete_test. It does not name sibling tools explicitly, so it stops short of full sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The word 'permanently' signals the destructive nature, but there is no explicit statement of when to use this versus xray_delete_test, xray_delete_test_plan, or xray_remove_tests_from_test_set, and no prerequisites or warnings about irreversibility are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_export_cucumber_featuresXray Export Cucumber FeaturesBRead-onlyIdempotent
Export Cucumber Tests as .feature files; returns each file's name and content.
| Name | Required | Description | Default |
|---|---|---|---|
| keys | No | Issue keys (Tests, Test Sets, Test Plans, Test Executions, requirements). | |
| filter_id | No | Jira saved filter id. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true and destructiveHint=false, so the safety profile is fully covered elsewhere. The description adds the return content shape ('each file's name and content'), but no behavioral context such as permission requirements, output format/archive details, or behavior when no selector is supplied.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the action first and the return preview second; nothing is padded. The return clause is largely redundant with the output schema, which keeps it just short of a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need no prose, and all parameters are documented in-schema. What's missing is selection behavior for a zero-required-parameter export (what is exported by default, and how keys vs filter_id are combined).
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% for both parameters, so the schema carries the semantics of keys and filter_id; the description adds nothing about their interaction or precedence. Baseline 3 applies when structured fields do the work.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Export Cucumber Tests as .feature files') and even previews the return shape. It does not explicitly contrast itself with the sibling xray_import_cucumber_features, but the export direction is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives, and no hint about how to choose between the two mutually competing selection inputs (keys vs filter_id). The name suggests the context but the description states no conditions or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_get_attachmentXray Get AttachmentARead-onlyIdempotent
Download an attachment or evidence file; text files come back as text, others as base64.
| Name | Required | Description | Default |
|---|---|---|---|
| attachment_id | Yes | Attachment or evidence id. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds genuinely useful behavioral context beyond that: return encoding varies by content type (text files as text, others as base64), which affects how the agent must handle the response.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the action, followed immediately by the return-format caveat. No filler and nothing important is buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With a rich annotation set, a fully documented single parameter, and an output schema present, the description only needs to add the caller-facing behavior it does (text vs base64 payload). Minor omissions such as error behavior or size limits remain, but nothing essential for correct invocation is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single parameter with 100% schema description coverage ('Attachment or evidence id.'), so the schema carries the semantics. The description adds no format, source, or lookup guidance for attachment_id, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (download) and resource (attachment or evidence file), and the name pairs it against its sibling xray_upload_attachment. It is clear what the tool does, though it never explicitly contrasts itself with any sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied: it is a fetch-by-id tool, so the agent infers it is used to retrieve a known attachment id. There is no statement of when to prefer this over, say, xray_get_test_run or xray_get_step_library, nor any prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_get_coverable_issueXray Get Coverable IssueBRead-onlyIdempotent
Get the coverage status of a requirement / story and the Tests covering it.
| Name | Required | Description | Default |
|---|---|---|---|
| issue | Yes | Jira issue key (e.g. 'CALC-12') or numeric issue id. | |
| fields | No | Jira fields to return in 'jira'. Default: ['key', 'summary', 'status', 'issuetype', 'project', 'assignee', 'labels']. | |
| version | No | Calculate coverage for this version name. | |
| is_final | No | Whether final statuses take precedence. | |
| test_plan | No | Calculate coverage within this Test Plan (key or id). | |
| environment | No | Calculate coverage for this Test Environment. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=true, so the safety profile is fully covered. The description adds the useful fact that the response includes both coverage status and the covering Tests, but says nothing about how coverage is computed or what happens when version/test_plan/environment scoping is combined.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single efficient sentence with the core action and result front-loaded and no filler. It is arguably too terse for a six-parameter tool with several scoping options, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is not required, and annotations cover safety. However, for a tool with six parameters offering version, test-plan, environment and final-status scoping, one sentence leaves the agent without guidance on how those modes interact or which combinations are meaningful.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter including version, is_final, test_plan and environment is already documented in the schema itself. The description adds no syntax, formatting, or interaction hints beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb and resource ('Get the coverage status of a requirement / story and the Tests covering it'), so an agent knows it returns coverage data plus linked tests. It does not, however, distinguish itself from the near-identical sibling xray_get_coverable_issues, leaving the singular/plural boundary to be inferred.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no indication of when to use this versus xray_get_coverable_issues, nor any mention of prerequisites such as needing a valid Jira issue key. The agent gets no routing guidance at all.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_get_coverable_issuesXray Get Coverable IssuesBRead-onlyIdempotent
Search coverable issues (requirements, stories...) with their coverage status (paginated).
| Name | Required | Description | Default |
|---|---|---|---|
| jql | No | JQL filter, e.g. 'project = CALC AND issuetype = Story'. | |
| limit | No | Page size (1-100). | |
| start | No | Page offset (index of the first result). | |
| fields | No | Jira fields to return in 'jira'. Default: ['key', 'summary', 'status', 'issuetype', 'project', 'assignee', 'labels']. | |
| issues | No | Restrict to these issue keys or ids. | |
| version | No | Calculate coverage for this version name. | |
| is_final | No | Whether final statuses take precedence. | |
| test_plan | No | Calculate coverage within this Test Plan (key or id). | |
| environment | No | Calculate coverage for this Test Environment. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds that results are paginated, which annotations do not convey, but says nothing about how coverage is computed, default page size, or the interaction between version/test_plan/environment scoping.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the resource stated first and the pagination constraint last; no filler. It is efficient, though arguably too terse given the tool's nine parameters.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and the schema fully documents parameters. But for a 9-parameter search tool whose behavior depends on coverage-scoping inputs (version, test_plan, environment, is_final), the description omits how these shape the result, leaving an inference gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (jql, limit, start, fields, issues, version, is_final, test_plan, environment) is already documented in the schema. The description contributes no additional parameter meaning, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Search) and resource (coverable issues: requirements, stories) plus the returned dimension (coverage status). However, it never distinguishes itself from the sibling xray_get_coverable_issue (singular), leaving the agent to infer the plural/singular split from the name alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use guidance, no conditions or exclusions, and no mention of the singular get_coverable_issue alternative or when a bulk/paginated search is preferable to a single-issue lookup. Only the '(paginated)' hint implies usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_get_datasetsXray Get DatasetsARead-onlyIdempotent
Get the datasets (parameters and rows) of data-driven Tests, optionally as overridden on Test Plans / Executions.
| Name | Required | Description | Default |
|---|---|---|---|
| tests | Yes | Test keys or ids. One Test (with at most one Test Plan / Execution) returns the dataset in effect there; otherwise all datasets stored for these Tests and the given overrides. | |
| test_plans | No | Dataset overrides on these Test Plans. | |
| test_executions | No | Dataset overrides on these Test Executions. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, non-destructive, open-world behavior. The description adds that the payload contains parameters and rows and that overrides may apply, but it does not expand on permissions, limits, or edge cases beyond what annotations/schema provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence with no filler; verb and resource lead. Every phrase contributes to scoping the request.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists, annotations state safety, and schema thoroughly documents parameters and the one-test-with-override edge case. The description supplies the essential purpose, so an agent has enough to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all three parameters are documented in the schema. The description only repeats the override idea and does not add syntax or format details beyond schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: retrieving datasets (parameters and rows) of data-driven Tests, with optional override context. It clearly separates this from sibling get_test tools that fetch test metadata rather than dataset content, though it does not explicitly name an alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Only implies when to use it through 'optionally as overridden on Test Plans / Executions'; it does not say when to choose this over other getters or when not to request overrides. The schema carries the real conditional logic.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_get_expanded_testXray Get Expanded TestARead-onlyIdempotent
Get a Test with 'Call Test' steps expanded inline into the steps of the called tests.
| Name | Required | Description | Default |
|---|---|---|---|
| test | Yes | Jira issue key (e.g. 'CALC-12') or numeric issue id. | |
| fields | No | Jira fields to return in 'jira'. Default: ['key', 'summary', 'status', 'issuetype', 'project', 'assignee', 'labels']. | |
| version_id | No | Test version id; default is the current version. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is covered. The description adds real behavioral context beyond them by disclosing that Call Test steps are flattened/expanded, which changes the shape of the returned step list. It stops short of noting any recursion depth, performance, or unresolvable-reference caveats.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with zero padding, front-loading the resource and the one behavior that matters. Nothing is repeated from the title or schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be explained, and annotations plus a fully-described schema cover the rest. The only gap is the lack of routing guidance versus xray_get_test, which would make this a standalone definition.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%: test, fields, and version_id are all documented in the schema itself. The description adds no parameter-level detail (e.g. how fields interacts with the expanded steps), so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Get) and resource (Test) with a distinguishing behavioral detail: 'Call Test' steps are expanded inline into the called tests' steps. That expansion is exactly what separates this from the sibling xray_get_test, though the description never names that sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied by the expansion behavior – an agent can infer 'use this when you want Call Test steps resolved', but there is no explicit when-to-use, when-not-to-use, or pointer to xray_get_test as the non-expanded alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_get_folderXray Get FolderARead-onlyIdempotent
Get a folder with its counts and complete subfolder tree. Use xray_get_tests(folder_path=...) for its Tests.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | Folder path, e.g. '/Regression/API'. '/' is the root. | / |
| project | No | Jira project key or id (for Test Repository folders). | |
| test_plan | No | Test Plan key or id (for Test Plan folders instead of the repository). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so safety is covered. The description adds genuinely behavioral detail beyond that: the result includes counts and the entire recursive subfolder tree, which tells the agent this is a deep traversal rather than a shallow listing. It does not mention pagination or size limits for large repositories.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero waste. The core capability is front-loaded before the sibling routing sentence, so an agent skimming the opening clause already knows what the tool returns.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be spelled out; the description still usefully flags that the subfolder tree is complete, which is the key recursion property. Missing only guidance on project vs test_plan selection, which the schema descriptions already handle adequately.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, with all three parameters documented in the schema including the project vs test_plan distinction, so the baseline is 3. The description contributes only the folder_path usage hint via the sibling reference, adding no syntax or format detail beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get a folder') plus its scope ('with its counts and complete subfolder tree'), which distinguishes it from the mutation siblings like xray_create_folder and xray_move_folder. It also routes to xray_get_tests for the folder's Tests, though 'counts' is left unqualified (counts of what?).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly names the alternative (xray_get_tests with folder_path=...) and the condition that selects it, so the agent knows where to go when it wants a folder's Tests rather than the folder itself. No exclusions or guidance on choosing between project and test_plan modes for folder scope.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_get_import_tests_statusXray Get Import Tests StatusARead-onlyIdempotent
Get status, progress and per-test results/errors of a bulk test import job.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | jobId returned by xray_import_tests_bulk. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and non-destructive, covering the safety profile. The description adds the useful behavioral detail that it surfaces progress and per-test errors, but with an output schema present the shape of those returns is partly already documented.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that names the resource and the payload categories with no wasted words. Optimal length for a one-parameter status tool.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Because an output schema exists, the description need not explain return values; it correctly points at the job_id source. It is complete for calling the tool, though it could note polling/retry expectations for a job-status endpoint.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter, documented at 100% coverage in the schema with its origin (jobId from xray_import_tests_bulk). The description does not add format, type, or retrieval semantics beyond what the schema already provides, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Get) and resource (status/progress of a bulk test import job), and explicitly enumerates what is returned (progress, per-test results/errors). It is immediately distinguishable from the many get_* siblings because it scopes to an import job rather than a test, plan, or run.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The schema links job_id to the output of xray_import_tests_bulk, which implies this is a follow-up/polling call, but neither description nor schema states when to call it (frequency, after queuing) or contrasts it with other alternatives. Usage is implied rather than spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_get_issue_historyXray Get Issue HistoryARead-onlyIdempotent
Get the Xray change history of an issue: who changed what (steps, definition, links...) and when (paginated).
| Name | Required | Description | Default |
|---|---|---|---|
| issue | Yes | Jira issue key (e.g. 'CALC-12') or numeric issue id. | |
| limit | No | Page size (1-100). | |
| start | No | Page offset (index of the first result). | |
| issue_type | Yes | Xray issue type of the issue. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, lowering the bar. The description adds genuine behavioral context beyond that: the result is paginated and the history covers who/what/when including steps, definition and links.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One tight sentence that front-loads the resource and ends with the pagination qualifier. Every clause earns its place; nothing is redundant with the title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return-value detail is unnecessary, and annotations cover the safety profile. The description is nearly sufficient, though it could note the auth/scope requirement or that history may be empty for certain issue types.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters (issue, issue_type, limit, start) are fully documented with types, bounds, defaults and an enum. The description adds no parameter-level syntax or format detail, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (Get) plus precise resource (Xray change history of an issue) and an enumerated scope of what the history contains (steps, definition, links) and when. No sibling tool covers issue change history, so it is trivially distinguishable from the get_* family.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The purpose implies when to reach for it (auditing changes to an issue), but there is no explicit when-to-use/when-not guidance and no alternatives are named. Adequate but inferential.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_get_issue_link_typesXray Get Issue Link TypesARead-onlyIdempotent
List the Jira issue link types known to Xray (used for requirement coverage).
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is fully covered by structured data. The description adds only that the list is 'known to Xray' (i.e., an Xray-scoped subset rather than all Jira link types), which is a mild behavioral nuance. With an output schema present, return values need no further explanation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. Every clause (verb, resource, scope, purpose) earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only lookup with full annotations and an output schema, the description supplies everything an agent needs to decide to call it. No gaps remain that more text would usefully close.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4. Schema coverage is also 100% with an empty properties object, meaning there is nothing for the description to compensate for.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb ('List') and resource ('Jira issue link types known to Xray') plus a scoping qualifier ('used for requirement coverage'). This is clearly distinguishable from the many test/precondition/plan/run siblings and from the other lookup tools like xray_get_statuses.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase '(used for requirement coverage)' implies the context in which link types matter, but there is no explicit when-to-use, when-not-to-use, or alternative named. Usage is implied rather than stated, which is the minimum viable level.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_get_preconditionXray Get PreconditionBRead-onlyIdempotent
Get a Precondition with its type, definition, folder and the Tests it is linked to.
| Name | Required | Description | Default |
|---|---|---|---|
| fields | No | Jira fields to return in 'jira'. Default: ['key', 'summary', 'status', 'issuetype', 'project', 'assignee', 'labels']. | |
| precondition | Yes | Jira issue key (e.g. 'CALC-12') or numeric issue id. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true and destructiveHint=false, so the safety profile is covered. The description adds the useful detail that the result includes the precondition's type, definition, folder and linked Tests, but says nothing about error behavior for a missing/invalid key or the size of the linked-Test set.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the resource and the payload it returns are stated immediately with zero redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, the description need not describe the return shape, and the safety profile is carried by annotations, so what remains is adequate for a 2-parameter read tool. The only shortfall is the absence of any disambiguation from the sibling list tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% – both the precondition key/id format and the optional 'fields' Jira field list are fully documented in the schema. The description contributes no additional parameter meaning, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Get') and resource ('a Precondition') and enumerates the returned facets (type, definition, folder, linked Tests), which is more informative than the title alone. However, it never distinguishes itself from the sibling xray_get_preconditions (plural list tool), leaving the singular/plural routing to be inferred from the name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no stated prerequisites, and no mention of alternatives such as xray_get_preconditions for listing or xray_get_expanded_test for deeper traversal. The agent must infer that this is the single-entity fetch from the name and parameter shape.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_get_preconditionsXray Get PreconditionsARead-onlyIdempotent
Search Preconditions by JQL, keys, project or type (paginated).
| Name | Required | Description | Default |
|---|---|---|---|
| jql | No | JQL filter. | |
| limit | No | Page size (1-100). | |
| start | No | Page offset (index of the first result). | |
| fields | No | Jira fields to return in 'jira'. Default: ['key', 'summary', 'status', 'issuetype', 'project', 'assignee', 'labels']. | |
| project | No | Jira project key or id. | |
| preconditions | No | Restrict to these keys or ids. | |
| modified_since | No | Only return issues modified after this date (ISO 8601, e.g. '2026-01-31'). | |
| precondition_type | No | Precondition type name, e.g. 'Manual', 'Cucumber', 'Generic'. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, openWorld and non-destructive, so the safety profile is covered without the description. The description adds only the pagination trait, which the limit/start parameters already convey in the schema, and says nothing about return shape or result ordering.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler: verb, resource, filter dimensions and pagination trait in that order. Nothing could be removed without losing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and 100% parameter coverage plus a full annotation set, the description only needs to frame the search scope, which it does. It is close to complete, missing only a note on the singular-vs-list sibling relationship.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all eight parameters including defaults and the fields list; baseline is 3. The description's filter enumeration only restates what the schema supplies, adding no format or syntax detail beyond it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Search Preconditions') and enumerates the filter axes (JQL, keys, project, type) plus pagination, which is far more informative than the plural name alone. It reasonably separates itself from the singular xray_get_precondition, though it never explicitly states the list-vs-single distinction.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The enumerated filters imply when the tool is useful (you have a JQL, key list, project or type to search by), but there is no explicit guidance on when to prefer xray_get_precondition for a single known key, nor any exclusions or prerequisites. Usage is implied rather than stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_get_project_settingsXray Get Project SettingsARead-onlyIdempotent
Get Xray project settings: test types, test environments, coverage, step fields, test run custom fields.
| Name | Required | Description | Default |
|---|---|---|---|
| projects | Yes | Jira project keys or ids. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false and openWorldHint=true, so the safety and repeat-call profile is fully covered. The description adds only the set of configuration categories returned; it says nothing about auth requirements, rate limits, or whether settings are merged across multiple projects.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the verb and resource, with the enumerated contents doing useful work rather than padding. Nothing extraneous.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and annotations cover the safety profile. For a simple one-parameter read tool this is nearly sufficient, though a hint about valid project identifiers or when to prefer this over GraphQL-based queries would close the remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single 'projects' array parameter is documented as 'Jira project keys or ids'. The description adds no meaning beyond the schema, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Get) and resource (Xray project settings), then enumerates exactly what the settings comprise: test types, test environments, coverage, step fields, test run custom fields. No sibling tool retrieves project-level settings, so the agent can distinguish it immediately.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites (e.g., that a valid Jira project key is needed), and no mention of alternatives. Usage is only implied by the tool name and the listed return contents.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_get_statusesXray Get StatusesARead-onlyIdempotent
List the test run statuses and step statuses available (globally or for a project).
| Name | Required | Description | Default |
|---|---|---|---|
| project | No | Jira project key or id; omit for the global configuration. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds the global-vs-project scope distinction, which is useful, but discloses nothing further about behavior. With annotations carrying the load, a 3 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence with no wasted words; the scope qualifier is compactly parenthesized. Nothing extraneous, though it is terse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
A single optional parameter, full annotation coverage, and an output schema mean little more needs saying. The description covers the essential scope nuance, so an agent can invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the single 'project' parameter is fully documented in the schema, including the omit-for-global semantics. The description's mention of 'globally or for a project' restates rather than extends that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (test run statuses and step statuses), which is clear enough to distinguish it from the surrounding test/run CRUD siblings. It does not explicitly name an alternative tool, but the resource is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical '(globally or for a project)' implies the two usage modes, and the schema clarifies that omitting project yields the global configuration. There is no explicit when/when-not guidance or naming of alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_get_step_libraryXray Get Step LibraryCRead-onlyIdempotent
Search the reusable manual or BDD (Gherkin) step library.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | 'manual' step library or 'bdd' (Gherkin) step library. | manual |
| test | No | Only steps used by this Test (key or id). | |
| limit | No | Page size (1-100). | |
| start | No | Page offset (index of the first result). | |
| search | No | Full-text search on the steps. | |
| project | No | Jira project key or id; omit for the global configuration. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, openWorld, so the safety profile is covered. The description adds nothing further: no mention of pagination behavior, what happens with no filters applied, or the global vs project-scoped configuration distinction present in the schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no waste. It is efficient but arguably too terse for a six-parameter search tool, which keeps it just under a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and all parameters are documented in the schema. However, for a tool with six optional filters, the description never clarifies the default scoping behavior (all steps vs project-scoped) or how filters compose, leaving a modest gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all six parameters are already documented in the schema (kind, test, limit, start, search, project). The description adds no extra semantic detail beyond what the schema provides, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb ('Search') and resource ('the reusable manual or BDD (Gherkin) step library'), and the kind enum reinforces the two modes. It does not distinguish this from sibling tools like xray_get_test or xray_add_test_step, so it stops short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No statement of when to reach for this tool versus related siblings (e.g. xray_get_test for a specific test's steps, or the add/update step tools). The two library kinds are implied rather than framed as usage guidance, leaving the agent to infer selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_get_testXray Get TestBRead-onlyIdempotent
Get one Xray Test with its type, steps / Gherkin / unstructured definition, folder and Jira fields.
| Name | Required | Description | Default |
|---|---|---|---|
| test | Yes | Jira issue key (e.g. 'CALC-12') or numeric issue id. | |
| fields | No | Jira fields to return in 'jira'. Default: ['key', 'summary', 'status', 'issuetype', 'project', 'assignee', 'labels']. | |
| include_relations | No | Also list linked Preconditions, Test Sets, Test Plans and Test Executions. | |
| include_requirements | No | Also list the requirements (coverable issues) each Test covers. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is covered. The description adds useful content-level context by naming the definition variants returned (steps/Gherkin/unstructured) and the folder/Jira field payload, but says nothing about errors, permissions or missing-test behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that leads with the verb and resource and then enumerates the returned content. No filler, no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need not be re-explained, annotations cover the safety profile, and all four parameters are schema-documented. The description is sufficient to call the tool, with the only real gap being the lack of routing guidance against the near-identical expanded/versions siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so 'test', 'fields', 'include_relations' and 'include_requirements' are already fully documented in the schema. The description adds no syntax, format or default information beyond what the schema provides, which is the expected baseline when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Get one Xray Test') and enumerates what the payload contains: type, steps/Gherkin/unstructured definition, folder and Jira fields. The word 'one' implicitly separates it from the plural sibling xray_get_tests, but it never addresses xray_get_expanded_test, which an agent would likely confuse it with.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus xray_get_expanded_test, xray_get_test_versions or xray_get_tests. The single-item framing gives a weak implied context (fetch one test by key), but no explicit conditions or exclusions are offered.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_get_test_executionXray Get Test ExecutionBRead-onlyIdempotent
Get a Test Execution with environments, Test Plans and its test runs (test + status).
| Name | Required | Description | Default |
|---|---|---|---|
| fields | No | Jira fields to return in 'jira'. Default: ['key', 'summary', 'status', 'issuetype', 'project', 'assignee', 'labels']. | |
| runs_limit | No | Max test runs to include. | |
| runs_start | No | Offset into the test runs. | |
| test_execution | Yes | Jira issue key (e.g. 'CALC-12') or numeric issue id. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is covered. The description adds useful payload context (it returns environments, Test Plans, and test runs), but says nothing about pagination behavior beyond the schema, authentication, or performance cost.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that states the resource and its returned contents with zero filler. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values needn't be fully explained, and annotations carry the safety profile. The description covers the resource and its main contents, though it could note how runs pagination interacts with the runs_limit/runs_start parameters.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters (test_execution key/id, fields, runs_limit, runs_start). The description adds no syntax, format, or default detail beyond what the schema provides, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb (Get) and resource (Test Execution) and enumerates the returned sub-resources (environments, Test Plans, test runs with test + status). The singular name plus 'a Test Execution' distinguishes it from the plural sibling xray_get_test_executions, though this distinction is implicit rather than stated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance, no mention of alternatives like xray_get_test_executions or xray_get_test_run, and no stated prerequisites. The agent can infer it fetches a single execution by key/id, but nothing is spelled out.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_get_test_executionsXray Get Test ExecutionsARead-onlyIdempotent
Search Test Executions by JQL, keys or project (paginated).
| Name | Required | Description | Default |
|---|---|---|---|
| jql | No | JQL filter. | |
| limit | No | Page size (1-100). | |
| start | No | Page offset (index of the first result). | |
| fields | No | Jira fields to return in 'jira'. Default: ['key', 'summary', 'status', 'issuetype', 'project', 'assignee', 'labels']. | |
| project | No | Jira project key or id. | |
| modified_since | No | Only return issues modified after this date (ISO 8601, e.g. '2026-01-31'). | |
| test_executions | No | Restrict to these keys or ids. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is fully covered. The description adds that results are paginated, which is useful context, but it does not disclose authentication needs, rate limits, or other behavioral traits beyond what annotations and the output schema provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with zero filler, front-loading the search action, the resource, and the accepted query forms. It is appropriately sized for a paginated list tool whose parameters and return shape are defined elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given full schema coverage, an output schema, and annotations that cover the operation's safety profile, the description provides enough orientation to call the tool correctly. It could be improved by explicitly differentiating from the singular retrieval sibling or mentioning filtering options like modified_since, but it is complete enough for a straightforward search endpoint.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all seven parameters are already documented in the input schema. The description mentions JQL, keys, and project, which maps to a subset of parameters, but it adds no syntax or format detail beyond what the schema already provides. A baseline 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('Search') and resource ('Test Executions'), and lists the primary query modes (JQL, keys, project) plus pagination. It is clear what the tool does, but it does not explicitly distinguish itself from the singular sibling xray_get_test_execution or other list tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage for bulk search and filtering by JQL, keys, or project, but it gives no explicit when-to-use guidance, no exclusions, and no named alternatives to compare against siblings. The context is clear enough to infer a list/search operation, but guidance is minimal.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_get_test_planXray Get Test PlanBRead-onlyIdempotent
Get a Test Plan with its Tests, Test Executions and folder structure.
| Name | Required | Description | Default |
|---|---|---|---|
| fields | No | Jira fields to return in 'jira'. Default: ['key', 'summary', 'status', 'issuetype', 'project', 'assignee', 'labels']. | |
| test_plan | Yes | Jira issue key (e.g. 'CALC-12') or numeric issue id. | |
| include_test_status | No | Include each Test's consolidated status within this plan. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the safety profile is covered. The description adds only that the result bundles Tests, Executions and folders — useful, but the output schema already conveys the shape, so incremental value is modest.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence that front-loads the verb and resource. No filler, though it is arguably terse given the tool's compositional behavior.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With a full input schema, a rich annotation set, and an output schema, the description only needs to frame scope — which it does by noting the plan is returned with its Tests, Executions and folders. Adequate for an agent to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema documents all three parameters (fields, test_plan, include_test_status) with defaults and formats. The description adds no parameter meaning beyond that, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Get') and resource ('Test Plan'), plus what the response composes (Tests, Test Executions, folder structure). This distinguishes it from the list-style xray_get_test_plans and the singular xray_get_test, though those siblings are not named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no statement of when to use this tool versus xray_get_test_plans or xray_get_test_executions, nor any prerequisites. The agent must infer selection from the name and resource alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_get_test_plansXray Get Test PlansBRead-onlyIdempotent
Search Test Plans by JQL, keys or project (paginated).
| Name | Required | Description | Default |
|---|---|---|---|
| jql | No | JQL filter. | |
| limit | No | Page size (1-100). | |
| start | No | Page offset (index of the first result). | |
| fields | No | Jira fields to return in 'jira'. Default: ['key', 'summary', 'status', 'issuetype', 'project', 'assignee', 'labels']. | |
| project | No | Jira project key or id. | |
| test_plans | No | Restrict to these keys or ids. | |
| modified_since | No | Only return issues modified after this date (ISO 8601, e.g. '2026-01-31'). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is fully covered structurally. The description's only behavioral addition is 'paginated', which the limit/start schema parameters already convey, so it adds little beyond the structured data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One compact sentence with the resource and filter modes front-loaded and zero filler. It is appropriately sized, though the parenthetical '(paginated)' is somewhat redundant with the schema.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need no explanation, and annotations cover safety. However, for a 7-parameter collection tool sitting among many similarly named get_* tools, the description does not explain how jql, project, and test_plans interact or when to prefer one filter over another, leaving a modest gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and each parameter (jql, project, test_plans, modified_since, fields, limit, start) is documented in the schema itself. The description adds no syntax or precedence detail beyond what the schema already provides, which is the baseline case.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Search) and resource (Test Plans) with the filter dimensions (JQL, keys, project) and pagination. It is clearly a bulk/collection query, contrasting with the singular xray_get_test_plan, though it never names that sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Search' implies list/multi-result usage, which distinguishes it loosely from single-entity siblings, but there is no explicit when-to-use or when-not-to-use guidance and no mention of how it relates to xray_get_test_plan or the other get_* variants.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_get_test_progressXray Get Test ProgressARead-onlyIdempotent
Count the Tests of a Test Plan or Test Execution per status, e.g. to answer 'how far along is it?'.
| Name | Required | Description | Default |
|---|---|---|---|
| max_tests | No | Stop counting after this many Tests (one API call per 100). | |
| test_plan | No | Test Plan key or id: counts each Test's consolidated status in the plan. | |
| environment | No | With test_plan: status in this Test Environment. | |
| list_statuses | No | List the Tests with these statuses (up to 100). Default: ['FAILED']. | |
| test_execution | No | Test Execution key or id: counts its test runs by status. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is fully covered and the description need not repeat it. The description adds the aggregation-by-status behavior, but says nothing about pagination, the API-call cost of max_tests, or how test_plan vs test_execution scope interact (both appear optional) — useful context that is left to the schema. Adequate with annotations carrying the load, hence a 3.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler: the action, the target resources, and the aggregation key are all stated before the optional example. Nothing could be trimmed without losing meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is unnecessary, and annotations cover the safety profile. What remains slightly under-specified is whether test_plan and test_execution are mutually exclusive or combinable and what happens when neither is supplied (all five params optional). For a read-only counting tool this is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and every parameter has a substantive description (including the default ['FAILED'], the environment-with-test_plan coupling, and the one-call-per-100 cost), so the schema does the heavy lifting. The description only loosely implies that test_plan or test_execution is the scope selector, adding no format or exclusivity detail beyond the schema. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb ('Count') and resource ('Tests of a Test Plan or Test Execution') with the aggregation dimension ('per status'), which unambiguously separates it from siblings like xray_get_tests (list) and xray_get_test_runs (raw runs). An agent can distinguish its role without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase "e.g. to answer 'how far along is it?'" gives a clear motivating use case, implying progress/status reporting. However, it never states when to prefer this over xray_get_test_plans, xray_get_tests, or xray_get_test_runs, and offers no exclusions or prerequisites. Usage is implied rather than specified.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_get_test_runXray Get Test RunARead-onlyIdempotent
Get one test run in full detail (steps, evidence, defects, examples, iterations), by id or by Test + Test Execution.
| Name | Required | Description | Default |
|---|---|---|---|
| test | No | Test key or id (with test_execution). | |
| test_run_id | No | Test run id. | |
| test_execution | No | Test Execution key or id (with test). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint, and destructiveHint=false, so the safety profile is fully covered structurally. The description adds only 'in full detail' as behavioral context; it says nothing about pagination, permission requirements, or behavior when the id is unknown.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the detail enumeration and the two identification modes are packed efficiently with no wasted clauses.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With a full output schema and complete annotations, the description only needs to identify the resource and how to target it, both of which it does. Minor gap: it never clarifies that all three parameters are optional or that exactly one of the two modes must be supplied.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and the schema descriptions already state '(with test)' and '(with test_execution)'. The description restates the same id-or-(Test + Test Execution) pairing, adding no new syntax, format, or constraint detail. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Get) and resource (one test run) and enumerates the detail returned (steps, evidence, defects, examples, iterations). It implicitly contrasts with the plural xray_get_test_runs via 'one test run', but never names a sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives the two valid lookup modes (by id, or by Test + Test Execution), which is useful invocation context. However, it offers no guidance on when to prefer this over xray_get_test_runs or xray_get_test_execution, and no exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_get_test_runsXray Get Test RunsBRead-onlyIdempotent
Search test runs by Tests, Test Executions, Test Plans, assignees or statuses (paginated).
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Page size (1-100). | |
| start | No | Page offset (index of the first result). | |
| tests | No | Filter by Test keys or ids. | |
| statuses | No | Filter by status names, e.g. ['FAILED']. | |
| assignees | No | Filter by assignee Atlassian account ids. | |
| test_plans | No | Filter by Test Plan keys or ids. | |
| test_run_ids | No | Fetch exactly these test run ids. | |
| include_steps | No | Include step results and evidence. | |
| modified_since | No | Only return issues modified after this date (ISO 8601, e.g. '2026-01-31'). | |
| test_executions | No | Filter by Test Execution keys or ids. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint and openWorldHint, so the safety profile is covered. The description adds the useful note that results are paginated, but says nothing about how pagination interacts with the limit/start params or about result ordering. Adequate but thin beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler, and the pagination caveat is appended compactly. Nothing is wasted, though it is arguably under-specified rather than maximally concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present the description needn't explain return values, and the safety profile is in annotations. But for a 10-parameter search tool it omits how filters combine (AND vs OR), whether test_run_ids overrides other filters, and result ordering, leaving moderate gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% across all 10 parameters, so the schema already documents each filter, its format, and defaults. The description only echoes the filter axes and adds no syntax, combining semantics, or note that test_run_ids bypasses the other filters. Baseline 3 is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb (Search) plus resource (test runs) and enumerates the five filter dimensions, so an agent can distinguish it from xray_get_test_run (single run) and xray_get_test_executions (different resource). It never names those siblings explicitly, so differentiation relies on inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no mention of when to prefer xray_get_test_run or xray_get_test_progress instead, and no statement about combining filters. Usage is only implied by the filter list.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_get_testsXray Get TestsARead-onlyIdempotent
Search Xray Tests by JQL, keys, project, test type or Test Repository folder (paginated).
| Name | Required | Description | Default |
|---|---|---|---|
| jql | No | JQL filter, e.g. 'project = CALC AND labels = smoke'. | |
| limit | No | Page size (1-100). | |
| start | No | Page offset (index of the first result). | |
| tests | No | Restrict to these Test keys or ids. | |
| fields | No | Jira fields to return in 'jira'. Default: ['key', 'summary', 'status', 'issuetype', 'project', 'assignee', 'labels']. | |
| project | No | Restrict to a Jira project (key or id). | |
| test_type | No | Restrict to a test type name, e.g. 'Manual'. | |
| folder_path | No | Restrict to a Test Repository folder path, e.g. '/Regression/API'. | |
| include_steps | No | Include manual steps of each test. | |
| modified_since | No | Only return issues modified after this date (ISO 8601, e.g. '2026-01-31'). | |
| expand_call_tests | No | Include steps with 'Call Test' steps replaced by the called Tests' steps (implies include_steps). | |
| include_definition | No | Include Gherkin / unstructured definitions of each test. | |
| include_subfolders | No | With folder_path: include descendant folders. | |
| include_requirements | No | Also list the requirements (coverable issues) each Test covers. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, destructiveHint=false and openWorldHint, so the safety profile is fully covered. The description adds only the pagination trait, and says nothing about the expandable payload behavior the tool offers.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the search verb, the resource, the filter facets and the pagination caveat all appear in one pass.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 14-parameter read tool with an output schema and rich annotations, the description is minimally adequate: it covers the search entry points but gives no hint that results can be enriched via include_steps, include_definition, expand_call_tests or include_requirements, which are meaningful behavioral choices an agent faces here.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and every parameter carries its own description (JQL example, page size/offset, field list default, folder path format), so the schema does the heavy lifting. The description's facet list largely restates schema content rather than adding syntax or interaction detail.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Search) and resource (Xray Tests) plus the facets it can filter on, which cleanly separates it from the singular sibling xray_get_test and xray_get_expanded_test. It stops short of explicitly naming which sibling to use for a single-test lookup.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the plural 'Search ... (paginated)' phrasing, which signals a list/multi-result retrieval, but the description never states when to prefer this over xray_get_test or xray_get_expanded_test, nor any prerequisite or exclusion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_get_test_setXray Get Test SetBRead-onlyIdempotent
Get a Test Set with the Tests it contains.
| Name | Required | Description | Default |
|---|---|---|---|
| fields | No | Jira fields to return in 'jira'. Default: ['key', 'summary', 'status', 'issuetype', 'project', 'assignee', 'labels']. | |
| test_set | Yes | Jira issue key (e.g. 'CALC-12') or numeric issue id. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false, so the safety profile is covered. The description adds useful retrieval context by stating that the returned Test Set includes the Tests it contains. It does not add details about permissions, rate limits, or output shape, but with rich annotations and an output schema, this is a reasonable contribution.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words. It efficiently states the action and the key returned relationship. For a simple retrieval tool, this is appropriately sized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only retrieval tool with full schema coverage, rich annotations, and an output schema, the description is largely complete. It conveys the core return relationship and avoids redundant return-value documentation. The main gap is the absence of when-to-use guidance relative to similar getter siblings.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both parameters are fully documented in the input schema. The description adds no parameter-level detail beyond what the schema already provides. Per the calibration rule, a baseline of 3 is appropriate when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Get a Test Set with the Tests it contains.' It is clear that this retrieves a single Test Set including its contained Tests, distinguishing it implicitly from the list-oriented sibling xray_get_test_sets. However, it does not explicitly call out how it differs from siblings such as xray_get_test or xray_get_expanded_test.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no explicit guidance on when to use this tool versus alternatives. It does not mention the relevant siblings xray_get_test_sets, xray_get_test, or xray_get_expanded_test, nor does it state prerequisites or exclusions. Usage is only implied by the verb 'Get'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_get_test_setsXray Get Test SetsBRead-onlyIdempotent
Search Test Sets by JQL, keys or project (paginated).
| Name | Required | Description | Default |
|---|---|---|---|
| jql | No | JQL filter. | |
| limit | No | Page size (1-100). | |
| start | No | Page offset (index of the first result). | |
| fields | No | Jira fields to return in 'jira'. Default: ['key', 'summary', 'status', 'issuetype', 'project', 'assignee', 'labels']. | |
| project | No | Jira project key or id. | |
| test_sets | No | Restrict to these keys or ids. | |
| modified_since | No | Only return issues modified after this date (ISO 8601, e.g. '2026-01-31'). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, so the safety profile is fully covered. The description adds only 'paginated' as behavioral context, which is useful but minimal; it says nothing about defaults, result caps, or ordering.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with the verb and scoping constraints front-loaded and zero filler. Nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists and all parameters are documented in the schema, so the description need not explain return values. However, for a flexible multi-filter search tool with seven optional parameters, the absence of any usage context leaves a gap that the structured fields do not fill.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all seven parameters (jql, project, test_sets, fields, modified_since, limit, start) are documented in the schema. The description merely names three of the filter concepts already covered, adding no syntax or format detail beyond the schema. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Search) and resource (Test Sets) plus the three filter modes (JQL, keys, project) and notes pagination. This distinguishes it reasonably well from the singular xray_get_test_set sibling, though it never explicitly names an alternative.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance and no routing to siblings such as xray_get_test_set (single) or xray_get_tests. The agent must infer that this is the bulk/plural listing tool from the name and description alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_get_test_versionsXray Get Test VersionsARead-onlyIdempotent
List the versions of a Test (ids for the version_id parameters of the other tools; paginated).
| Name | Required | Description | Default |
|---|---|---|---|
| test | Yes | Jira issue key (e.g. 'CALC-12') or numeric issue id. | |
| limit | No | Page size (1-100). | |
| start | No | Page offset (index of the first result). | |
| include_archived | No | Also list archived versions. | |
| include_definition | No | Include steps and Gherkin / unstructured definition of each version. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveness=false, so the safety profile is covered. The description adds that results are paginated and that they feed version_id parameters, which is useful but brief. It doesn't explain pagination semantics or mention include_archived/include_definition behavior beyond what the schema says.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single tight sentence in parentheses, front-loaded with the action and resource. No wasted words, though it is very terse and could carry slightly more structure without bloat.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values needn't be explained, and annotations cover safety. The one notable gap is that pagination behavior isn't detailed, but for a read-only list tool with full schema coverage, the description is nearly complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters are fully documented in the schema. The description adds only the conceptual link that returned ids map to version_id parameters elsewhere, which is context rather than parameter syntax. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (List) and resource (versions of a Test), clearly distinguishable from sibling xray_get_test which returns the test itself. The parenthetical explaining the version_id purpose adds useful routing context, though it doesn't explicitly differentiate from other get tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The parenthetical hints at when to use it (to obtain version_id values for other tools), which is implicit guidance. But there's no explicit when-not-to-use or named alternative within the large sibling set.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_graphql_mutationXray Graphql MutationB
Run an Xray GraphQL mutation not covered by a dedicated tool.
| Name | Required | Description | Default |
|---|---|---|---|
| mutation | Yes | GraphQL mutation document. | |
| variables | No | GraphQL variables. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the full safety profile (readOnlyHint=false, openWorldHint=true, idempotentHint=false, destructiveHint=false), so the description's burden is reduced. But it adds no behavioral context whatsoever: no auth/permission notes, no statement that arbitrary mutations can alter or remove data, and no error semantics for an escape-hatch tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler or repetition. It is efficient, though the extreme terseness borders on under-specification for a generic mutation entry point.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and 100% param coverage, the essentials are covered by structured fields. For a generic GraphQL escape hatch, however, the description omits useful navigational context — e.g., that the schema can be discovered via xray_graphql_schema or that queries should use xray_graphql_query — leaving a gap for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% (both mutation and variables are documented inline), so the schema carries the burden and a 3 baseline applies. The description adds no format, syntax, or variable-binding guidance beyond what the schema already states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ("Run an Xray GraphQL mutation") and scopes it with "not covered by a dedicated tool," which meaningfully separates it from the many dedicated xray_create_*/update_*/delete_* siblings. It is clear but does not name the closest sibling (xray_graphql_query) it complements.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
"Not covered by a dedicated tool" implies a fallback condition, so the agent can infer when to reach for it rather than a purpose-built mutation tool. However, it names no explicit alternative, gives no exclusions, and never mentions the sibling xray_graphql_query or xray_graphql_schema, leaving selection to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_graphql_queryXray Graphql QueryARead-onlyIdempotent
Run a read-only Xray GraphQL query. Entities are addressed by issue id; every list needs limit (max 100).
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | GraphQL query document (no mutations). | |
| variables | No | GraphQL variables. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, openWorldHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description repeats the read-only nature and adds the important operational rule that lists require a limit of at most 100, but it does not describe auth needs, rate limits, or error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler. The read-only restriction is front-loaded, followed immediately by the two most useful query-construction constraints.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and annotations cover the safety profile. The description supplies the key query-construction constraints, but it does not clarify when to use this over xray_graphql_schema or xray_graphql_mutation, which would make it fully complete for a complex GraphQL tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the schema already documents both parameters, setting a baseline of 3. The description goes beyond the schema by specifying that entities are addressed by issue id and that every list needs a limit of max 100, which are concrete constraints for building the query parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: run a read-only Xray GraphQL query. It distinguishes itself from the sibling xray_graphql_mutation via 'read-only' and the schema's 'no mutations' constraint, so an agent can identify its role immediately.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage by saying it is read-only and by enforcing 'every list needs limit (max 100)', which helps construct a valid query. However, it does not explicitly say when to prefer this tool over xray_graphql_schema or xray_graphql_mutation, leaving alternatives to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_graphql_schemaXray Graphql SchemaBRead-onlyIdempotent
Look up the Xray GraphQL schema: signatures and field docs of queries, mutations and types.
| Name | Required | Description | Default |
|---|---|---|---|
| names | No | Operation or type names, e.g. ['getTests', 'TestResults', 'createTest']. Omit to list all queries and mutations. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, openWorldHint and destructiveHint=false, covering the safety profile. The description adds only that the result contains signatures and field docs, which is largely return-value information already served by the output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One front-loaded sentence with no waste: the resource and the returned artifact are stated immediately. It is arguably too terse to carry routing information, but nothing is padded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return format needs no explanation, and annotations cover safety. The remaining gap is contextual: in a 70+ sibling toolset, the description omits the key workflow relation to xray_graphql_query and xray_graphql_mutation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single 'names' parameter is already fully documented in the schema, including the example array and the 'omit to list all' behavior. The description adds no parameter meaning of its own, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('look up') and resource ('the Xray GraphQL schema') plus what it returns (signatures, field docs of queries, mutations, types). It is distinguishable from the many entity-CRUD siblings, though it never distinguishes itself from the closely related xray_graphql_query/xray_graphql_mutation tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied — an agent can infer this is the discovery step before calling xray_graphql_query or xray_graphql_mutation, but the description never says so, nor when to omit names vs. pass them (that hint lives only in the schema).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_import_cucumber_featuresXray Import Cucumber FeaturesB
Create or update Cucumber Tests and Preconditions from .feature files (single file or zip).
| Name | Required | Description | Default |
|---|---|---|---|
| source | No | Name of the import source, used to match previously imported scenarios. | |
| feature | No | Content of a single .feature file. | |
| project | Yes | Jira project key or id. | |
| filename | No | File name, e.g. 'login.feature'. | |
| test_info | No | Jira payload applied to Tests created by the import, e.g. {"fields": {"labels": ["auto"]}}. | |
| zip_base64 | No | Base64 zip containing several .feature files. | |
| precondition_info | No | Jira payload applied to created Preconditions. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already state readOnlyHint=false, destructiveHint=false, openWorldHint=true, and idempotentHint=false, so the agent knows this is a non-idempotent write to an open world; the description's 'create or update' is consistent with that. It adds the zip-vs-single-file input dimension but does not explain update matching semantics, what happens to existing scenarios, or permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence that states the action, the resulting entities, and the accepted input forms with zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and the annotations cover the safety profile. For a 7-parameter bulk import with scenario-matching behavior, however, more context on how updates are identified (via source) and which payloads apply to Tests vs Preconditions would be expected.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all 7 parameters (including source, test_info, precondition_info) are already documented inline. The description only echoes the single-file/zip distinction already carried by the feature and zip_base64 schema entries, adding no syntax or format detail beyond that.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb pair (create or update) and resource (Cucumber Tests and Preconditions), plus the accepted input format (.feature single file or zip). It is clearly distinguishable from sibling importers like xray_import_tests_bulk and xray_import_execution_results_json, though it never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to choose this over xray_import_tests_bulk or xray_import_execution_results_*, and no mention of prerequisites such as project permissions. Usage is only implied by the name and the 'Cucumber' scope.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_import_execution_results_jsonXray Import Execution Results JsonA
Import JSON test results into Xray, creating or updating a Test Execution.
Xray JSON can target an existing execution via "testExecutionKey" or describe a new one in "info". Pass test_execution_info to control all Jira fields of the created Test Execution (multipart import).
| Name | Required | Description | Default |
|---|---|---|---|
| format | No | 'xray' = Xray JSON, 'cucumber' = Cucumber JSON report, 'behave' = Behave JSON report. | xray |
| results | Yes | The results document (JSON object/array or its text). | |
| test_execution_info | No | Jira issue payload for the Test Execution created by the import, e.g. {"fields": {"project": {"key": "CALC"}, "summary": "Nightly run", "issuetype": {"name": "Test Execution"}}}. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare the write/safety profile (readOnly=false, destructive=false, idempotent=false, openWorld=true), so the description only needs to add beyond that. It usefully discloses create-or-update behavior, the testExecutionKey/info targeting mechanism, and that test_execution_info works via multipart import. It does not warn about non-idempotent re-imports (duplicate executions), a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the purpose and scoping follow-ups. Every sentence carries information; no padding.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation. The description covers the main behavioral nuance (create vs update, multipart info). The one inconsistency is that the name/description say "JSON" while the format enum accepts cucumber and behave reports, which an agent must infer from the schema alone.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents format, results, and test_execution_info. The description reinforces test_execution_info's role but adds no syntax or format detail beyond the schema. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Import JSON test results into Xray") and adds scope ("creating or updating a Test Execution"). The JSON framing distinguishes it from xray_import_execution_results_xml, though it doesn't name that sibling explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Describes the create-vs-update targeting conditions ("target an existing execution via testExecutionKey or describe a new one in info"), which is useful usage context. However, it never states when to choose this tool over the xml/cucumber import siblings or when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_import_execution_results_xmlXray Import Execution Results XmlA
Import an XML test report (JUnit, TestNG, NUnit, xUnit, Robot Framework) into Xray.
Either use the simple parameters (project_key, test_execution_key, ...), or pass test_execution_info / test_info for full control over the created issues (multipart import).
| Name | Required | Description | Default |
|---|---|---|---|
| format | Yes | Report format. | |
| results | Yes | The XML report content. | |
| revision | No | Source code revision. | |
| test_info | No | Jira payload applied to Tests created by the import, e.g. {"fields": {"labels": ["auto"]}}. | |
| fix_version | No | Fix version of the Test Execution. | |
| project_key | No | Project for a new Test Execution. | |
| test_plan_key | No | Link the Test Execution to this Test Plan. | |
| test_environments | No | Test Environment names. | |
| test_execution_key | No | Import into this existing Test Execution. | |
| test_execution_info | No | Jira issue payload for the Test Execution created by the import, e.g. {"fields": {"project": {"key": "CALC"}, "summary": "Nightly run", "issuetype": {"name": "Test Execution"}}}. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=false, destructiveHint=false and openWorldHint=true, so the safety profile is covered. The description adds that a 'multipart import' mode exists and that issues are 'created' by the import, but it omits practical consequences such as duplicate Test/Test Execution creation on re-import or auth/permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, no filler. The core purpose and the supported formats come first, and the mode guidance follows – appropriately front-loaded and easy to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and annotations cover the safety profile. The description supplies the mode distinction, but could note side effects like creating Tests/Test Executions and the non-idempotent re-import risk for a write-heavy import tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description goes further by grouping the 10 parameters into two coherent usage modes and clarifying that test_execution_info/test_info are for full control over created issues. That framing meaningfully helps an agent choose a parameter combination.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (import) and resource (XML test report into Xray) and enumerates the supported formats (JUnit, TestNG, NUnit, xUnit, Robot). This cleanly distinguishes it from the JSON sibling import tool and the Cucumber import tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly describes two invocation modes: the 'simple parameters' path (project_key, test_execution_key, ...) versus the 'full control' multipart path with test_execution_info/test_info. It doesn't name the trade-offs (e.g. exactly when to prefer full control) or directly contrast with xray_import_execution_results_json, but the mode selection guidance is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_import_tests_bulkXray Import Tests BulkA
Start an asynchronous bulk import of Tests. Poll xray_get_import_tests_status with the returned jobId.
| Name | Required | Description | Default |
|---|---|---|---|
| tests | Yes | Tests to create/update (max 1000). Each: {'testtype': 'Manual'|'Cucumber'|'Generic', 'fields': {'summary': ..., 'project': {'key': ...}}, 'steps': [{'action','data','result'}], 'gherkin_def': ..., 'unstructured_def': ..., 'xray_test_repository_folder': '/path', 'xray_test_sets': [...], 'xray_preconditions': [...], 'update': {...}}. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=false), so the bar is lower. The description still adds real value by disclosing that the operation is asynchronous and that it returns a jobId to be polled, which the annotations do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and then the required follow-up step. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and the async polling requirement is captured. Only minor gaps remain, such as how partial failures or the 1000-item limit are reported back to the caller.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single 'tests' parameter is documented in exhaustive detail in the schema itself. The description adds nothing about parameter shape, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (import), resource (Tests) and scope (bulk, asynchronous). The word 'bulk' implicitly separates it from the single-record sibling xray_create_test, but no sibling is named explicitly, so it falls short of a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a follow-up workflow (poll xray_get_import_tests_status with the returned jobId), which is useful operational guidance, but never states when to prefer this bulk tool over xray_create_test or the other import variants (cucumber, execution results). Usage is implied rather than framed as a choice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_list_toolsetsXray List ToolsetsARead-onlyIdempotent
List all Xray toolsets (read and write halves), which tools are enabled, and whether the server is read-only.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and openWorldHint=false, so safety is covered. The description's only added behavioral hint is the 'read and write halves' toolset structure and that read-only status is reported, both of which largely restate deliverable content the output schema already covers.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single well-formed sentence with the verb and scope front-loaded and no filler. Nothing is wasted or buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters, rich annotations, and an output schema handling return details, the description is sufficient for correct invocation. It could have added a one-line routing hint in a 70+ tool catalog, which is the only real gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the baseline is 4; there is nothing for the description to clarify, and schema coverage is trivially 100%.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a concrete verb and resource ('List all Xray toolsets') plus the specific content returned ('which tools are enabled, and whether the server is read-only'). It is unambiguous, but it never contrasts itself with siblings (no other toolset-listing tool exists, yet no routing hint is given in this crowded catalog).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is only implied by the return content: an agent can infer it should call this to discover toolsets and enabled tools. There is no explicit when-to-use statement, no prerequisite, and no mention of alternatives or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_move_folderXray Move FolderBIdempotent
Move a folder (with its content) below another folder.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Folder path, e.g. '/Regression/API'. '/' is the root. | |
| index | No | Position among sibling folders / issues. | |
| project | No | Jira project key or id (for Test Repository folders). | |
| test_plan | No | Test Plan key or id (for Test Plan folders instead of the repository). | |
| destination_path | Yes | Parent folder to move into, e.g. '/Archive'. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare non-readOnly, idempotent, non-destructive, open-world. The description adds one genuinely useful behavioral fact — that the folder's contents move with it — but omits what happens when the destination is missing, name collisions, and whether cross-project/test-plan moves are supported.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the parenthetical placement of the content-move behavior is well positioned. It is arguably too terse for a five-parameter mutation, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With annotations and an output schema present, the description need not cover return values. However, for a move operation spanning repository and test-plan folder scopes (per the project/test_plan params), it says nothing about destination existence, ordering via index, or failure behavior, leaving real gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all five parameters including the ambiguous index, project, and test_plan fields are documented in the schema. The description adds no parameter meaning beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Move a folder') and adds the important scope detail that contents move along with it. It does not explicitly distinguish itself from the sibling xray_move_to_folder, which sounds similar but moves issues/tests into a folder, so the agent must infer the difference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this versus xray_move_to_folder, xray_remove_from_folder, or xray_rename_folder. The use case is implied by the verb but no context, prerequisites, or exclusions are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_move_to_folderXray Move To FolderAIdempotent
Move Tests (and, in the Test Repository, Preconditions) into a folder.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Folder path, e.g. '/Regression/API'. '/' is the root. | |
| index | No | Position among sibling folders / issues. | |
| tests | No | Test keys or ids. | |
| project | No | Jira project key or id (for Test Repository folders). | |
| test_plan | No | Test Plan key or id (for Test Plan folders instead of the repository). | |
| preconditions | No | Precondition keys or ids (repository folders only). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already disclose the safety profile (readOnlyHint=false, destructiveHint=false, idempotentHint=true, openWorldHint=true). The description adds one useful behavioral constraint—that Preconditions can only be moved in the Test Repository—but does not elaborate on permissions, side effects, or idempotency behavior beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with no wasted words. It communicates the core action and scope efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has six parameters, an output schema, and rich annotations, the description is minimally adequate. It omits how tests can be moved into Test Plan folders (covered only by the schema), and lacks usage guidance, though the schema and annotations carry much of the burden.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all six parameters are fully documented in the schema. The description does not add meaning beyond what the schema already provides (e.g., path format, index semantics, test_plan vs repository context).
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Move'), the resources being moved ('Tests' and, conditionally, 'Preconditions'), and the destination ('into a folder'). This clearly distinguishes the tool from siblings like xray_move_folder (which moves folders) and xray_remove_from_folder (which removes items from folders).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no explicit guidance on when to use this tool versus alternatives, nor does it state prerequisites, exclusions, or when-not-to-use conditions. The parenthetical about Preconditions being limited to the Test Repository provides some scope context, but no broader usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_remove_all_test_stepsXray Remove All Test StepsBDestructiveIdempotent
Remove all steps from a manual Test.
| Name | Required | Description | Default |
|---|---|---|---|
| test | Yes | Jira issue key (e.g. 'CALC-12') or numeric issue id. | |
| version_id | No | Test version id; default is the current version. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true, readOnlyHint=false and openWorldHint=true, so the safety profile is covered. The description adds one useful scoping fact – steps only exist on manual Tests – but says nothing about irreversibility, whether the operation is scoped to a version, or what remains after removal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with the action and scope front-loaded and zero filler. Nothing to trim.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple two-parameter destructive operation with full annotation coverage and an output schema, the description supplies enough to call it correctly, and the 'manual Test' qualifier is a meaningful constraint. It stops just short of stating irreversibility or the version-scoping behavior an agent might want to know before wiping all steps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters (test, version_id) are documented in the schema, including accepted formats and the current-version default. The description adds no parameter meaning beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Remove all steps from a manual Test'), and the scope word 'all' implicitly distinguishes it from the sibling xray_remove_test_step, which removes a single step. However, it never names that sibling explicitly, so the differentiation is left to the agent to infer from the name alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance, no mention of prerequisites, and no pointer to the closely related xray_remove_test_step, xray_update_test_step, or xray_add_test_step siblings. Usage is only implied by the tool name and the phrase 'manual Test'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_remove_from_folderXray Remove From FolderAIdempotent
Take Tests / Preconditions out of their folder. The issues themselves are not deleted.
| Name | Required | Description | Default |
|---|---|---|---|
| tests | No | Test keys or ids. | |
| project | No | Jira project key or id (for Test Repository folders). | |
| test_plan | No | Test Plan key or id (for Test Plan folders instead of the repository). | |
| preconditions | No | Precondition keys or ids (repository folders only). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (destructiveHint=false, idempotentHint=true), and the description adds genuinely useful context beyond them: 'The issues themselves are not deleted,' clarifying that the entities survive the operation. It still omits permission requirements or side effects on folder contents, but for a simple relationship-removal tool this is solid added value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and immediately followed by the key non-destructive clarification. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a four-optional-parameter removal tool with an output schema (so return values need not be explained) and covered annotations, the description conveys the essential action and outcome. Its only real gap is the absence of usage/routing guidance, which is scored separately.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all four parameters (tests, project, test_plan, preconditions). The description only alludes to Tests and Preconditions and adds no syntax or format detail beyond the schema, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (remove) and resource (Tests/Preconditions from their folder), so the action is unambiguous. It does not differentiate from close siblings like xray_move_to_folder or xray_delete_folder, so an agent must infer the inverse relationship.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance, and no alternatives are named. It never clarifies how this differs from xray_move_to_folder (move into a folder) or xray_delete_folder, leaving the agent to infer routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_remove_test_associationsXray Remove Test AssociationsADestructiveIdempotent
Unlink one Test from Preconditions, Test Sets, Test Plans and/or Test Executions.
| Name | Required | Description | Default |
|---|---|---|---|
| test | Yes | Jira issue key (e.g. 'CALC-12') or numeric issue id. | |
| test_sets | No | Issue keys or ids. | |
| test_plans | No | Issue keys or ids. | |
| version_id | No | Test version, for preconditions only. | |
| preconditions | No | Issue keys or ids. | |
| test_executions | No | Issue keys or ids. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true, readOnlyHint=false and openWorldHint=true, so the safety profile is covered. The description usefully clarifies that only the association is severed ('Unlink'), not the Test itself, but says nothing about irreversibility, required permissions, or what happens when no container list is supplied.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the scope (one Test, four container types) is delivered before any detail. Nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a destructive unlink mutation with full annotation coverage, a rich schema and an output schema, the description covers what is unlinked from what. Remaining gaps are minor: behavior when all container lists are omitted, and whether removal is atomic across the listed containers.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so every parameter (test key/id, test_sets, test_plans, preconditions, test_executions, version_id) is already documented in the schema, placing the baseline at 3. The description only adds the implied combinability of the four association lists and does not clarify the precondition-only nature of version_id beyond what the schema states.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Unlink') and enumerates exactly which resources are affected (Preconditions, Test Sets, Test Plans, Test Executions) for a single Test. This distinguishes it from the look-alike bulk siblings, though it never names them explicitly (e.g. xray_remove_tests_from_test_set), so an agent must infer the difference from the singular 'one Test'.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'and/or' phrasing implies the container lists are optional and combinable in a single call, which hints at usage. However, there is no explicit when-to-use statement, no prerequisite (e.g. must the test exist / be linked), and no routing to the plural per-container alternatives like xray_remove_tests_from_test_set.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_remove_test_environmentsXray Remove Test EnvironmentsCDestructiveIdempotent
Remove Test Environments from a Test Execution.
| Name | Required | Description | Default |
|---|---|---|---|
| test_execution | Yes | Jira issue key (e.g. 'CALC-12') or numeric issue id. | |
| test_environments | Yes | Test Environment names, e.g. ['Chrome', 'staging']. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true, and openWorldHint=true, so the safety profile is covered. The description adds nothing beyond restating 'Remove'—it does not explain reversibility, permission requirements, or what happens to the execution's state after removal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single efficient sentence with the action front-loaded and zero filler. It is appropriately sized for the tool's simplicity, though it borders on under-specification rather than being optimally informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema exists, so return values need not be described, and annotations cover the safety profile. However, the description omits any operational context (e.g., when to use it versus add_test_environments, whether removal is scoped per execution), leaving a minor gap for an agent choosing between related tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and both parameters are fully documented in the schema, including issue key/id format and example environment names. The description adds no parameter-level detail, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Remove) and resource (Test Environments from a Test Execution). It is clear what the tool does, though it does not explicitly name or contrast with the sibling xray_add_test_environments, which would have earned a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives, no prerequisites, and no mention of the complementary add operation. The agent must infer usage from the tool name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_remove_test_executions_from_test_planXray Remove Test Executions From Test PlanCDestructiveIdempotent
Remove Test Executions from a Test Plan.
| Name | Required | Description | Default |
|---|---|---|---|
| test_plan | Yes | Jira issue key (e.g. 'CALC-12') or numeric issue id. | |
| test_executions | Yes | Jira issue keys (e.g. 'CALC-12') or numeric issue ids. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly=false, destructive=true, idempotent=true, and openWorld=true, so the safety profile is covered externally. The description adds no behavioral context such as irreversibility or effect on linked executions. With no added disclosure, a 2 is appropriate.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler. It is appropriately sized for a simple removal action, though it is essentially a title restatement. A 4 reflects concision with minimal added structure.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema, rich annotations, and complete parameter docs cover most call requirements. The description still lacks usage guidance and sibling differentiation, especially against xray_remove_tests_from_test_plan. A 3 reflects adequate but incomplete context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and both parameters are well documented in the input schema. The description provides no additional parameter meaning, so the baseline score of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Remove Test Executions from a Test Plan'), so the action is clear. However, it adds nothing beyond the name/title and does not differentiate from siblings like xray_remove_tests_from_test_plan. A 4 is appropriate because it is clear but lacks sibling-specific routing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, prerequisites, or alternatives are provided. The agent must infer usage entirely from the name and schema. This is 'no guidance' rather than misleading, so 2.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_remove_test_run_evidenceXray Remove Test Run EvidenceBDestructiveIdempotent
Remove evidence from a test run by id and/or filename.
| Name | Required | Description | Default |
|---|---|---|---|
| test_run_id | Yes | Test run id, from xray_get_test_execution or xray_get_test_runs. | |
| evidence_ids | No | ||
| evidence_filenames | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false, so the safety profile is covered. The description adds the id/filename selection clause but says nothing about permanence, behavior when an id matches nothing, or whether passing both selectors is allowed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero filler; verb, resource, and selectors are all stated immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists so return values need no explanation, and annotations carry the destructive/idempotent profile. Still, for a destructive mutation the description leaves unclear what happens on partial matches or how id and filename selectors interact.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%: evidence_ids and evidence_filenames have no schema descriptions. The phrase 'by id and/or filename' does explain that either or both selectors may drive the removal, which is real added meaning, but it doesn't clarify array semantics, matching behavior, or mutual exclusion.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (remove), resource (evidence), and scope (from a test run), with the selectors named. It is clearly the inverse of the sibling xray_add_test_run_evidence, though it doesn't name that sibling to make the pairing explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, and no reference to alternatives such as xray_add_test_run_evidence or xray_get_test_runs for finding evidence. The agent must infer context entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_remove_tests_from_preconditionXray Remove Tests From PreconditionBDestructiveIdempotent
Unlink Tests from a Precondition.
| Name | Required | Description | Default |
|---|---|---|---|
| tests | Yes | Jira issue keys (e.g. 'CALC-12') or numeric issue ids. | |
| precondition | Yes | Jira issue key (e.g. 'CALC-12') or numeric issue id. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true, and openWorldHint=true, so the safety profile is covered. The description adds no behavioral context beyond restating the removal, such as what happens to missing links, permission requirements, or whether only the association is deleted.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single front-loaded sentence with zero filler. It is efficiently sized for a simple unlink operation, though it offers no structural detail beyond the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present and annotations covering safety, the description need not explain return values or mutation risk. However, it remains minimal for a destructive association-removal tool and does not clarify operational caveats or edge cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% and both parameters clearly describe accepted Jira issue keys or numeric ids. The description adds no parameter meaning beyond the schema, so the baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description 'Unlink Tests from a Precondition' states a specific verb and both resources involved, and the verb 'Unlink' cleanly distinguishes this tool from its sibling xray_add_tests_to_precondition. An agent can identify the operation without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, prerequisites, or alternatives are given. The name implies usage, but the description never says when an agent should remove associations versus adding them or handling other precondition mutations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_remove_tests_from_test_executionXray Remove Tests From Test ExecutionADestructiveIdempotent
Remove Tests from a Test Execution, deleting their test runs.
| Name | Required | Description | Default |
|---|---|---|---|
| tests | Yes | Jira issue keys (e.g. 'CALC-12') or numeric issue ids. | |
| test_execution | Yes | Jira issue key (e.g. 'CALC-12') or numeric issue id. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true and idempotentHint=true, so safety is covered. The description adds a specific consequence beyond the annotations: deleting the associated test runs. This is useful context, though it stops short of covering irreversibility or permission requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words. It immediately communicates the action and its key side effect.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that an output schema exists and annotations cover the safety profile, the description only needs to add behavioral context. It does so by mentioning test-run deletion, but it could be slightly more complete by noting side effects on the execution's status or confirming irreversibility.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents both parameters. The description adds no additional meaning about parameter formats or constraints, making the baseline score of 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Remove'), resource ('Tests'), and target scope ('from a Test Execution'), and it explicitly notes the deletion of test runs. This clearly distinguishes it from sibling operations like add_tests_to_test_execution or remove_tests_from_test_set.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description offers no explicit when-to-use, when-not-to-use, or alternative guidance. It merely restates the action, which is insufficient given the many similar remove_* siblings. An agent gets no help choosing between this and, for example, remove_test_associations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_remove_tests_from_test_planXray Remove Tests From Test PlanCDestructiveIdempotent
Remove Tests from a Test Plan.
| Name | Required | Description | Default |
|---|---|---|---|
| tests | Yes | Jira issue keys (e.g. 'CALC-12') or numeric issue ids. | |
| test_plan | Yes | Jira issue key (e.g. 'CALC-12') or numeric issue id. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare destructiveHint=true, idempotentHint=true, readOnlyHint=false, and openWorldHint=true, so the safety profile is covered. The description adds no behavioral context beyond restating the removal action — it does not say what gets deleted, whether permissions are required, or what side effects occur. With annotations already carrying the safety information, the description's lack of any additional context is a clear gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no wasted words. However, it merely echoes the title and does not earn its place by adding any actionable information. It is concise but under-informative rather than efficiently structured.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple removal tool with a 100% documented schema and rich annotations, the minimal description is survivable. However, it provides no routing guidance among many similar sibling tools and no caution beyond what the annotations already state. The structured data compensates enough that the description is not wholly inadequate, but clear gaps remain.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents both parameters (test_plan and tests) with examples and accepted formats. The description adds no parameter meaning beyond what the schema provides, which matches the baseline score of 3 when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description restates the tool name almost verbatim: 'Remove Tests from a Test Plan.' It states a specific verb and resource, but it adds no distinguishing information beyond the title and does not differentiate from sibling tools such as xray_remove_tests_from_test_set or xray_remove_tests_from_precondition. This fits the rubric's 'tautology' category more than a genuinely informative purpose statement.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no indication of when to use this tool versus alternatives. It does not mention the related 'add_tests_to_test_plan' tool, nor does it clarify that it is for removing existing tests rather than test executions (a distinction covered by the sibling 'remove_test_executions_from_test_plan'). The agent is left to infer usage entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_remove_tests_from_test_setXray Remove Tests From Test SetCDestructiveIdempotent
Remove Tests from a Test Set.
| Name | Required | Description | Default |
|---|---|---|---|
| tests | Yes | Jira issue keys (e.g. 'CALC-12') or numeric issue ids. | |
| test_set | Yes | Jira issue key (e.g. 'CALC-12') or numeric issue id. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false, so the safety profile is covered. The description adds nothing beyond that — no note that the tests themselves are not deleted (only the association), no auth/permission requirements, and no indication of partial-failure behavior when a subset of keys is invalid.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with the verb and resource front-loaded and zero filler. It is appropriately sized, though the terseness comes at the cost of any contextual detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists and rich annotations cover the mutation profile, so return values and safety are not the description's burden. However, for a destructive, open-world mutation against a test set, the total absence of when-to-use guidance and of the crucial 'only unlinks, does not delete tests' clarification leaves a meaningful gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both 'tests' and 'test_set' documented as Jira issue keys or numeric ids. The description adds no parameter detail beyond what the schema already provides, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Remove') and resource ('Tests' from a 'Test Set'), which clearly distinguishes it from the sibling xray_add_tests_to_test_set. It does not, however, mention any of the other removal siblings (test plan, test execution, precondition), so differentiation is implicit rather than explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives, no prerequisites, and no mention of the counterpart add operation. The agent must infer usage purely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_remove_test_stepXray Remove Test StepBDestructiveIdempotent
Remove one step from a manual Test.
| Name | Required | Description | Default |
|---|---|---|---|
| step_id | Yes | Step id, from xray_get_test. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true, and openWorldHint=true, so the safety profile is covered. The description adds only the 'manual Test' scope constraint; it says nothing about irreversibility, permissions, or ordering effects on remaining steps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short, front-loaded sentence with zero filler. It is efficient, though arguably thin rather than optimally sized for a destructive operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema and rich annotations exist, so return values and safety need not be described. However, for a destructive tool in a family of step operations, the lack of routing guidance against xray_update_test_step and xray_remove_all_test_steps leaves a gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Single parameter with 100% schema description coverage ('Step id, from xray_get_test'), so the schema carries the semantics. The description adds no information about step_id beyond what is already documented.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: remove a single step from a manual Test. The word 'one' implicitly distinguishes it from xray_remove_all_test_steps, but no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this versus xray_update_test_step, xray_remove_all_test_steps, or how it relates to adding steps. The agent must infer the context from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_rename_folderXray Rename FolderCIdempotent
Rename a folder.
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Folder path, e.g. '/Regression/API'. '/' is the root. | |
| project | No | Jira project key or id (for Test Repository folders). | |
| new_name | Yes | New folder name (not a path). | |
| test_plan | No | Test Plan key or id (for Test Plan folders instead of the repository). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the full safety profile (readOnlyHint=false, idempotentHint=true, destructiveHint=false, openWorldHint=true), so the description is not obligated to repeat them. It adds no behavioral context of its own – nothing about how the rename resolves paths, whether it overwrites, or what happens across the repository/test plan variants.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with no waste, but its brevity comes at the cost of substance rather than being efficient compression of useful detail. It is front-loaded but says almost nothing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is a mutation with two distinct folder contexts (Test Repository via project, or Test Plan via test_plan), and the description never mentions this branching or any constraints. Annotations and the output schema cover safety and return values, but the description leaves the operational context too thin for a folder-renaming mutation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters including path, new_name, project, and test_plan are documented in the schema. The description contributes no additional parameter meaning, which matches the baseline 3 when the schema does the heavy lifting.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (Rename) and resource (folder), so the basic action is unambiguous. It does not, however, distinguish this tool from close siblings like xray_move_folder or xray_delete_folder, nor does it scope which folder types are affected.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use rename vs move_folder or move_to_folder, and no prerequisites such as required permissions or the repository-vs-test-plan distinction. The agent must infer usage entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_reset_test_runXray Reset Test RunADestructiveIdempotent
Reset a test run to its initial state, discarding status, comments, evidence and step results.
| Name | Required | Description | Default |
|---|---|---|---|
| test_run_id | Yes | Test run id, from xray_get_test_execution or xray_get_test_runs. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare destructiveHint=true, idempotentHint=true, and readOnlyHint=false, so the safety profile is covered. The description adds genuine value beyond that by enumerating exactly what is lost: status, comments, evidence, and step results. It stops short of saying whether the reset is reversible or what permissions are needed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One sentence, front-loaded with the action and immediately followed by the destructive consequence. No filler and nothing an agent needs to skip past.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-param mutation with rich annotations and an output schema, the description covers the essential risk (what data is discarded). The remaining gap is routing guidance relative to xray_update_test_run, which is a minor omission given the annotations already flag destructiveness.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter, and the schema already documents it at 100% coverage, including where to obtain the id (xray_get_test_execution or xray_get_test_runs). The description adds no parameter-level detail, so the schema carries the load and baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (reset) and resource (test run) plus the exact scope of the reset. It reads clearly against siblings like xray_update_test_run, but the description never names or contrasts with those alternatives explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied: the agent can infer this is the tool for wiping a run back to a clean state rather than making a targeted edit. There is no explicit when-to-use vs when-not, no mention of xray_update_test_run as the non-destructive alternative, and no prerequisite (e.g., completed runs) stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_set_test_run_timerXray Set Test Run TimerB
Start, pause or reset the execution timer of a test run.
| Name | Required | Description | Default |
|---|---|---|---|
| reset | No | True resets the timer. | |
| running | No | True starts, False pauses the timer. | |
| test_run_id | Yes | Test run id, from xray_get_test_execution or xray_get_test_runs. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=false, and openWorldHint=true, so the safety profile is covered. The description adds only the three timer modes (start/pause/reset), which are already implied by the parameters; it says nothing about permission requirements, whether an in-progress timer survives a pause, or reversibility.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
One short sentence, front-loaded with the action and resource, with no filler. It is perhaps too terse for the tool's branching behavior, but nothing in it is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists and all three parameters are documented, so return values need not be explained. However, for a mutation tool with branching modes, the absence of any guidance on mode combination or preconditions leaves the description at minimum viability.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% – reset, running, and test_run_id are each documented in the schema, including the source of the test run id. The description adds no syntax, default, or interaction detail (e.g., precedence when both reset and running are supplied) beyond what the schema already provides, so baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb set (start/pause/reset) and a precise resource (the execution timer of a test run), which distinguishes it from the nearby xray_reset_test_run and xray_update_test_run. It does not explicitly name those siblings, so an agent still has to infer the boundary, but the resource is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use context, no prerequisites (e.g., must the run be in a particular state, must the timer not already be running), and no routing away from xray_reset_test_run or xray_update_test_run. The agent is left to infer everything from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_update_preconditionXray Update PreconditionCIdempotent
Update type, definition and/or folder of a Precondition.
| Name | Required | Description | Default |
|---|---|---|---|
| definition | No | ||
| folder_path | No | Move to this Test Repository folder. | |
| precondition | Yes | Jira issue key (e.g. 'CALC-12') or numeric issue id. | |
| precondition_type | No | Precondition type name, e.g. 'Manual', 'Cucumber', 'Generic'. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the mutation profile (readOnlyHint=false, destructiveHint=false, idempotentHint=true), so the safety picture is covered. The description adds almost nothing beyond that: the 'and/or' slightly hints at partial updates, but it never says what happens to unspecified or null fields, nor what the update returns or whether the change is reversible.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler; the only weakness is that it is a fragment that never names the tool's subject beyond 'a Precondition,' leaving no room for guidance.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described, and the safety annotations cover the mutation profile. Still, for a mutation tool in a large family of near-identically named siblings, the definition omits partial-update semantics and any routing guidance, leaving it only minimally sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 75% and the schema already documents folder_path, precondition_type, and the precondition identifier. The description names the same three updatable fields but contributes no format, validation, or 'null means leave unchanged' semantics beyond what the structured fields provide, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Update) and resource (Precondition) and enumerates the mutable aspects (type, definition, folder), which cleanly separates it from the many xray_update_* siblings that target tests, test runs, or folders. It does not, however, explicitly contrast itself with those nearest siblings such as xray_update_test_definition.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no indication of when to reach for this tool versus xray_get_precondition, xray_create_precondition, or xray_delete_precondition, and states no prerequisites such as required Jira permissions. Usage must be inferred entirely from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_update_test_definitionXray Update Test DefinitionAIdempotent
Replace the Gherkin or unstructured definition of a Test. Use the step tools for manual tests.
| Name | Required | Description | Default |
|---|---|---|---|
| test | Yes | Jira issue key (e.g. 'CALC-12') or numeric issue id. | |
| gherkin | No | New Gherkin scenario (Cucumber tests). | |
| version_id | No | Test version id; default is the current version. | |
| unstructured | No | New free-text definition (Generic tests). |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=true, destructiveHint=false, so the safety profile is covered. The word 'Replace' usefully signals that the existing definition is overwritten, but the description adds nothing about permissions, what happens on partial input, or how versioning interacts with the update.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with zero filler, and the core action is front-loaded ahead of the routing hint. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is a mutation with an output schema and full parameter coverage, so the description need not explain returns. However, it omits whether gherkin and unstructured are mutually exclusive, how they relate to the test type, and version_id behavior, leaving meaningful gaps for an update tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents test, gherkin, version_id, and unstructured. The description's mention of 'Gherkin or unstructured' mirrors the schema rather than adding format details, mutual-exclusivity rules, or guidance on version_id. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Replace') and resource ('the Gherkin or unstructured definition of a Test'), so the agent knows exactly what is being modified. It does not name a sibling tool explicitly, but the phrase 'step tools' points away from the step-related siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides a clear when-not with an alternative category: 'Use the step tools for manual tests.' That routes the agent away from this tool for manual/step-based tests. It stops short of explicitly naming xray_update_test_step or telling the agent which of gherkin vs unstructured to pick.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_update_test_runXray Update Test RunCIdempotent
Update a test run: status, comment, timestamps, assignee, executor and custom fields.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | New status name, e.g. 'PASSED' or 'FAILED'. | |
| comment | No | ||
| started_on | No | ISO 8601 start timestamp. | |
| assignee_id | No | Atlassian account id of the assignee. | |
| finished_on | No | ISO 8601 finish timestamp. | |
| test_run_id | Yes | Test run id, from xray_get_test_execution or xray_get_test_runs. | |
| custom_fields | No | Test run custom fields. | |
| executed_by_id | No | Atlassian account id of the executor. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the safety profile (readOnlyHint=false, idempotentHint=true, destructiveHint=false, openWorldHint=true), so the description is not obliged to restate safety. It adds nothing else though: it never says whether omitted fields are left untouched (partial vs. full update), whether it needs write permission on the execution, or how custom fields merge.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence naming the verb, resource, and the field set, with no filler. It is perhaps too terse to be maximally useful, but nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value explanation is unnecessary, and annotations cover safety. Still, for an 8-parameter mutation the description omits the key operational fact an agent needs: whether this is a partial update (unspecified fields preserved) or a replacement, and how the custom_fields array interacts with existing values.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 88%, so the schema already documents nearly every parameter (status examples, ISO timestamps, Atlassian account ids, custom field shape). The description only re-lists the same field names in prose, adding no format or constraint detail beyond the schema. Baseline 3 is correct.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Update a test run') and enumerates the mutable fields, so the agent knows exactly what it changes. It does not, however, distinguish itself from close siblings like xray_update_test_run_step, xray_update_test_run_defects, or xray_reset_test_run.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no prerequisites, and no reference to the sibling tools that also mutate test runs. The agent must infer that this is the general-purpose run-update tool versus the more targeted step/defect/reset variants.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_update_test_run_defectsXray Update Test Run DefectsBIdempotent
Link and/or unlink defects on a test run.
| Name | Required | Description | Default |
|---|---|---|---|
| add | No | Defect issue keys or ids to link. | |
| remove | No | Defect issue keys or ids to unlink. | |
| test_run_id | Yes | Test run id, from xray_get_test_execution or xray_get_test_runs. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description's 'link and/or unlink' reinforces the mutation semantics but adds no context on permission requirements, duplicate handling, or what happens to existing defect links. Adequate but thin beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler; the 'and/or' phrasing efficiently conveys that both operations can occur in one call. It is on the terse side, but nothing is wasted or buried.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With an output schema present, return values need not be explained, and the input schema fully documents the three parameters. However, for a write tool that modifies defect associations, the description omits any statement about merge vs replace behavior or error conditions, leaving a real gap for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the definitions for add, remove, and test_run_id are already documented, including the source of the test_run_id. The description repeats the link/unlink mapping already implied by the schema without adding format or constraint detail, which is the expected baseline at full coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb pair (link/unlink) and a specific resource (defects on a test run), which cleanly separates it from adjacent siblings such as xray_add_test_run_evidence and xray_update_test_run. It stops short of naming any alternative tool, so it lands at a solid 4 rather than 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives, no note that 'add' and 'remove' may be supplied together, and no prerequisites or ordering advice. The agent must infer usage entirely from the name and schema.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_update_test_run_example_statusXray Update Test Run Example StatusBIdempotent
Set the status of one Cucumber Scenario Outline example in a test run.
| Name | Required | Description | Default |
|---|---|---|---|
| status | Yes | Status name as configured in Xray, e.g. 'PASSED', 'FAILED', 'TODO', 'EXECUTING'. | |
| example_id | Yes | Scenario Outline example id, from xray_get_test_run. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so the safety/idempotency profile is covered without the description. The description adds no further behavioral context (no error behavior, no effect on the parent run/progress counters, no permission notes), so it does not go beyond the structured data.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with the verb, scope, and unit of change, containing no redundant or filler text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a low-complexity two-parameter mutation with full schema coverage, an output schema, and annotations covering the safety profile, the description is minimally adequate. It omits any note on valid status transitions or what happens if the example_id does not belong to a run, which would help an agent call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both parameters documented (status naming convention and example_id provenance from xray_get_test_run). The description adds nothing beyond the schema, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Set the status') and a precisely scoped resource ('one Cucumber Scenario Outline example in a test run'). The word 'example' implicitly differentiates it from sibling tools like xray_update_test_run_iteration_status and xray_update_test_run_step, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus alternatives such as xray_update_test_run_iteration_status, nor any prerequisites (e.g. that the example must belong to an existing run). The agent must infer usage context from the name alone.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_update_test_run_iteration_statusXray Update Test Run Iteration StatusBIdempotent
Set the status of one iteration of a data-driven test run.
| Name | Required | Description | Default |
|---|---|---|---|
| status | Yes | Status name as configured in Xray, e.g. 'PASSED', 'FAILED', 'TODO', 'EXECUTING'. | |
| test_run_id | Yes | Test run id, from xray_get_test_execution or xray_get_test_runs. | |
| iteration_rank | Yes | Iteration rank, from xray_get_test_run. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=true, destructiveHint=false and openWorldHint=true, so safety and repeatability are covered. The description adds only the scoping fact that a single iteration is targeted; it says nothing about required permissions, what happens to the parent test run's overall status, or whether changes are reversible.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no redundant or filler text; the scope qualifier (one iteration of a data-driven test run) is exactly what matters.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and annotations cover the safety profile. However, for a mutation tool surrounded by several near-identical status/step siblings, the description is too thin to route the agent reliably and omits any note on how the iteration status relates to the test run's own status.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and the schema itself supplies examples for status and provenance hints for test_run_id and iteration_rank, so the baseline applies. The description contributes no additional parameter detail beyond what the schema already documents.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (Set), resource (status of one iteration of a data-driven test run) and scope, which is more precise than the bare tool name. It does not, however, differentiate itself from the closely-named sibling xray_update_test_run_example_status, which an agent could easily confuse with iteration status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this versus xray_update_test_run, xray_update_test_run_step, or xray_update_test_run_example_status. The context is implied by the phrase 'data-driven test run' but no alternatives or exclusions are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_update_test_run_stepXray Update Test Run StepBIdempotent
Record the result of one test run step: status, comment, actual result, defects and evidence.
| Name | Required | Description | Default |
|---|---|---|---|
| status | No | Step status name, e.g. 'PASSED'. | |
| comment | No | ||
| step_id | Yes | Step id of the test run step, from xray_get_test_run. | |
| add_defects | No | Defect keys or ids to link. | |
| test_run_id | Yes | Test run id, from xray_get_test_execution or xray_get_test_runs. | |
| add_evidence | No | ||
| actual_result | No | ||
| iteration_rank | No | Iteration rank for data-driven runs, from xray_get_test_run. | |
| remove_defects | No | Defect keys or ids to unlink. | |
| remove_evidence_ids | No | ||
| remove_evidence_filenames | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=true, destructiveHint=false, and openWorldHint=true, so safety is covered. The description adds that a step result is recorded, but does not disclose the incremental add/remove semantics implied by the schema or what happens on repeat calls.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no padding. Efficient, though the brevity is part of why the parameter and usage gaps go unfilled.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, but for a mutation tool with 11 parameters and partial schema coverage the description leaves the add/remove mechanics and iteration handling undocumented.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 55% across 11 parameters, and the description merely restates field names already present in the schema. It does not explain the add_/remove_ pairs, iteration_rank's role in data-driven runs, or the base64-vs-attachment_id choice for evidence.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Record the result of one test run step') and enumerates the fields it writes. This is enough to distinguish it from xray_update_test_run and xray_update_test_step, though it never names those siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no prerequisites, no mention of alternatives. With siblings like xray_update_test_run, xray_update_test_run_defects, and xray_add_test_run_evidence nearby, the agent gets no help deciding which tool applies.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_update_test_stepXray Update Test StepBIdempotent
Update action, data and/or expected result of a manual test step.
| Name | Required | Description | Default |
|---|---|---|---|
| data | No | ||
| action | No | ||
| result | No | ||
| step_id | Yes | Step id, from xray_get_test. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare the mutation is not read-only, not destructive, and idempotent. The description adds the useful detail that updates are partial ('and/or'), implying unspecified fields are untouched, but it does not explain permissions, whether step_id must already exist, or what happens on an unknown id.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single short sentence with the resource and fields front-loaded; nothing is wasted. It is arguably a touch terse given the gaps elsewhere.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. However, for a mutation tool with 25% parameter coverage, the definition omits step existence requirements and failure behavior, leaving the agent to guess how partial updates interact with omitted fields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 25%: step_id is documented in the schema, while data, action, and result have no schema descriptions. The description names those three fields and implies they are optional and independently updatable, which partially compensates for the gap but gives no format or length details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (update) and resource (a manual test step) plus the exact fields affected (action, data, expected result). It is distinguishable from the sibling xray_update_test_run_step, though it never says so explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use or when-not-to-use guidance, no prerequisites, and no mention of alternatives such as xray_add_test_step or xray_remove_test_step. Usage is left entirely to inference from the name.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_update_test_typeXray Update Test TypeBIdempotent
Change the test type of a Test (e.g. Manual -> Cucumber).
| Name | Required | Description | Default |
|---|---|---|---|
| test | Yes | Jira issue key (e.g. 'CALC-12') or numeric issue id. | |
| test_type | Yes | Xray test type name, e.g. 'Manual', 'Cucumber', 'Generic'. | |
| version_id | No | Test version id; default is the current version. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, destructiveHint=false, idempotentHint=true, and openWorldHint=true, so the safety profile is carried by structured fields. The description adds no behavioral context beyond the bare mutation purpose — it does not mention side effects, required permissions, or what happens to existing steps when the type changes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with zero wasted words. The core operation and an illustrative example are delivered efficiently.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a required-parameter mutation tool, the description plus rich annotations, full schema coverage, and an output schema make invocation clear. However, it omits usage guidance and sibling differentiation, which would help an agent navigate the large Xray toolset.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. The description's 'Manual -> Cucumber' example slightly reinforces the test_type values, but the schema already documents all three parameters with examples and defaults.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Change the test type of a Test', with a concrete example of the transformation. An agent can identify the operation, but the description does not explicitly differentiate it from nearby siblings such as xray_update_test_definition or xray_create_test.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no when-to-use guidance, no prerequisites, and no alternatives. It does not explain when to choose this tool over xray_update_test_definition or other test-editing siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
xray_upload_attachmentXray Upload AttachmentB
Upload a file; the returned id can be used as attachment_id in evidence / step attachments.
| Name | Required | Description | Default |
|---|---|---|---|
| text | No | File content as text (UTF-8), instead of base64. | |
| filename | Yes | ||
| mime_type | No | Content type, e.g. 'image/png'. | application/octet-stream |
| content_base64 | No | File content, base64 encoded. |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, idempotentHint=false, openWorldHint=true and destructiveHint=false, so the safety profile is covered. The description adds nothing behavioral beyond that — no size limits, no statement about whether re-uploading the same file creates duplicates, no auth requirements.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence with the action front-loaded and the downstream id usage immediately after. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value detail is unnecessary, and the description does mention the key returned id. But for a 4-parameter tool it omits the text-vs-base64 selection rule and any upload constraints, leaving the calling agent to infer them.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 75%, so the schema largely documents the parameters itself (text vs base64, mime_type default). The description adds no meaning about which of text/content_base64 to choose or that they are mutually exclusive, so it sits at the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a clear verb+resource ('Upload a file') and adds the downstream payoff: the returned id is usable as attachment_id in evidence / step attachments. That distinguishes it from xray_get_attachment without naming siblings explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It implies the workflow context (upload first, then reference the id when attaching evidence), which is usable guidance. However, it never names an alternative tool or states when not to use it, nor does it mention prerequisites such as project or auth scoping.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
82 tool updates
v0.1.0- First observed
xray_add_test_associations - First observed
xray_add_test_environments - First observed
xray_add_test_executions_to_test_plan - First observed
xray_add_test_run_evidence - First observed
xray_add_test_step - First observed
xray_add_tests_to_precondition - First observed
xray_add_tests_to_test_execution - First observed
xray_add_tests_to_test_plan - First observed
xray_add_tests_to_test_set - First observed
xray_create_folder - First observed
xray_create_precondition - First observed
xray_create_test - First observed
xray_create_test_execution - First observed
xray_create_test_plan - First observed
xray_create_test_set - First observed
xray_delete_folder - First observed
xray_delete_precondition - First observed
xray_delete_test - First observed
xray_delete_test_execution - First observed
xray_delete_test_plan - First observed
xray_delete_test_set - First observed
xray_export_cucumber_features - First observed
xray_get_attachment - First observed
xray_get_coverable_issue - First observed
xray_get_coverable_issues - First observed
xray_get_datasets - First observed
xray_get_expanded_test - First observed
xray_get_folder - First observed
xray_get_import_tests_status - First observed
xray_get_issue_history - First observed
xray_get_issue_link_types - First observed
xray_get_precondition - First observed
xray_get_preconditions - First observed
xray_get_project_settings - First observed
xray_get_statuses - First observed
xray_get_step_library - First observed
xray_get_test - First observed
xray_get_test_execution - First observed
xray_get_test_executions - First observed
xray_get_test_plan - First observed
xray_get_test_plans - First observed
xray_get_test_progress - First observed
xray_get_test_run - First observed
xray_get_test_runs - First observed
xray_get_test_set - First observed
xray_get_test_sets - First observed
xray_get_test_versions - First observed
xray_get_tests - First observed
xray_graphql_mutation - First observed
xray_graphql_query - First observed
xray_graphql_schema - First observed
xray_import_cucumber_features - First observed
xray_import_execution_results_json - First observed
xray_import_execution_results_xml - First observed
xray_import_tests_bulk - First observed
xray_list_toolsets - First observed
xray_move_folder - First observed
xray_move_to_folder - First observed
xray_remove_all_test_steps - First observed
xray_remove_from_folder - First observed
xray_remove_test_associations - First observed
xray_remove_test_environments - First observed
xray_remove_test_executions_from_test_plan - First observed
xray_remove_test_run_evidence - First observed
xray_remove_test_step - First observed
xray_remove_tests_from_precondition - First observed
xray_remove_tests_from_test_execution - First observed
xray_remove_tests_from_test_plan - First observed
xray_remove_tests_from_test_set - First observed
xray_rename_folder - First observed
xray_reset_test_run - First observed
xray_set_test_run_timer - First observed
xray_update_precondition - First observed
xray_update_test_definition - First observed
xray_update_test_run - First observed
xray_update_test_run_defects - First observed
xray_update_test_run_example_status - First observed
xray_update_test_run_iteration_status - First observed
xray_update_test_run_step - First observed
xray_update_test_step - First observed
xray_update_test_type - First observed
xray_upload_attachment
TDQS
Scored across 82 tools
Descriptions are generally precise, but the large surface creates several close pairs and groups: get_test/get_tests/get_expanded_test, get_test_run/get_test_runs, and the multiple test-run update variants. These boundaries are mostly explained, yet an agent can still misselect among similar read/update operations.
Tool names follow a highly consistent xray_verb_noun pattern across the entire set, including create_, get_, add_, remove_, update_, import_, and export_ actions. The convention is predictable and readable throughout.
At 82 tools, the set is far beyond a practical MCP surface and hits the extreme-mismatch threshold. Even for a broad domain like Xray test management, this volume makes discovery, selection, and maintenance unnecessarily difficult.
The surface covers the full Xray lifecycle: tests, preconditions, test sets, plans, executions, runs, folders, coverage, imports/exports, attachments, history, settings, and GraphQL escape hatches. There are no obvious dead ends for the stated domain.
Maintenance
Related MCP Connectors
Remote MCP for 1,500+ APIs. Vault-managed credentials; OAuth or API key. Search, load, and execute.
- mcpOAuthcom.vibgrate
Query your team's drift, vulnerability, and upgrade data from any AI assistant. OAuth 2.1, 51 tools.
AI-callable tools for API mocking, testing, monitoring, security, and automation.
Governed MCP gateway: one endpoint for your tools, with credential custody and audit log.
Related MCP Servers
- AlicenseBqualityDmaintenanceEnables integration with Xray Cloud APIs for comprehensive test management including creating and managing test cases, test executions, test plans, and test sets. Supports CI/CD automation and test result tracking through GraphQL APIs.2320 npm3ISC
- AlicenseBqualityDmaintenanceEnables comprehensive test management in Zephyr Scale Cloud, including creating and managing test cases, executing tests with step-by-step results, organizing test cycles and plans, and performing advanced JQL searches.1816 npmMIT
- FlicenseBqualityFmaintenanceEnables AI assistants to interact with Xray Test Management for both Cloud and Server deployments. Supports test execution, importing results from multiple formats (JUnit, Cucumber, Robot Framework, TestNG), and managing test plans and executions.83-
- AlicenseAqualityDmaintenanceEnables interaction with Xray Cloud and Data Center for test management, including authentication, GraphQL queries, test execution/plan management, and result import via MCP.98 npmMIT