Skip to main content
Glama

create_test_results_bulk

Record new test executions for several cases in one test run with a single call, returning created execution ids aligned to the input entries.

Instructions

Record NEW executions for several items of ONE test run in a single call (POST /testrun/{runKey}/testresults). The body is the results array itself; in each entry only the fields you pass are sent, and the returned ids are positionally aligned with it. The test case should already be an item of the run; if it is not, the behavior is VERSION-SPECIFIC — some Server builds silently ADD it to the run as a new item (verified live: testCaseCount grows; the new item's POSITION in items[] is not the head and not the tail — it landed second of three and second of four in two separate runs, so do not rely on where it appears), others reject the call with 400/404. Default statuses: 'Not Executed', 'In Progress', 'Pass', 'Fail', 'Blocked' — case-sensitive internal names; instances may define custom ones. scriptResults carry per-step outcomes of a STEP_BY_STEP script as { index (0-based), status, comment? }. An overall status sent TOGETHER with scriptResults is stored as sent (verified live: 'Blocked' with three 'Pass' steps stored 'Blocked') and is NEVER derived from the step statuses — scriptResults without a status leave the execution at the project default ('Not Executed'), so pass status in the SAME call. Some older builds may instead ignore the overall status: read the result back with get_test_run_results rather than sending a second update_last_test_result, which replaces the whole execution. A scriptResults entry whose index is past the last step of the case is discarded silently (HTTP 200, no error). matchEnvironment / matchUserKey apply to the WHOLE batch and only SELECT which existing run item to append to — they never set the created result's environment. NOT ATOMIC, and it does NOT abort at the failing entry: when one entry is rejected (unknown testCaseKey, bad status/iteration/version value, unknown custom field) the call answers an error and returns no ids, yet that entry alone is SKIPPED while EVERY other valid entry is COMMITTED — the entries AFTER the failing one just as much as those before it (verified live twice, with ids: [valid, valid, unknown key, valid] committed entries 1, 2 and 4; [unknown key, valid] committed entry 2). So after an error that names ONE entry the batch may already be fully written except for that entry: re-read the run with get_test_run_results BEFORE resending anything and resend only the entries that are genuinely missing — resending the batch or its tail creates DUPLICATE executions. An error that rejects the payload as a whole (a malformed body, an unknown run key, a batch-wide matchEnvironment/matchUserKey the API cannot resolve) is different: it is refused before any entry is processed, so nothing was committed and nothing needs re-reading. Two entries for the same test case create two independent executions (no upsert). Returns the array of created executions ([{ id }, …]).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultsYesOne entry per execution to record; each targets a run item by testCaseKey and carries that execution's fields
testRunKeyYesTest run (cycle) key, e.g. PROJ-R123 (PROJ-C123 on older instances)
matchUserKeyNoRun-item selector, sent as the 'userKey' QUERY parameter (never in the body): targets the run item by its executor's Jira user key, e.g. 'JIRAUSER10000'.
matchEnvironmentNoRun-item selector, sent as the 'environment' QUERY parameter (never in the body): targets the run item with this environment (case-sensitive). Distinct from the 'environment' body field, which sets the environment recorded on the result.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv1.0.5

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so exceptionally: it discloses non-atomic skip-and-commit semantics with concrete id-level evidence, version-specific behavior for non-member test cases, silent discarding of out-of-range scriptResults indices, status defaulting, batch-wide matchEnvironment/matchUserKey scoping, and no-upsert duplicate semantics. These are exactly the behaviors an agent cannot infer from the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

It is long, but the length is largely earned by the tool's genuine complexity and it is front-loaded with purpose before behavior. The repeated 'verified live' parentheticals with example id arrays add evidence but are verbose, and a few clauses could be tightened without losing meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, yet the description explicitly states the return shape ([{ id }, …]) and its positional alignment with the request. Combined with the mutation hazards, failure modes, and version caveats it covers, an agent has everything needed to call this safely and correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds real meaning beyond it: the returned ids are positionally aligned with the input array, only the fields passed per entry are sent, and matchEnvironment/matchUserKey select the run item for the whole batch rather than setting the created result's environment. The status/scriptResults interplay is also a semantic clarification the schema does not express.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a specific verb (Record), resource (NEW executions for several items of ONE test run), and the endpoint (POST /testrun/{runKey}/testresults). It is immediately distinguishable from the singular create_test_result and from update_last_test_result. An agent knows exactly what class of operation this is.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives strong operational routing: re-read with get_test_run_results instead of sending a second update_last_test_result, and re-read before resending after a partial failure. It does not explicitly contrast itself with create_test_result for single executions, so the selection boundary versus that sibling is left implicit, but recovery and read-back alternatives are named precisely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools