Skip to main content
Glama

correct_trial_status

Correct a terminal trial's recorded status when executor output contradicts the recorded outcome, appending an audited correction with a mandatory reason.

Instructions

Correct a terminal trial's status — the recorded repair act.

For mislabeled records: e.g. the executor's outer wrapper exited 0 while the recorded output documents a crash ("status": "error" / nonzero inner exit code in the payload). Source must be completed|failed|retryable (all terminal — a retryable source is a correction made in error or evidence re-read after the fact); target must be completed|failed|retryable. reason is mandatory. Correcting TO completed requires the record to evidence a completed run — a cancellation receipt, reaper note, or failure output does not qualify, so that direction is refused unless the executor record documents an actual completion.

The correction is appended to the trial's executor_output_json under 'corrections' — the record shows both what was claimed and what it was corrected to. Never hand-edit the database: a direct sqlite3 UPDATE bypasses this audit trail.

Enforcement: commitment 1 — the loop is the unit (orphan check).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
reasonYesWhy the correction is recorded — mandatory, appended to the trial's audit trail.
trial_idYesID of the target trial.
to_statusYesTarget status: completed | failed | retryable.
programme_idYesID of the target research programme.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
toNo
fromNo
trial_idNo
corrected_atNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.28

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden and does so: the correction is appended to executor_output_json under 'corrections', both claimed and corrected states are retained, and the direction-toward-completed refusal rule is disclosed. The 'reason' mandate and the ban on direct DB edits add concrete behavioral context an agent cannot infer from the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose is front-loaded in the first line, and the subsequent sentences all carry weight (valid statuses, refusal rule, audit destination, DB warning). It is somewhat verbose and the closing 'commitment 1 — the loop is the unit (orphan check)' line is opaque, keeping it just short of maximal.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, return values need not be described, and the description instead covers everything else an agent needs: valid status directions, mandatory reason, the refusal condition on correcting to completed, and the audit-trail side effect. Nothing required to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all four parameters and the baseline is 3. The description adds real meaning: it explicates that the source status is effectively terminal-only (completed|failed|retryable) and constrains to_status semantically rather than just enumerating it, plus reinforces that reason is mandatory and fed into the audit trail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with a specific verb+resource ('Correct a terminal trial's status') and characterizes it as 'the recorded repair act', which distinguishes it from siblings like mark_retryable, run_trial, and cancel_trial. An agent immediately understands this is an audited correction of an existing terminal status, not a status transition driver.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit motivating scenario (executor wrapper exited 0 while output documents a crash) and states the valid source/target status pairs. It also states a negative condition — correcting TO completed is refused unless the record evidences an actual completion — and warns against direct sqlite3 UPDATE. When-to-use, when-refused, and the alternative-to-avoid are all covered.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.