Skip to main content
Glama

mark_retryable

Flag a stuck running trial as retryable for infrastructure failures, not scientific ones, so terminal trials are recorded without counting as evidence.

Instructions

Mark a running trial as retryable (infrastructure failure).

This is for trials that are stuck in 'running' due to infrastructure issues (server restart, executor state lost, timeout) — NOT for scientific failures. A retryable trial is terminal and does not count as evidence.

It does NOT unlock re-capture or re-run: a terminal trial's record is finished. To retry the work, design_experiment a new trial and capture_bundle on it (file-path or code://).

Enforcement: commitment 1 — the loop is the unit (orphan check).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
reasonYesWhy the trial is marked retryable — becomes part of the record.
trial_idYesID of the target trial.
programme_idYesID of the target research programme.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
reasonNo
statusNo
trial_idNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.28

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations, so the description carries the full burden — and it delivers: it states the trial becomes terminal, does not count as evidence, does not unlock re-capture/re-run, and references an enforcement rule (commitment 1, orphan check). This is exactly the behavioral context an agent needs before an irreversible mutation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and the critical when/when-not distinction, then the do-not-unlock warning, then the enforcement note. Well organized, though slightly long and the trailing enforcement sentence is a bit cryptic without more context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained. For a mutation tool with no annotations, the description supplies everything an agent needs: what qualifies, what it excludes, the terminal/evidence semantics, the non-unlock caveat, and the correct alternative route.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters (programme_id, trial_id, reason). The description adds the useful detail that reason 'becomes part of the record', but no syntax or format guidance. Baseline 3 is correct when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb (mark) + resource (running trial as retryable) with explicit scoping: it names the exact condition (infrastructure failure: server restart, executor state lost, timeout) and distinguishes it from scientific failures. Among siblings like correct_trial_status and cancel_trial, the retryable-vs-scientific-failure distinction lets an agent pick it correctly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use (stuck-in-running infrastructure failures) and when-not (scientific failures), plus a named alternative path: design_experiment a new trial and capture_bundle on it. It even warns that this does NOT unlock re-capture/re-run, closing the main misuse route.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.