Skip to main content
Glama

run_trial

Execute a controlled experiment trial by invoking the executor with the bundle's training config; long jobs return running so you can poll status.

Instructions

Run a trial by calling the executor role.

Imports run_training from the bundle's code_ref and calls it with the trial config. The code_ref must be a path to a .py file exposing def run_training(config: dict) -> dict returning {"metrics": {...}, "variance": {...}}. Read the executor://contract resource for the full contract.

When the bundle carries data_refs, the config handed to run_training is extended with two injected keys: 'data_paths' ({split: resolved read-only path}) and 'data_ref_paths' ({data_ref_id: resolved read-only path}). The stored config_json keeps the designed config verbatim — the injection is runtime-only.

For long-running jobs, this returns quickly with status "running". Use get_trial_status to poll for completion. Note the return is not immediate — there is a consistent ~10 s handshake/settle before status "running" comes back; that settle window is what makes cancel-race behaviour reproducible.

Enforcement: commitment 6 — the bundle must be controlled. Rejects if the bundle is not fully captured. Concurrency: rejects if trial is already running (no double execution).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
trial_idYesID of the target trial.
programme_idYesID of the programme owning the trial — required; the programme_id returned by design_experiment/list_trials (orphan check). The programme must be status=active — completed/abandoned/archived programmes are closed, immutable records.
timeout_secondsNoPer-trial hard deadline override in seconds — the executor kills the process past it. None uses the server default ([executor] timeout_seconds). Bounded by [executor] max_timeout_seconds; the applied value is recorded in executor_output.timeout_seconds.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
errorNo
statusNo
messageNo
trial_idNo
executor_outputNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.28

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers: it discloses the runtime-only injection of 'data_paths'/'data_ref_paths' while config_json stays verbatim, the non-immediate ~10 s handshake/settle window, the enforcement rule (commitment 6, bundle must be controlled), and the no-double-execution concurrency guard. This is unusually rich behavioral context for a mutation-style executor call.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The purpose is front-loaded and each paragraph carries distinct information (contract, injection, async behavior, enforcement, concurrency). It is longer than average, and the aside about the settle window making cancel-race behaviour reproducible is marginal, but no sentence is dead weight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be spelled out; the description still explains the async return shape and the polling path. Combined with rejection conditions, timeout semantics, and the executor contract reference, nothing an agent needs to invoke this correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so trial_id, programme_id, and timeout_seconds are already fully documented in the schema, making 3 the baseline. The description does not add parameter-level detail beyond that (its detail concerns the bundle contract, not the arguments), so it neither compensates nor detracts.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Run a trial by calling the executor role') and immediately defines the mechanism (imports run_training from the bundle's code_ref and calls it with the trial config). It is cleanly distinguishable from siblings like get_trial_status, wait_trial, and cancel_trial.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly routes the agent: for long-running jobs this returns quickly with status 'running', and 'Use get_trial_status to poll for completion.' It also names the conditions under which the call is rejected (bundle not fully captured, trial already running), which is exactly the when/when-not guidance an agent needs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.