Skip to main content
Glama

Run Experiment

run_experiment

Launch 1-25 pipelines as one paced, non-blocking experiment and get an experimentId instantly; poll get_experiment for per-run status and aggregated metrics.

Instructions

Run 1-25 pipelines as ONE paced experiment (NON-BLOCKING). Returns an experimentId immediately; a background thread submits at most 2 runs at a time (min(max_concurrent, 2)), retries queue-full up to 3 times per run, and polls each execution to completion. Poll get_experiment() for per-run status and, once finished, aggregated metrics.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
runsYes
max_concurrentNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.5.1

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden and does so well: non-blocking return of an experimentId, a background thread capped at min(max_concurrent, 2), retry of queue-full up to 3 times per run, and polling to completion. It omits auth requirements and what happens on partial failure, but the concurrency/retry semantics are unusually well disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the critical behavioral fact '(NON-BLOCKING)' and the 1-25 bound, and every clause carries operational information. It is dense to the point of being slightly run-on, but no sentence is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a batch-execution tool with an output schema, the description correctly covers non-blocking semantics, scheduling behavior, and the handoff to get_experiment, so return-value detail is not needed. The only real gap is the shape of the `runs` array elements, which is not covered anywhere.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and one parameter (`runs`) is an untyped array of open objects whose per-run shape is never explained. The description partially compensates by bounding `runs` at 1-25 pipelines and by clarifying that `max_concurrent` is effectively clamped to 2, but the structure of each run entry remains undocumented in both schema and description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (run) and resource (1-25 pipelines as one paced experiment), and explicitly contrasts with the sibling polling tool get_experiment(). An agent can immediately tell this apart from run_pipeline (single run) and get_experiment (status polling) without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It names the follow-up tool and the condition that selects it ('Poll get_experiment() for per-run status and, once finished, aggregated metrics'), and '(NON-BLOCKING)' tells the agent it can continue working after the call. It stops short of saying when to prefer run_pipeline over this batch tool, so it is clear context rather than full when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.