get_experiment_run
Retrieve an experiment run's report, including cell outcomes, per-metric summaries, and grid rows for debugging and analysis.
Instructions
Reads one Experiment run: how its cells ended, a summary per Metric, then the grid rows.
A run judges each chosen Session with each Evaluator; one cell per pair. counts says how
the cells ended: errored is the Evaluator's own code breaking (see the cell's
error_detail), not a low Metric; failed is a cell that never ran and can be retried.
metrics summarises each Evaluator's Metric over every row, rows past max_rows
included: mean, min and max for a Score, a count per label for a Label, and how often it did
not apply. Each row is a Session with its cells; a cell's metrics carry the value per
address (session, turn/2, ...) and the rationale the Evaluator wrote, which is what to
group failures by. run.selection says whether the Sessions were SAMPLED or HAND_PICKED;
only sampled runs say anything about traffic as a whole.
:param pipeline_name: Name of the pipeline.
:param experiment_id: The Experiment's id.
:param run_id: The run's id.
:param max_rows: Most rows to return; the summaries still cover every row.
:returns: The run report, or an error message.
The output is automatically stored and can be referenced in other functions.
Returns a formatted preview with an object ID (e.g., @obj_123).
Use the object store tools in combination with the object ID to view nested properties of the object.
Use the returned object ID to pass this result to other functions.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | ||
| max_rows | No | ||
| experiment_id | Yes | ||
| pipeline_name | Yes |