Skip to main content
Glama

list_experiments

List experiments for a pipeline, each with its most recent run. Use the returned object ID to view nested properties like status and selection.

Instructions

Lists a pipeline's Experiments, each with its most recent run.

An Experiment is a named set of Evaluators that runs against chosen Sessions of one pipeline. last_run is the newest run without its grid (status, selection, timing), or null when the Experiment never ran; read a run's results with get_experiment_run. :param pipeline_name: Name of the pipeline. :returns: The Experiments, or an error message.

The output is automatically stored and can be referenced in other functions. Returns a formatted preview with an object ID (e.g., @obj_123). Use the object store tools in combination with the object ID to view nested properties of the object. Use the returned object ID to pass this result to other functions.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
pipeline_nameYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv0.1.29

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does reasonable work: it discloses that last_run is the newest run stripped of its grid, that it is null when the Experiment never ran, and that the tool returns 'an error message' on failure. It also explains the object-store side effect (result auto-stored, preview plus @obj_ id). It does not mention pagination or permission requirements.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with purpose, then semantics, then return information, which is a sensible order. The four-sentence object-store block is generic boilerplate, but in this tool family it carries operational value; only minor trimming is possible.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must describe returns, and it does: Experiments with last_run, null semantics, error case, and the object ID for follow-up calls. Combined with the pointer to get_experiment_run, an agent has what it needs for a one-parameter list tool; only pagination/volume expectations are unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single parameter 'pipeline_name' has no schema description, so the description must compensate. The ':param pipeline_name: Name of the pipeline.' line is present but adds only trivially obvious meaning; it does not explain that the name must refer to an existing pipeline or whether it is case-sensitive.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states a specific verb and resource with scope: 'Lists a pipeline's Experiments, each with its most recent run.' It also defines what an Experiment is (a named set of Evaluators against Sessions of one pipeline), which is useful domain grounding. It distinguishes itself from get_experiment_run for the follow-up task, but does not explicitly separate itself from the other list_* siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'read a run's results with get_experiment_run' is a genuine routing hint to the correct alternative for a related need. However, there is no explicit statement of when to use this tool versus list_evaluators / list_sessions, nor any precondition or exclusion. Usage is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools