Skip to main content
Glama

Compare Experiments

compare_experiments

Compare experiment outcomes side by side to choose the best approach. Get test results, failure details, exit codes, duration, and a recommendation.

Instructions

Compare two or more experiments side by side.

USE THIS when you tried several approaches -- three candidate fixes, two runtime versions, a couple of dependency upgrades -- and have to recommend one. Run each approach in its own experiment, then compare.

RETURNS per-experiment test results, failing test names, exit codes, change size, duration and artifact counts, plus a recommendation when the evidence supports one. Destroyed experiments are still comparable: their findings were recorded before teardown.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
labelsNoOptional experiment_id -> human label, e.g. {'exp_a': 'Node 22'}.
experiment_idsYesTwo or more experiment ids.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
notesNo
dimensionsNoPer-dimension map of experiment_id -> value.
experimentsYes
recommendationNo

Schema Changelog

Changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. First observedv0.1.0

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden, and it does well: it discloses the return fields (test results, failing test names, exit codes, change size, duration, artifact counts), conditional recommendation behavior, and that destroyed experiments remain comparable. It does not explicitly state whether the operation is read-only or whether comparing triggers any side effects, but the described behavior is otherwise transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with clear sections: a terse one-line purpose, a concrete 'USE THIS' scenario, and a focused 'RETURNS' list. Every sentence earns its place, and the most decision-relevant information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has moderate complexity, an output schema, complete parameter schema coverage, and no annotations. The description covers purpose, usage context, return behavior, and the destroyed-experiment edge case. It could be more complete by explicitly stating that this is a read-only analytic comparison that does not run or modify experiments, but it is otherwise sufficient for an agent to select and call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema already documents experiment_ids and labels with examples. The description adds context about comparing destroyed experiments, which relates to how experiment_ids may be interpreted, but it does not add meaningful parameter-level semantics beyond what the schema provides. Baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool compares two or more experiments side by side, with a specific verb ('compare') and resource ('experiments'). It also names the core differentiating output (per-experiment results plus a recommendation), which separates it from siblings like get_experiment, list_experiments, and run_tests.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives an explicit scenario: use this when several approaches were tried and a recommendation is needed. It advises running each approach in its own experiment first. However, it does not explicitly state when NOT to use it or name alternative tools such as run_tests or get_experiment for simpler lookup needs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/Dhananjay-JSR/sandbox-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server