Skip to main content
Glama

sequential_two_sample_mean_test

Read-onlyIdempotent

Compare two groups' means with always-valid p-values, safe to re-check after each new observation as data accumulates in live experiments. Avoids false positives from repeated testing.

Instructions

Always-valid test of whether two groups' means differ, safe to call again after every new observation in either group -- unlike two_sample_t_test, which needs a sample size decided in advance and gives no such guarantee if checked repeatedly and stopped at the first significant look (that repeated-checking failure mode is exactly what inflates false positives; see naive_peeking_inflation for a demonstration). Use this instead of two_sample_t_test whenever a result will be (or already has been) checked more than once as data accumulates, e.g. monitoring a live experiment. Returns the current effect estimate, its standard error, the mixture likelihood ratio and always-valid p-value, and assumption warnings. tau does not need to be exact -- reuse the minimum-detectable-effect you'd otherwise plug into sample_size_for_two_sample_t_test.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
aYesfirst group's observations so far -- can be re-checked as more come in
bYessecond group's observations so far, same units as a
tauYesmixing prior's standard deviation over the true mean difference, in a/b's own units -- e.g. the smallest difference worth caring about. Not a threshold; see the tool's docstring
alphaNosignificance level for the test (and any confidence interval); default 0.05
equal_varNoassume equal population variances (pooled) instead of Welch's, same meaning as two_sample_t_test's equal_var

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv0.5.0

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already provide readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is covered. The description adds meaningful behavioral context: the always-valid guarantee, safety to re-call after each observation, the repeated-checking failure mode, and the return payload (effect estimate, standard error, mixture likelihood ratio, always-valid p-value, assumption warnings).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose and then builds through contrast, usage guidance, return values, and param guidance. It is longer than strictly necessary, especially the parenthetical about naive_peeking_inflation, but each sentence contributes useful decision-relevant information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description wisely covers the return values. It also addresses the key selection context (repeated checking), the main alternative (two_sample_t_test), and tau tuning. It could be slightly more explicit about data-shape expectations or assumptions, but the current detail is adequate for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds value by explaining tau practically: 'tau does not need to be exact -- reuse the minimum-detectable-effect you'd otherwise plug into sample_size_for_two_sample_t_test.' This helps an agent choose a real tau value rather than treating it as an exact threshold.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Always-valid test of whether two groups' means differ.' It clearly differentiates itself from two_sample_t_test by emphasizing repeated checking and sequential safety, so an agent can distinguish this from its main sibling without reading schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says when to use this tool: 'Use this instead of two_sample_t_test whenever a result will be (or already has been) checked more than once as data accumulates.' It also names the alternative and explains the failure mode of repeated peeking, giving clear decision criteria.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.