Skip to main content
Glama

sequential_two_proportion_test

Read-onlyIdempotent

Determine if two proportions differ in a live experiment, supporting repeated checks after each new observation while preserving statistical validity.

Instructions

Always-valid test of whether two proportions (e.g. two conversion rates in a live A/B test) differ, safe to call again after every new observation in either group -- the sequential-monitoring counterpart to two_proportion_z_test. Use this instead whenever the result will be checked more than once before the experiment ends, which is the normal case for a live dashboard rather than a one-shot analysis. p1 and p2 are interchangeable (only their difference matters). tau does not need to be exact -- reuse the minimum-detectable-effect you'd otherwise plug into sample_size_for_two_proportion_test.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
n1Yesobservations so far in group 1
n2Yesobservations so far in group 2
tauYesmixing prior's standard deviation over the true proportion difference -- e.g. 0.02 for 'I mainly care about a 2-point-or-larger swing'. Not a threshold; see the tool's docstring
alphaNosignificance level for the test (and any confidence interval); default 0.05
successes1Yessuccesses (e.g. conversions) observed so far in group 1
successes2Yessuccesses observed so far in group 2

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv0.5.0

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark readOnly/idempotent/non-destructive, and the description adds meaningful behavioral context: it is 'safe to call again after every new observation', producing an 'always-valid' sequential test. It also discloses that p1/p2 are interchangeable and tau is not a hard threshold, which goes beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, front-loaded with the core purpose and the most important sequential property, followed by when-to-use and parameter guidance. No filler or repetition of schema details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the main conceptual complexity, the live-dashboard use case, the alternative, and the non-obvious tau parameter, so an agent can call it correctly. It does not describe the return shape, but for a statistical hypothesis test the return is conventional and the schema covers all arguments; the slight gap is acceptable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is a 3; the description adds extra semantic value by explaining tau as a reuse of the minimum-detectable-effect and noting that p1 and p2 are interchangeable. This helps an agent choose and set tau without requiring the full statistical derivation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description opens with an active verb phrase: 'Always-valid test of whether two proportions ... differ'. It clearly identifies the resource (two proportions, e.g. conversion rates), and distinguishes itself as the 'sequential-monitoring counterpart to two_proportion_z_test' rather than merely restating the name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit use condition: 'Use this instead whenever the result will be checked more than once before the experiment ends', and contrasts it with a one-shot analysis. It also names the sibling alternative (two_proportion_z_test) and gives tau guidance via sample_size_for_two_proportion_test.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.