sequential_two_sample_mean_test
Compare two groups' means with always-valid p-values, safe to re-check after each new observation as data accumulates in live experiments. Avoids false positives from repeated testing.
Instructions
Always-valid test of whether two groups' means differ, safe to call again after every new observation in either group -- unlike two_sample_t_test, which needs a sample size decided in advance and gives no such guarantee if checked repeatedly and stopped at the first significant look (that repeated-checking failure mode is exactly what inflates false positives; see naive_peeking_inflation for a demonstration). Use this instead of two_sample_t_test whenever a result will be (or already has been) checked more than once as data accumulates, e.g. monitoring a live experiment. Returns the current effect estimate, its standard error, the mixture likelihood ratio and always-valid p-value, and assumption warnings. tau does not need to be exact -- reuse the minimum-detectable-effect you'd otherwise plug into sample_size_for_two_sample_t_test.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| a | Yes | first group's observations so far -- can be re-checked as more come in | |
| b | Yes | second group's observations so far, same units as a | |
| tau | Yes | mixing prior's standard deviation over the true mean difference, in a/b's own units -- e.g. the smallest difference worth caring about. Not a threshold; see the tool's docstring | |
| alpha | No | significance level for the test (and any confidence interval); default 0.05 | |
| equal_var | No | assume equal population variances (pooled) instead of Welch's, same meaning as two_sample_t_test's equal_var |