naive_peeking_inflation
Simulate how checking a standard test after each new observation and stopping at the first significant result inflates false-positive rates, showing why always-valid sequential tests are needed.
Instructions
Demonstrates, by simulation, why sequential_two_sample_mean_test / sequential_two_proportion_test exist: the actual false-positive rate of checking an ordinary fixed-sample test (two_sample_t_test, two_proportion_z_test, ...) after every new observation and stopping the first time it clears alpha, versus the alpha actually intended. Call this to show a skeptical stakeholder concretely what "just peek at the dashboard and stop early" costs before recommending the always-valid alternative. Returns the estimated true false-positive rate, its Monte Carlo standard error, and a citation.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| alpha | No | significance level for the test (and any confidence interval); default 0.05 | |
| trials | No | Monte Carlo trials -- higher is more precise but slower; the result reports its own standard error | |
| n_looks | Yes | how many times the result gets checked as data accumulates |