Checks whether a number is real or just noise: peek-safe A/B tests you can look at as often as you like without inflating false positives, two-sided change detection, and a guard for when a metric moved only because its sample size did. Zero dependencies, standard library only.
It enables solo builders to automatically log product changes to PostHog and later get statistically honest verdicts on whether those changes moved key metrics, without needing enough traffic for A/B tests.
Eval-integrity statistics for AI benchmark claims — multiple-testing correction, power/MDE for model gaps, judge-bias and leaderboard-rank checks. Catches a
benchmark number that won't survive a second look.
MCP-first A/B testing server that enables agents to manage experiments (create, update traffic splits, read results, apply winning variants) from tools like Claude Code, with self-hosted Cloudflare backend.
Checks whether a trading backtest survives its own statistics: deflated Sharpe, multiple-testing correction against a best-of-N-noise benchmark, minimum track record length, and fill realism. Takes no market data and no API keys, and cannot recommend a trade — it only reports that a result is weaker than claimed or not yet provable.
Provides verified statistical inference and hypothesis testing tools, including t-tests, effect sizes, power analysis, and multiple comparisons correction, with assumption checks and citations.