Eval-integrity statistics for AI benchmark claims — multiple-testing correction, power/MDE for model gaps, judge-bias and leaderboard-rank checks. Catches a
benchmark number that won't survive a second look.
Provides honest A/B test verdicts using interactive cards, with statistical calculations for conversion rates, sample sizes, and Bayesian probability-to-beat.
Checks whether a trading backtest survives its own statistics: deflated Sharpe, multiple-testing correction against a best-of-N-noise benchmark, minimum track record length, and fill realism. Takes no market data and no API keys, and cannot recommend a trade — it only reports that a result is weaker than claimed or not yet provable.
Provides verified statistical inference and hypothesis testing tools, including t-tests, effect sizes, power analysis, and multiple comparisons correction, with assumption checks and citations.
Answers "sales dropped since last week — where?" by comparing a target period against a weekday-adjusted baseline and localizing which attribute combinations (e.g. channel=web, or Tuesday nights) explain the shift. Runs fully offline on your own CSV — no API key, no ML training, read-only.
Glass-box statistical analysis for time series and business data: 19 research-grade methods, cited findings, and measured detector false-fire rates shipped as a calibration corpus. Agents cite real math instead of inventing it.