A
licenseA
qualityA
maintenanceEval-integrity statistics for AI benchmark claims — multiple-testing correction, power/MDE for model gaps, judge-bias and leaderboard-rank checks. Catches a
benchmark number that won't survive a second look.
9
MIT
No user-submitted related servers found.
Scored across 6 tools
Each tool targets a distinct statistical question: A/B testing, change detection, forecasting, denominator shift, forecast evaluation, and multiple testing correction. No overlap in purposes.
All tool names follow a consistent verb_noun pattern in lowercase snake_case, e.g., did_it_change, forecast_next, which_metrics_matter. No deviations.
6 tools is an ideal number for a focused statistics toolkit, covering essential operations without being overwhelming or sparse.
The set covers key statistical tasks (A/B testing, change detection, forecasting, multiple testing), but lacks a sample size/power analysis tool, which is a minor gap.