rigor-mcp
Related Servers
Alternatives to rigor-mcp
No user-submitted related servers found.
Related Servers
- AlicenseAqualityAmaintenanceEval-integrity statistics for AI benchmark claims — multiple-testing correction, power/MDE for model gaps, judge-bias and leaderboard-rank checks. Catches a benchmark number that won't survive a second look.9MIT
- FlicenseNot gradedqualityCmaintenanceEnables comprehensive statistical analysis including descriptive statistics, hypothesis testing, regression, and more via a FastMCP-based API.3-
- FlicenseNot gradedqualityBmaintenanceEnables auditing scientific papers for methodological biases such as selection bias and p-hacking, and assessing citation credibility and research consensus.8-
- AlicenseAqualityAmaintenanceChecks whether a number is real or just noise: peek-safe A/B tests you can look at as often as you like without inflating false positives, two-sided change detection, and a guard for when a metric moved only because its sample size did. Zero dependencies, standard library only.6MIT
- AlicenseAqualityAmaintenanceEnables AI agents to perform reproducible, verifiable statistical analysis through 25 deterministic tools for descriptive statistics, hypothesis testing, regression, clustering, time-series forecasting, and Chinese-labeled plotting.301MIT
- FlicenseAqualityDmaintenanceProvides tools for managing quantitative research knowledge graphs, enabling structured representation of research projects, datasets, variables, hypotheses, statistical tests, models, and results.69-
TDQS
Scored across 37 tools
Every tool targets a distinct statistical procedure or design (e.g., paired vs. independent proportions, fixed-sample vs. sequential means, chi-square vs. exact test), and the descriptions explicitly cross-reference sibling tools to prevent misselection. Even the paired power/sample-size tools are clearly differentiated as inverse operations.
All tool names use consistent snake_case and follow a predictable pattern: tests are named after the procedure (e.g., two_sample_t_test, mann_whitney_u), effect sizes after the statistic, and power/sample-size tools share the power_for_/sample_size_for_ prefix. The few non-test utilities (recommend_test, naive_peeking_inflation) are still stylistically consistent.
37 tools is well beyond the 25+ threshold for a heavy surface, even though each tool is individually useful. The inclusion of recommend_test and pairwise_group_comparisons mitigates the burden, but the set would benefit from consolidation or a more focused scope.
The suite covers the common lifecycle of a statistical analysis: assumption checks, parametric/nonparametric tests, effect sizes, power/sample-size planning, multiple-comparison corrections, and pairwise follow-ups. Minor gaps remain, such as no power/sample-size calculations for ANOVA or chi-square and no explicit normality test, but agents can work around these.