Skip to main content
Glama

Related Servers

Alternatives to rigor-mcp

No user-submitted related servers found.

    Related Servers

    • A
      license
      A
      quality
      A
      maintenance
      Eval-integrity statistics for AI benchmark claims — multiple-testing correction, power/MDE for model gaps, judge-bias and leaderboard-rank checks. Catches a benchmark number that won't survive a second look.
      9
      MIT
    • F
      license
      Not graded
      quality
      C
      maintenance
      Enables comprehensive statistical analysis including descriptive statistics, hypothesis testing, regression, and more via a FastMCP-based API.
      3
      -
    • F
      license
      Not graded
      quality
      B
      maintenance
      Enables auditing scientific papers for methodological biases such as selection bias and p-hacking, and assessing citation credibility and research consensus.
      8
      -
    • A
      license
      A
      quality
      A
      maintenance
      Checks whether a number is real or just noise: peek-safe A/B tests you can look at as often as you like without inflating false positives, two-sided change detection, and a guard for when a metric moved only because its sample size did. Zero dependencies, standard library only.
      6
      MIT
    • A
      license
      A
      quality
      A
      maintenance
      Enables AI agents to perform reproducible, verifiable statistical analysis through 25 deterministic tools for descriptive statistics, hypothesis testing, regression, clustering, time-series forecasting, and Chinese-labeled plotting.
      30
      1
      MIT

    TDQS

    A4.1/5.0

    Scored across 37 tools

    Disambiguation5/5

    Every tool targets a distinct statistical procedure or design (e.g., paired vs. independent proportions, fixed-sample vs. sequential means, chi-square vs. exact test), and the descriptions explicitly cross-reference sibling tools to prevent misselection. Even the paired power/sample-size tools are clearly differentiated as inverse operations.

    Naming Consistency5/5

    All tool names use consistent snake_case and follow a predictable pattern: tests are named after the procedure (e.g., two_sample_t_test, mann_whitney_u), effect sizes after the statistic, and power/sample-size tools share the power_for_/sample_size_for_ prefix. The few non-test utilities (recommend_test, naive_peeking_inflation) are still stylistically consistent.

    Tool Count2/5

    37 tools is well beyond the 25+ threshold for a heavy surface, even though each tool is individually useful. The inclusion of recommend_test and pairwise_group_comparisons mitigates the burden, but the set would benefit from consolidation or a more focused scope.

    Completeness4/5

    The suite covers the common lifecycle of a statistical analysis: assumption checks, parametric/nonparametric tests, effect sizes, power/sample-size planning, multiple-comparison corrections, and pairwise follow-ups. Minor gaps remain, such as no power/sample-size calculations for ANOVA or chi-square and no explicit normality test, but agents can work around these.

    Maintenance

    ActivityMaintained
    ResponsivenessNo issues