Skip to main content
Glama

Related Servers

Alternatives to Autoeval

No user-submitted related servers found.

    Related Servers

    • A
      license
      Not graded
      quality
      B
      maintenance
      Enables evidence-first regression testing for AI agents by turning production traces into reviewed, replayable cases that gate releases. It supports reproducible, auditable agent evaluation with controlled tool execution, evidence-based judging, and versioned quality gates.
      MIT
    • A
      license
      Not graded
      quality
      C
      maintenance
      Enables evidence-first release readiness assessment by running or accepting build, API, browser, visual, performance, and security evidence, then returning SHIP, REVIEW, or HOLD recommendations with clustered regressions.
      9 npm
      MIT
    • A
      license
      Not graded
      quality
      C
      maintenance
      Provides advanced evaluation tools for assessing AI safety, alignment, and performance of LLM outputs. Enables programmatic evaluation of quality, safety metrics like toxicity and PII detection, and operational metrics including carbon footprint and cost estimation.
      4
      Apache 2.0
    • A
      license
      Not graded
      quality
      F
      maintenance
      Provides SDLC compliance verification as tools that AI agents can invoke, continuously monitoring and evaluating development processes.
      MIT
    • A
      license
      A
      quality
      A
      maintenance
      Provides coding agents with visibility into test health through tools for flaky test detection, test quality linting, and LLM evaluation harness, enabling them to triage failures, review test quality, and check prompt changes for regressions.
      9
      Apache 2.0