Skip to main content
Glama

Related Servers

Alternatives to agent-eval

No user-submitted related servers found.

    Related Servers

    • A
      license
      A
      quality
      C
      maintenance
      Deterministic time-series statistics for AI agents. This MCP server gives any LLM agent unit-tested statistical tools — anomaly detection, changepoint detection, seasonal decomposition, stationarity/trend tests, data-quality audits, baseline forecasts — with schema-validated structured output and no arbitrary code execution.
      17
      MIT
    • F
      license
      Not graded
      quality
      D
      maintenance
      MCP server that diagnoses RAG pipeline regression by detecting which failure modes (e.g., retrieval miss, hallucination) significantly regressed, not just score drops, using statistical methods to avoid false alarms.
      -
    • A
      license
      A
      quality
      A
      maintenance
      MCP server that lets coding agents test AI agents. Create YAML test cases, snapshot golden baselines, check for regressions, and generate visual reports all from inside Claude Code or any MCP-compatible tool. Works with LangGraph, CrewAI, OpenAI, Claude, Mistral, and any HTTP API.
      10
      63 npm
      540 PyPI
      134
      Apache 2.0
    • F
      license
      Not graded
      quality
      C
      maintenance
      An MCP server that exposes causal inference methods (difference-in-differences, synthetic control, propensity matching, and assumption checks) as callable tools, enabling AI agents to run deterministic statistical analyses instead of computing them inline.
      -
    • A
      license
      Not graded
      quality
      B
      maintenance
      A local MCP server that provides tools for AI coding agents to execute, evaluate, and diff regression tests for LLM agents.
      2 npm
      MIT

    TDQS

    A3.7/5.0

    Scored across 1 tool

    Disambiguation5/5

    With only one tool available, there is no possibility of confusing it with another; its purpose is clearly described.

    Naming Consistency4/5

    The tool name 'run' is a simple, direct verb that matches its function. While it lacks a noun modifier, there is no inconsistency to penalize.

    Tool Count3/5

    The server has a single tool, which feels thin for a general evaluation platform but is appropriate for a focused CLI wrapper. It is borderline but not excessive.

    Completeness4/5

    The server's stated purpose is to run the agent-regress CLI, and the single tool fulfills that. However, it lacks auxiliary capabilities like listing available tests or parsing results separately.

    Maintenance

    ActivityMaintained
    ResponsivenessUnresponsive