Skip to main content
Glama

Related Servers

Alternatives to agent-eval-mcp

No user-submitted related servers found.

    Related Servers

    • A
      license
      A
      quality
      B
      maintenance
      An MCP server that lets any AI agent evaluate RAG outputs -- faithfulness scoring, hallucination detection, and retrieval quality metrics -- with zero API keys, using MCP sampling.
      6
      MIT
    • A
      license
      A
      quality
      A
      maintenance
      MCP server that lets coding agents test AI agents. Create YAML test cases, snapshot golden baselines, check for regressions, and generate visual reports all from inside Claude Code or any MCP-compatible tool. Works with LangGraph, CrewAI, OpenAI, Claude, Mistral, and any HTTP API.
      10
      57 npm
      543 PyPI
      137
      Apache 2.0
    • A
      license
      Not graded
      quality
      B
      maintenance
      A local MCP server that provides tools for AI coding agents to execute, evaluate, and diff regression tests for LLM agents.
      3 npm
      MIT
    • A
      license
      A
      quality
      A
      maintenance
      An MCP server that exposes governance, trust-scoring, compliance, guardrail, cost, drift, and supply-chain scanning tools and resources to any MCP client over stdio, Streamable HTTP, or legacy HTTP+SSE. It lets agents route every tool call through a deterministic five-way decision (allow, redact, require approval, deny, or quarantine) with hash-chained evidence, human approval workflows, and in-agent trust gates for LangChain, LangGraph, and Google ADK.
      30
      2
      MIT

    TDQS

    B3/5.0

    Scored across 3 tools

    Disambiguation5/5

    Each tool has a distinct purpose: running evaluations, comparing results, and applying deployment policy. There is no overlap.

    Naming Consistency5/5

    All tools follow a consistent verb_noun pattern (run_evaluation_suite, compare_regression, decide_canary), making them predictable.

    Tool Count4/5

    With 3 tools, the server covers the core evaluation workflow without being bloated. Slightly minimal but appropriate for a focused server.

    Completeness4/5

    The tools cover the main lifecycle: run suite, compare results, decide promotion. Minor gaps like suite management exist, but the core workflow is complete.

    Maintenance

    ActivityInactive
    ResponsivenessNo issues