Skip to main content
Glama
louislaurent1

Kryve Agent Evaluation MCP

Related Servers

Alternatives to Kryve Agent Evaluation MCP

No user-submitted related servers found.

    Related Servers

    • A
      license
      Not graded
      quality
      A
      maintenance
      Enables LLM evaluation and observability by uploading documents, building test sets, running RAG pipelines, and automatically scoring answers for groundedness, hallucination risk, retrieval quality, latency, and cost, with tools exposed to MCP-compatible clients.
      1
      MIT
    • A
      license
      A
      quality
      A
      maintenance
      Provides coding agents with visibility into test health through tools for flaky test detection, test quality linting, and LLM evaluation harness, enabling them to triage failures, review test quality, and check prompt changes for regressions.
      9
      Apache 2.0
    • A
      license
      A
      quality
      B
      maintenance
      Enables AI assistants to interact with Coval's evaluation platform for launching and monitoring evaluation runs, managing agents and test sets, and retrieving evaluation metrics.
      18
      8 npm
      2
      MIT
    • A
      license
      Not graded
      quality
      B
      maintenance
      Enables evidence-first regression testing for AI agents by turning production traces into reviewed, replayable cases that gate releases. It supports reproducible, auditable agent evaluation with controlled tool execution, evidence-based judging, and versioned quality gates.
      MIT
    • A
      license
      Not graded
      quality
      C
      maintenance
      Provides advanced evaluation tools for assessing AI safety, alignment, and performance of LLM outputs. Enables programmatic evaluation of quality, safety metrics like toxicity and PII detection, and operational metrics including carbon footprint and cost estimation.
      4
      Apache 2.0

    TDQS

    A3.9/5.0

    Scored across 3 tools

    Disambiguation5/5

    Each tool addresses a distinct phase of evaluation: generating test plans, scoring an agent run, and retrieving a scorecard structure. There is no functional overlap or ambiguity between them.

    Naming Consistency4/5

    Two tools follow a clear verb_noun pattern (generate_test_plan, score_agent_run), but evaluation_scorecard is noun-led rather than verb-led. The inconsistency is minor and all names remain readable and predictable.

    Tool Count5/5

    Three tools is an appropriate, focused scope for an evaluation-oriented MCP server. Each tool serves a core need without bloat, making the surface easy to navigate.

    Completeness4/5

    The set covers planning, assessment, and structure retrieval, but lacks an explicit verification or finalization tool. The 'unverified' caveat on score_agent_run hints at this gap, but the core workflow is still functional.

    Maintenance

    ActivitySlowing
    ResponsivenessNo issues