judge-audit-mcp
Related Servers
Alternatives to judge-audit-mcp
No user-submitted related servers found.
Related Servers
- AlicenseAqualityCmaintenanceMCP server that provides tools for evaluating LLM agent reliability, including adversarial task generation, automated LLM-as-judge assessment, and confidence statistics.31MIT
- AlicenseAqualityBmaintenanceAn MCP server that lets any AI agent evaluate RAG outputs -- faithfulness scoring, hallucination detection, and retrieval quality metrics -- with zero API keys, using MCP sampling.6MIT
- FlicenseNot gradedqualityCmaintenanceMCP server that scans multi-language codebases to detect stubs, missing imports, and structural incompleteness, assigning drift scores to quantify LLM context erosion.-
- AlicenseBqualityBmaintenanceAn MCP-style stdio server for evaluating AI agent outputs, enabling CI-friendly quality gates, regression comparisons, and canary promotion decisions.3MIT
- AlicenseBqualityCmaintenanceAn MCP server implementing the Flourishing-Justice-Autonomy (FJA) alignment framework, enabling FJA evaluation and fine-tuning of LLMs through examples like cultural diet, medical autonomy, and hiring fairness.2MIT
- AlicenseAqualityCmaintenanceAn MCP server that runs curated adversarial prompts against local Ollama models to test guardrails, scoring responses with heuristic verdicts for human review.5MIT
TDQS
Scored across 6 tools
Each tool serves a clearly distinct purpose: comprehensive audit, specific bias testing, calibration against humans, drift detection, data generation, and metric explanation. No two tools overlap in functionality.
All tool names follow a consistent verb_noun pattern in snake_case (e.g., audit_judge, bias_probe, explain_metric). The convention is uniform and predictable.
6 tools cover the essential aspects of LLM judge auditing without redundancy or bloat. Each tool is justified and serves a specific role in the workflow.
The set covers comprehensive auditing, bias probes, calibration, drift detection, data generation, and explanation. Minor gap: no tool for running a judge on raw inputs, but that is likely external to this server's scope.