MCP server that provides tools for evaluating LLM agent reliability, including adversarial task generation, automated LLM-as-judge assessment, and confidence statistics.
An MCP server that lets any AI agent evaluate RAG outputs -- faithfulness scoring, hallucination detection, and retrieval quality metrics -- with zero API keys, using MCP sampling.
MCP server that scans multi-language codebases to detect stubs, missing imports, and structural incompleteness, assigning drift scores to quantify LLM context erosion.
An MCP server implementing the Flourishing-Justice-Autonomy (FJA) alignment framework, enabling FJA evaluation and fine-tuning of LLMs through examples like cultural diet, medical autonomy, and hiring fairness.
An MCP server that runs curated adversarial prompts against local Ollama models to test guardrails, scoring responses with heuristic verdicts for human review.