MCP server exposing statistical regression testing for LLM agents as a "run" tool: p-value, effect size, and confidence interval on whether agent behavior actually changed.
Vendor-neutral CLI and MCP server that verifies the token and cost savings AI-coding-agent context-reduction proxies actually deliver, measured against a real labeled corpus.
Enables comparison of responses from multiple LLMs (OpenAI, Anthropic, Gemini) to the same prompt, returning a validated divergence score based on sentence embeddings.
Exposes a run_suite tool to evaluate whether an AI agent is safe to operate internal web apps, scoring task completion and forbidden-action violations to gate CI/CD pipelines.