Enables building and running custom LLM benchmarks with multi-judge evaluation, supporting GUI, MCP client, and CLI usage for ranked, auditable results.
Provides deployable, stateless MCP services for rigorous inference evaluation and benchmarking, including a profiled lm-evaluation-harness controller/worker and a bounded vLLM forward-pass benchmark adapter.
This MCP server provides a stateful, resettable, verifiable API runtime that gates every tool call, enabling agents to run long workflows against provider-shaped environments without live provider write access. It records decisions, side effects, and outcome evidence for replayable, verifiable benchmark runs.
Enables automated adversarial safety probing and jailbreak fuzzing of AI systems through MCP, returning structured execution telemetry and output dossiers.