A
licenseNot graded
qualityB
maintenanceEnables agents to turn a real model failure into a deterministic benchmark task by forging it with exact assertions, grading the task itself across six weighted factors, and inspecting a tamper-evident SHA-384 audit trail. Exposes eleven typed JSON-RPC 2.0 tools that call the same idempotent service layer as the UI, so mutations and decisions stay sealed and reproducible.
MIT