A
licenseNot graded
qualityC
maintenanceEnables a coding agent to iteratively rewrite a misbehaving prompt and rerun it, with every attempt graded against a frozen baseline by judges from other AI vendors so improvements are proven rather than assumed. It also classifies failures into nine origins, showing when the problem lies in the test set, rubric, judges, or surrounding code instead of the prompt.
Apache 2.0