Revert benchmark reference updates
revert_benchmark_reference_updatesUndo automatic edits a scoring pass made to a scenario's reference. Every scoring pass folds what the scored models showed the reference should be (a candidate the judge found better, a rule the samples prove, a value the reference lacked) and writes it into the reference — an edit a pass already made is only replaced by stronger evidence (samples, a wrong verdict, more agreeing models), never by one more model's better verdict; the log is reference_meta.auto_applied on get_benchmark_scenario, each entry with its inverse patch. Pass the revision ids to undo: the inverse is applied, the entry is marked reverted and its (path, attribute) is pinned so no later pass re-applies it (a manual set_benchmark_reference lifts every pin). Scores never go stale from this. Requires owner and a plan with benchmarks. No LLM call.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| scenario_id | Yes | UUID of the scenario. | |
| revision_ids | Yes | Ids of the reference_meta.auto_applied entries to undo (1..200). |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||