run_eval
Measure a rule set against hand-graded calls before shipping: check per-rule agreement, quote validity, and pass/fail against min_agreement; pass rules_yaml to test drafts unsaved.
Instructions
Measure a rule set against the hand-graded calls before it ships: agreement with human labels per rule, quote validity, and pass or fail. Every rule and the total have to reach min_agreement. Pass rules_yaml to test a draft rule set without saving it. A rule with no human labels cannot pass.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| scorer | No | baseline | |
| rules_yaml | No | ||
| min_agreement | No | ||
| rules_version | No |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||