Evaluate text
evaluateScore any text against Overwing rule sets to classify it as pass, fail, or review, with per-rule probability and confidence for blocking, redacting, or human routing.
Instructions
Score any text (typically an LLM's output) against an Overwing rule set. Returns an aggregate verdict of pass, fail, or review plus per-rule answers with probability and confidence. fail means a rule's fail condition matched: block or redact. review means a rule was unsure: route to a human or slower model. The prebuilt 'content-safety' set checks toxicity, PII, self-harm, sexual content, and severity.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| input | Yes | The text to evaluate | |
| metadata | No | Opaque context stored with the evaluation (max 8 KB) | |
| rule_set | No | Rule set slug | content-safety |