get_score_trust
Check if probability scores match recorded red-team or BAS outcomes, and if the priority order truly discriminates. Use this before quoting any score as a probability.
Instructions
Report how well the engine's probabilities have matched reality, measured against recorded red-team or BAS outcomes: the verdict (well-calibrated / calibrated-on-average / overconfident / underconfident / insufficient-data; calibrated-on-average means only the average matches, so no individual score may be quoted as a probability), the predicted-versus-observed rates, and what to do about the gap. Call this before quoting any score as a probability. If it reports insufficient-data, the numbers are expert estimates and must be presented as the model's own estimate, not as odds. It also reports discrimination - whether the score, and separately the triage Priority order, put confirmed paths above refuted ones (AUC with a 95% interval). Do not present the ranking as evidence of which path is most dangerous unless priorityDiscrimination reads 'discriminates'; 'insufficient-data' or 'indistinguishable-from-chance' means the order has not been shown to beat a coin, and must be said so.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
No arguments | |||