Grade a decision against what happened after it
grade_decisionGrades one recommendation and decision against before/after observations, returning holding, not-holding, or refused to measure whether advice actually worked.
Instructions
Grades one recommendation-and-decision pair against a before/after observation log. Compares a pre-decision BASELINE window to a post-decision, exposure-aligned RESULT window on the same subjectId + checkKey, and returns 'holding' (the exposed post-decision window stayed under refuteThreshold bad observations AND its bad rate is not higher than the baseline's), 'not-holding' (EITHER the exposed post-decision window hit refuteThreshold bad observations, OR its bad rate is higher than the baseline's, even below that count -- 1 bad of 10 before, 1 bad of 3 exposed since is 'not-holding', not 'holding', despite the count staying under the default bar of 2), or 'refused' (the evidence didn't clear a floor -- see refusalCodes for exactly which one, never a vague 'unproven'). The rate comparison is exact -- it cross-multiplies the raw counts rather than comparing the rounded badRate fields or the sign of badRateDelta, so a not-holding can occur even when both windows display the same 3-place badRate. Always show badRateDelta and the raw baseline/result counts next to the verdict, don't quote 'holding' on its own. The verdict grades the DECISION, not just whether advice was taken: a DISMISSED recommendation whose problem later surfaced also grades 'not-holding', because the evidence sided with the advice either way. Two things this tool refuses to let slide: (1) 'nothing has gone wrong since' only counts as a result if the baseline shows the problem existed before -- otherwise it's 'baseline_lacks_negative_signal'; (2) only observations where the advice could actually have applied (exposed === true) can create or reverse the headline verdict -- everything else is reported separately as secondary (whose own wouldBeVerdict can itself be 'refused'), labelled non-headline, and can never become the headline. Use this to close the loop on any recommender/advisor system that logs recommendations, human accept/reject calls, and later outcomes, instead of letting advice quality go unmeasured forever. For grading a whole population of engine-vs-human calls at once (rather than one decision against its own before/after window), use compute_divergence instead.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| config | No | Floors and thresholds, every field optional with a library default (see advice-ledger-kit's DEFAULT_GRADE_CONFIG). Omit entirely to use the library's defaults. | |
| decision | Yes | The human's call on the recommendation. | |
| observations | Yes | Every observation available. This tool does the filtering by subjectId + checkKey itself, so passing the whole ledger (observations for other subjects/checks too) is fine. | |
| recommendation | Yes | The advice being graded. |