Skip to main content
Glama

Grade a decision against what happened after it

grade_decision

Grades one recommendation and decision against before/after observations, returning holding, not-holding, or refused to measure whether advice actually worked.

Instructions

Grades one recommendation-and-decision pair against a before/after observation log. Compares a pre-decision BASELINE window to a post-decision, exposure-aligned RESULT window on the same subjectId + checkKey, and returns 'holding' (the exposed post-decision window stayed under refuteThreshold bad observations AND its bad rate is not higher than the baseline's), 'not-holding' (EITHER the exposed post-decision window hit refuteThreshold bad observations, OR its bad rate is higher than the baseline's, even below that count -- 1 bad of 10 before, 1 bad of 3 exposed since is 'not-holding', not 'holding', despite the count staying under the default bar of 2), or 'refused' (the evidence didn't clear a floor -- see refusalCodes for exactly which one, never a vague 'unproven'). The rate comparison is exact -- it cross-multiplies the raw counts rather than comparing the rounded badRate fields or the sign of badRateDelta, so a not-holding can occur even when both windows display the same 3-place badRate. Always show badRateDelta and the raw baseline/result counts next to the verdict, don't quote 'holding' on its own. The verdict grades the DECISION, not just whether advice was taken: a DISMISSED recommendation whose problem later surfaced also grades 'not-holding', because the evidence sided with the advice either way. Two things this tool refuses to let slide: (1) 'nothing has gone wrong since' only counts as a result if the baseline shows the problem existed before -- otherwise it's 'baseline_lacks_negative_signal'; (2) only observations where the advice could actually have applied (exposed === true) can create or reverse the headline verdict -- everything else is reported separately as secondary (whose own wouldBeVerdict can itself be 'refused'), labelled non-headline, and can never become the headline. Use this to close the loop on any recommender/advisor system that logs recommendations, human accept/reject calls, and later outcomes, instead of letting advice quality go unmeasured forever. For grading a whole population of engine-vs-human calls at once (rather than one decision against its own before/after window), use compute_divergence instead.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
configNoFloors and thresholds, every field optional with a library default (see advice-ledger-kit's DEFAULT_GRADE_CONFIG). Omit entirely to use the library's defaults.
decisionYesThe human's call on the recommendation.
observationsYesEvery observation available. This tool does the filtering by subjectId + checkKey itself, so passing the whole ledger (observations for other subjects/checks too) is fine.
recommendationYesThe advice being graded.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and discharges it: it enumerates the three verdicts, the exact cross-multiplication semantics, the exposure (exposed === true) rule, the baseline_negative_signal refusal, the secondary/non-headline handling, and the downgrade of a dismissed-but-correct recommendation. This is far beyond what any structured field provides.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core purpose and routing are front-loaded in the first sentence, and every sentence is load-bearing (verdict logic, exposure rule, secondary handling, alternative tool). It is dense and long, however, with the parenthetical worked example pushing it toward the verbose side, so it is efficient but not maximally tight.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex nested-object tool with no annotations and no output schema, the description supplies the missing return semantics: the three verdict strings, the fields to surface (badRateDelta, baseline/result counts), secondary.wouldBeVerdict, and refusalCodes. An agent has enough to interpret results without an output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description genuinely adds meaning: it explains how refuteThreshold drives 'holding' vs 'not-holding', that the rate comparison cross-multiplies raw counts rather than badRate fields, and how exposure gates the headline verdict. It stops short of restating per-field defaults, which is fine since the schema already covers them.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a precise verb (grades) plus resource (a recommendation-and-decision pair against a before/after observation log) and immediately defines the scope: same subjectId + checkKey, baseline vs result window. It is trivially distinguishable from siblings like compute_divergence, which it names as the population-level alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Use this to close the loop on any recommender/advisor system that logs recommendations, human accept/reject calls, and later outcomes' and names when NOT to use it: 'For grading a whole population of engine-vs-human calls at once ... use compute_divergence instead.' Both the when and the alternative are stated outright.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.