Skip to main content
Glama
Mipiti
by Mipiti

Recompute Verdicts

recompute_verdicts

Estimate, queue, or retry a model's coverage and group-sufficiency verdict evaluations. Use it when control-to-mapping verdicts look wrong or transient failures left verdicts parked.

Instructions

Estimate, force, or retry a model's verdict evaluation.

  • mode="quote" (default): read-only. The cost of a recompute, enqueueing nothing: {estimated_credits, computed_at, rate_version, informational, total_enqueueable, already_evaluated, governor}. Subjects already carrying a verdict cost nothing, so it is an upper bound. Show the operator this number before recomputing: on a large model it runs to thousands of credits.

  • mode="recompute": mutating. Queues a fresh evaluation of every control's coverage verdict and every live objective's group-sufficiency verdict, bypassing quiet-period batching; usage is metered as it runs. Returns {model_id, model_version, enqueued_coverage, enqueued_group_sufficiency, total_enqueued, estimated_credits, quote, governor}.

  • mode="retry_parked": re-runs only the verdicts a transient failure (outage, exhausted credits, timeout) parked, of every kind, including per-control sufficiency and coherence; nothing else is touched. Returns {model_id, model_version, retried_slots, governor}.

Work runs in the background; when governor.exhausted it is queued and resumes at governor.resets_at, never dropped.

A recompute evaluates COVERAGE and GROUP SUFFICIENCY only. A control at partially_verified, or coherence_status: "pending" on an assertion, is not a reason to recompute: that verdict is computed on assertion write and read with get_sufficiency. Recompute when control-to-CO mappings look wrong (get_verdict_divergence). One objective awaiting judgement is judge_objective.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
modeNoquote
model_idYes
server_versionYes

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed3 schema fields changedv0.84.0
    • removedInput schema / properties / dry_run
      Removed value: -{
      -  "default": false,
      -  "description": "When True, return only the pre-flight estimate and enqueue\nnothing. When False (default), enqueue the recompute.",
      -  "type": "boolean"
      -}
    • addedInput schema / properties / mode
      Added value: +{
      +  "default": "quote",
      +  "enum": [
      +    "quote",
      +    "recompute",
      +    "retry_parked"
      +  ],
      +  "type": "string"
      +}
    • removedInput schema / properties / model_id / description
      Removed value: -"ID of the threat model to re-evaluate (or estimate for)."
  2. Changed2 schema fields changedv0.68.2
    • addedInput schema / properties / dry_run
      Added value: +{
      +  "default": false,
      +  "description": "When True, return only the pre-flight estimate and enqueue\nnothing. When False (default), enqueue the recompute.",
      +  "type": "boolean"
      +}
    • changedInput schema / properties / model_id / description
      Previous value: -"ID of the threat model to re-evaluate."New value: +"ID of the threat model to re-evaluate (or estimate for)."
  3. Addedv0.62.2
  4. Removedv0.62.1
  5. Addedv0.60.1

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the full burden and does so: it labels mode='quote' read-only and mode='recompute' mutating, explains that usage is metered as it runs, that subjects already carrying a verdict cost nothing, that work runs in the background, and that a governor.exhausted run is queued and resumes at governor.resets_at rather than dropped. It also warns the operator to see the credit estimate first, which is exactly the kind of cost/side-effect disclosure an agent needs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the verb and the three mode names, then bulleted per mode, with the caveats in a final block. Efficient for the amount of behavior being disclosed, though the per-mode return-object listings border on verbose and overlap the output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-mode, mutating, cost-incurring tool with an output schema, the description covers mode selection, mutation/read-only semantics, cost, background execution, governor exhaustion behavior, and sibling alternatives. The only omission is model_id/server_version semantics, which is minor relative to the completeness elsewhere.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It fully documents the 'mode' enum (including the default and the distinct behavior/return shape of each value), but model_id and server_version — both required-adjacent inputs — get no explanation at all. Mode is the highest-value parameter and is well served, but the remaining gap keeps this at baseline-plus rather than high.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb ('Estimate, force, or retry') over a specific resource ('a model's verdict evaluation'), then breaks the three modes apart so an agent knows exactly which operation it is invoking. It explicitly differentiates itself from siblings like get_sufficiency, get_verdict_divergence, and judge_objective, so the agent can route without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use per mode (quote = read-only estimate, recompute = mutating full re-evaluation, retry_parked = only transient-failure-parked verdicts) and explicit when-NOT-to-use ('A control at partially_verified, or coherence_status: pending ... is not a reason to recompute') plus the correct alternatives for those cases. Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools