Skip to main content
Glama
Mipiti
by Mipiti

Judge Imported Controls

judge_imported_controls

Evaluate imported controls awaiting judgement: preview charges first, then queue judgements for their objectives by confirming the estimate.

Instructions

Have the imported controls awaiting their judgement judged. Mutating only with confirm_estimate=True; may consume credits then.

Controls saved by import_controls are not judged unprompted: the mitigation groups they join credit nothing until this runs. It judges every objective those controls join, priced and charged as judge_objectives prices and charges them.

  1. Call with confirm_estimate=False (the default). Nothing is queued and nothing is charged; the answer carries awaiting_judgement (the control ids), co_ids (the objectives they join), scope, ungrouped and estimate. Show the user the estimate.

  2. Call again with confirm_estimate=True once they agree. The judgements are queued (confirmed: true, queued) and the controls stop awaiting (awaiting_judgement comes back empty).

ungrouped lists objectives with no mitigation group: nothing can be judged there until their controls are grouped with set_mitigation_groups. When nothing awaits, the answer says so and does nothing.

Refusals come back as data, {confirmed: false, queued: 0, http_status, ...}: 409 while a control build is running, 402 when the balance this workspace bills to cannot cover the estimate, 503 when judging is unavailable on this deployment.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
model_idYesID of the threat model.
server_versionYes
confirm_estimateNoFalse (default) returns the estimate and queues nothing; True queues the judgements.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv0.84.0

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so: it discloses that it is mutating only when confirm_estimate=True, that it may consume credits, that pricing matches judge_objectives, that it queues judgements, and it enumerates refusal statuses 409/402/503 returned as data rather than exceptions. This is unusually rich behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the mutating/credit caveat, then uses a numbered two-step list that is easy to follow. Dense but each paragraph (workflow, ungrouped, refusals) earns its place; only minor tightening is possible.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Even though an output schema exists, the description names the relevant return fields (awaiting_judgement, co_ids, scope, ungrouped, estimate, confirmed, queued) and explains refusal shapes, leaving no gap for correct invocation or interpretation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 67%; the description explains confirm_estimate's two modes more fully than the schema does, and ties model_id implicitly to the target threat model. It does not add meaning for server_version, but that param is boilerplate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Have the imported controls awaiting their judgement judged') and scopes it to controls saved by import_controls, distinguishing it from judge_objectives / judge_objective. An agent can identify the distinct role this plays in the control workflow.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives an explicit two-step protocol (estimate first with confirm_estimate=False, then confirm with True), names prerequisites (controls must be grouped, otherwise use set_mitigation_groups), and states the no-op case ('when nothing awaits, the answer says so and does nothing').

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools