Skip to main content
Glama

Verdict on whether to hand this to a human

should_escalate
Read-onlyIdempotent

Decide whether a human must review an action, plan, diff, or message by checking risk, reversibility, and human-required constraints; use the boolean verdict as a gate before irreversible steps.

Instructions

One call, one boolean answer: does a human need to see this?

Asks three questions at once (risk, reversibility, whether a human is required), then reduces them to the strictest verdict. Use it as a gate in front of an irreversible step.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
stateYesThe action, plan, diff or message under consideration
thresholdNoAs in `decide`

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
gradeYesOverall grade; act autonomously only when this is auto_execute
usageNo
cachedNoTrue when served from the in-process cache
answersYes
routingYes
summaryYes
contractYesDirective: overall grade, per-question grades, and the reasons for them
autonomousYesTrue when every answer cleared the auto-execute bar
latency_msYesServer-side inference time
requires_humanYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.2.0

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already establish read-only, idempotent, closed-world behavior, so the description need not restate safety. It adds genuine behavioral detail the schema cannot: the internal composition of three questions and the reduction rule ('strictest verdict'), which tells the agent how the boolean is derived.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core answer and followed by mechanism and usage. No filler; each sentence carries distinct information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With full schema coverage, rich annotations, and an output schema, the description only needs to convey purpose, mechanism, and usage, all of which it does. The only soft spot is the unexplained `threshold` cross-reference, which is minor given the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so both parameters are documented structurally, and an output schema exists. The description adds no syntax or semantics for `state` or `threshold` beyond a cross-reference ('As in `decide`'), so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific outcome ('one boolean answer: does a human need to see this?') that separates it from the decide_* siblings, which produce richer verdicts. It does not explicitly name which sibling to prefer, but the boolean framing is distinctive enough for selection.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Use it as a gate in front of an irreversible step' gives a clear trigger condition. No explicit when-not or named alternative (e.g., 'if you want the full rationale, use decide'), so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.