Skip to main content
Glama

Plan the gauntlet

run_gauntlet_plan
Read-onlyIdempotent

Build the full pre-submission review plan: get all 8 stages in order with their profiles, tools, required files and Context fields, plus the final packet call.

Instructions

Call when the researcher wants the full pre-submission run. Returns the 8 stages in order, each with its profile, the tool that runs it, the files and Context fields it reads and the instruction for its call, then the Context fields still empty and the build_packet call that ends the run. A hosted stage runs through run_review and needs the connection token; a core stage is answered by your own model. The plan itself needs no account.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
contextNoWhat the researcher states about the finding: the same context object prepare_review takes.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
askNoContext fields a stage reads that are still empty.
stepsYesWhat to do, in order.
finishYesThe build_packet call that ends the run.
stagesYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.7.0

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive, closed-world), but the description adds execution-relevant behavior: hosted stages need the connection token, core stages are answered by the model itself, and the plan requires no account. That auth/execution context is genuinely beyond the structured fields, though it stops short of describing ordering guarantees or failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the invocation condition, then the return shape, then the hosted/core execution caveat. Dense but every clause carries information; it could be marginally tightened but nothing is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given an output schema exists (so return values needn't be spelled out), rich annotations, and a nested optional parameter, the description covers when to call, what the plan contains, and the token/account requirements for following through. An agent has everything needed to call it and act on the result.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the single `context` parameter is documented in-schema as the same object prepare_review takes. The description adds only that empty Context fields are reported back, which is useful framing but not syntax or structure beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (plan/returns the 8 stages) and resource (the gauntlet pre-submission run), and enumerates the returned contents: stages, profiles, running tool, files and Context fields read, call instructions, empty fields, and the terminating build_packet call. It also distinguishes its role from siblings by naming run_review and build_packet.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Call when the researcher wants the full pre-submission run" gives a clear triggering condition and implies this is the entry point for the whole flow. It clarifies which stages route to run_review versus the agent's own model, but offers no explicit when-not or contrast against prepare_review/list_profiles beyond the shared context reference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.