Skip to main content
Glama

scout_verify

Re-test open findings from previous runs and record verdicts (gone, present, changed) to clear or confirm them. Use after a fix wave to separate resolved issues from ones still failing.

Instructions

Re-test findings earlier runs left open. With no arguments, returns the open findings in the order to re-test them — worst route first, grouped so a route is walked once — each with its evidence and repro steps. Pass ids to narrow it to specific findings. After re-testing one, call again with id and verdict to record what you saw: "gone" resolves it, "present" stamps it confirmed so the report stops calling it unverified, "changed" keeps it open and says the behaviour differs. Use after a fix wave, or at the start of a run against an app this project has tested before.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
idNoThe finding being verified. Omit to get the worklist.
idsNoNarrow the worklist to these finding ids.
noteNoWhat you saw, in a sentence. Shown in the report beside the verdict.
sessionNoTarget this session directly instead of the active one — pass it explicitly when dispatching to MULTIPLE sessions in one turn (e.g. two scout_click calls with different `session`), which then run CONCURRENTLY rather than queueing. Omit for single-session sequential use.
verdictNoWhat the re-test found: "gone", "present" or "changed". Requires id.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv3.4.0

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden, and it delivers: it discloses ordering (worst route first), grouping, returned content (evidence and repro steps), and the state effects of each verdict—'gone' resolves, 'present' confirms, 'changed' keeps open. This is unusually transparent about side effects.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Every sentence earns its place: the tool's action, the worklist behavior, narrowing, verdict recording, and usage timing. The information is front-loaded with the core purpose and flows logically without repetition or fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with five parameters, no annotations, and no output schema, the description is remarkably complete. It covers what the worklist returns, how to narrow it, how to record verdicts, and what each verdict means, leaving no critical gap for an agent trying to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaning beyond the schema by explaining the interaction between id, ids, and verdict—including that a verdict is recorded by calling again with id and verdict—and clarifies what the verdict values do to report state.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: re-test findings left open by previous runs and record verdicts. It distinguishes itself from siblings like scout_scan or scout_resolve by describing a specific verification workflow with 'gone', 'present', and 'changed' outcomes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit when-to-use context: 'Use after a fix wave, or at the start of a run against an app this project has tested before.' It also explains the call pattern for listing, narrowing, and recording verdicts. It does not mention alternatives or when-not-to-use, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.