Skip to main content
Glama

ddflow_gate_verify

Verify a gate actually fails on mutated input by applying each registered mutation, requiring a non-zero exit, and restoring the file. Catches vacuous passes that report success on every change.

Instructions

Break what a gate guards and require it to NOTICE. Applies each mutation registered on the gate, runs it, requires a non-zero exit, and restores the file.

This is the anti-vacuous-pass check turned on the checks themselves. A gate that cannot fail is worse than no gate: it reports success on every change and everyone downstream reads that as evidence. Exit 1 means the gate did NOT catch its mutation — or that nobody has registered one, which is the same problem earlier.

A mutation whose old text is absent or ambiguous is a FAILURE, not a skip: the edit never happened, so the gate ran on pristine source and passing proves the opposite of what it claims.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
idYesItem whose worktree to mutate in.
gateYesGate id. Must be a command gate.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full responsibility for behavioral disclosure. It clearly states that it mutates the file, runs the gate, expects a non-zero exit (indicating the gate caught the mutation), and restores the file. It also explains the failure semantics: a mutation with absent or ambiguous 'old' text is a failure, and Exit 1 from the tool indicates the gate did not catch the mutation. This is thorough and unambiguous.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is divided into three paragraphs. The first is concise and action-focused. The second explains the rationale behind the tool, which is useful context but could be shortened. The third clarifies a failure condition. Overall, every sentence adds value, but the middle paragraph is slightly verbose. The structure front-loads the core action and then elaborates, which is effective.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with two parameters, no output schema, and no annotations, the description covers the key behaviors: mutation application, exit-code expectations, file restoration, and failure conditions. It does not explicitly state what the tool returns on success (presumably 0) or how multiple mutations are handled sequentially, but these are inferable from the description. The absence of an output schema reduces the need to explain return values, so the description is sufficiently complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides descriptions for both parameters (id and gate), achieving 100% coverage. The description does not add any additional parameter-specific guidance, such as examples or further constraints beyond 'Must be a command gate.' Since the schema handles parameter documentation, the description's lack of extra detail keeps this at the baseline score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a concrete, action-oriented statement: 'Break what a gate guards and require it to NOTICE.' It then specifies the exact procedure (apply mutations, run, require non-zero exit, restore file). This clearly distinguishes it from sibling tools like ddflow_gate_run (which presumably runs gates normally) and ddflow_gate_record (which registers mutations).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description establishes clear context: it is the 'anti-vacuous-pass check turned on the checks themselves.' This implies it should be used when verifying that a gate actually catches its registered mutations. While it doesn't explicitly name alternative tools or state when not to use it, the purpose and scenario are evident from the wording and sibling names. It stops short of explicit exclusions or direct comparisons.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.