Skip to main content
Glama

verify_repo

Judge a git working tree against a baseline by running pre-existing tests, then return MERGE, BLOCK, or INCONCLUSIVE to gate AI coding changes.

Instructions

Judge the working tree of repo against a baseline, by execution.

claims: pytest node ids (``path::test``) to verify; empty -> discovered from
the tests that executed a changed line. pytest_args: test paths or node ids
the coverage run collects (default: the whole suite); options are refused.
base_ref: used only when the operator did not pin ADVERSARY_GATE_BASE_REF.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
repoYes
claimsNo
base_refNoHEAD
pytest_argsNo
timeout_secondsNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv2.8.0

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It usefully discloses behavior beyond the schema: pytest options are refused, base_ref is ignored when ADVERSARY_GATE_BASE_REF is operator-pinned, and execution is the verification mechanism. However, it says nothing about side effects, permissions, sandboxing, or cost/latency of running a suite.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The one-line purpose is front-loaded, followed by compact per-parameter notes. Formatting is a bit idiosyncratic (line breaks in prose, doubled backticks), but nearly every sentence carries semantic weight and there is little redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need not be explained, and the description covers most invocation semantics for a 5-parameter execution tool. It still omits timeout behavior, side-effect/permission profile, and any relationship to `gate_policy`, leaving meaningful gaps for a tool that executes code.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must document all parameters, and it largely does: claims format (`path::test`) and the empty-value discovery rule, pytest_args scope plus the 'options are refused' constraint, and base_ref's conditional use. Only timeout_seconds is left undocumented, which is a minor gap against solid coverage of the other four.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening line gives a specific verb+resource+mechanism: judge the working tree of `repo` against a baseline by execution. An agent can infer this runs tests to validate changes. It does not differentiate itself from the sibling `gate_policy`, but the purpose itself is clear despite the idiosyncratic verb 'judge'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage context is implied through defaults (claims discovered from changed-line tests, pytest_args defaulting to the whole suite) rather than stated. There is no explicit when-to-use vs. when to prefer `gate_policy`, and no prerequisites for calling this tool are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools