Skip to main content
Glama

ddflow_tests

Run the tests your change reaches, in parallel, with a reason for each and a ready command. Use it for fast feedback after every edit; the full suite stays at the gate.

Instructions

AFTER EACH CHANGE, while you work: the tests your change reaches (changed test files, tests importing a changed module directly or one step removed, tests named after a changed file, everything under a changed conftest.py), each with why, and a command that runs them IN PARALLEL -- the project's own test command with its runner and worker flags, the files swapped in. Run it; do not reason about which tests matter. It is fast feedback, never a pass: the unit_tests gate still runs the WHOLE suite, in parallel. Pass item so the diff is taken in that item's worktree against its base. Exit 2 when no test reaches the change.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
baseNoCompare against this ref instead of the item's base.
itemNoThe item whose worktree and base to use.
as_agentNoSubagent sharing the parent's connection: your own stable name, for this call only (see ddflow_identify).

Schema Changelog

Changes observed during successful MCP inspections.

  1. Addedv0.1.10

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and largely meets it: it discloses parallel execution, that output is per-test with rationale plus a command, exit code 2 when nothing reaches the change, and that the diff is scoped to the item's worktree against its base. It does not state whether running tests can mutate the worktree or require prior setup/identification, which keeps it at a 4.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The content is relevant but the delivery is a dense run-on block with nested parentheses, em-dashes and caps, making it hard to scan. The trigger and purpose are front-loaded, but the rest is not structured for fast parsing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations and no output schema, the description covers output shape, exit semantics, execution mode and diff scoping, which is most of what an agent needs. It leaves implicit whether a session/item must exist first and what happens on a dirty worktree.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but the description adds real meaning for `item` (the diff is taken in that item's worktree against its base) beyond the schema's terse 'which item'. `base` and `as_agent` are only covered by the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states precisely what the tool produces: the tests a change reaches (changed files, direct/one-step imports, name matches, conftest.py contents), each with a reason and a runnable parallel command. It is clearly distinguishable from the unit_tests gate, which it names. It loses a point because the verb is buried in a dense clause rather than front-loaded as a clean 'run X' statement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit timing ('AFTER EACH CHANGE, while you work'), an explicit constraint ('Run it; do not reason about which tests matter'), and an explicit boundary against the alternative ('never a pass: the unit_tests gate still runs the WHOLE suite'). This is close to ideal when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools