Skip to main content
Glama

ddflow

A portable, agent-agnostic work-queue kernel for AI coding agents.

You keep a queue of phases and tasks with declared dependencies. You say "implement phase P2". Independent tasks fan out to parallel agents in isolated git worktrees; dependent ones wait. Every task passes a quality pipeline whose gates cannot be passed by assertion. If an agent crashes, its work is found rather than lost. If everything except the log is destroyed, the project's decision history rebuilds from the log alone.

One dependency beyond python3 and git (Jinja2, for the prompt templates; see Extending it by writing text, not code). Works with Claude Code, Gemini CLI, Codex, Copilot, Cursor, Kimi, opencode, Aider, a CI job, a Makefile, or a human at a terminal — over a CLI and an MCP server that are the same implementation.


How do I…?

Every row is a command you can run in a terminal and a tool an agent can call over MCP — the same implementation, so neither drifts from the other.

I want to…

CLI

MCP tool

see what the workflow is

ddflow workflow

ddflow_workflow

change the workflow

ddflow workflow pipeline task … · workflow gate <id> … · workflow drop <id>

ddflow_workflow_pipeline · _gate · _drop

change any setting

ddflow config --explain · --set <key> <value>

ddflow_configure

add a phase / a task

ddflow phase add P1 --title … · ddflow task add P1.T1 --phase P1 --globs 'src/**'

ddflow_phase_add · ddflow_task_add

get a plan into the queue

see From plan mode to the queue

same

know what to work on

ddflow next

ddflow_next

start a task

ddflow claim <id> → work → ddflow gate … → ddflow merge → ddflow complete

ddflow_claim, ddflow_gate_*, ddflow_merge, ddflow_complete

see progress / effort

ddflow progress · ddflow status · ddflow board

ddflow_progress · ddflow_status · ddflow_board

find out if we're going in circles

ddflow loops

ddflow_loops

record a lesson / decision / research / bug

ddflow lesson add · decision add · research · bug found|fixed

ddflow_lesson_add · ddflow_decision_add · ddflow_research · ddflow_bug_*

search everything the project remembers

ddflow recall '<regex>'

ddflow_recall

record what happened this session

ddflow session start|prompt|note|end

ddflow_session_*

read the engineering log

ddflow history

ddflow_history

check the tooling around the gates

ddflow companions

ddflow_companions

find work a crashed agent left

ddflow recover

ddflow_recover

check the project's integrity

ddflow doctor

ddflow_doctor

rebuild everything from the log

ddflow replay --verify

ddflow_replay

invoke a workflow / a mode of your own

ddflow prompts list · prompts show <name>

prompts/list · prompts/get

see what this project left undone

ddflow doctor · ddflow status

the footer on tool results

ask the tool to explain itself

ddflow help [topic]

ddflow_help

Every read command takes --json. Every exit code means the same thing everywhere: 0 healthy · 1 real failure · 2 could not run / nothing to do · 3 coordination refused. 2 is never collapsed into 0 — "nothing is ready" and "everything is fine" are different facts, and an agent that cannot tell them apart invents work.


Related MCP server: agent-tasks

Table of contents


Help: what it can do, and the workflow

$ ddflow help                 # what this is, the loop, every capability grouped
$ ddflow help workflow        # workflow · import · gates · parallel · memory · recovery · config

Reachable as ddflow_help over MCP, and that is the point: an agent connecting had 59 tool descriptions and a state-aware handshake, neither of which answers "what is this, and how am I meant to work here". A tool description explains one tool to someone who already picked it; the handshake describes this repository right now.

Two halves, deliberately:

  • The narrative is a template under ddflow/templates/prompts/help/, so ddflow prompts eject-style overriding applies — put your own .ddflow/prompts/help/workflow.md in place and the tool teaches your workflow.

  • The capability inventory is generated from the live tool table. A hand-kept command list in a second place is the documentation-drift class, and this project has paid for it twice.

Three ratchets keep the prose honest, because a page recommending a flag that was renamed is worse than no page — whoever finds nothing reads the code, and whoever finds a wrong answer trusts it. Every command a page names must exist as a CLI leaf or an MCP tool; every topic the index offers must resolve; and every tool must fall into a group, so a new capability has to be classified rather than quietly dropped from an inventory that claims to be complete.


Two ways to drive it

The CLI is the whole product. The MCP server is a second surface over the same commands, and tests/test_mcp_parity.py fails if the two diverge — every subcommand has a tool, every flag is reachable, and each exemption carries a written reason.

Standalone: a terminal, a Makefile, CI

$ ddflow init
$ ddflow config --set gate.unit_tests.command "python -m pytest -q"
$ ddflow phase add P1 --title "Billing" --globs "src/billing/**"
$ ddflow task add P1.T1 --phase P1 --title "Tax rules" --globs "src/billing/tax.py"

$ ddflow next                          # exit 2 = nothing actionable
$ ddflow claim P1.T1                   # exit 3 = refused, with the reason
leased P1.T1 · worktree .ddflow-worktrees/P1.T1 · branch ddflow/P1.T1

$ cd .ddflow-worktrees/P1.T1 && ...    # do the work
$ ddflow gate status P1.T1             # what the pipeline wants next
$ ddflow gate run P1.T1 unit_tests     # runs it; the exit code IS the evidence
$ ddflow gate record P1.T1 implement --outcome passed --evidence "added tax.py"
$ ddflow complete P1.T1                # exit 3 lists whatever is unsatisfied
$ ddflow merge P1.T1

You get everything except the judgement. Command gates run themselves; agent gates wait for a human to record an outcome, and ddflow gate skip <id> <gate> --reason "..." is the escape hatch — recorded as a skip, never as a pass.

In CI, the exit codes are the interface:

check:
	ddflow doctor        # 1 = integrity problems, each named
	ddflow workflow      # 1 = the pipeline does not hang together
	ddflow cadence       # 2 = no periodic pass is due

2 is never "no problem". A job that treats it as success reports a green build for a suite that never ran.

As an MCP server

ddflow mcp speaks newline-delimited JSON-RPC over stdio. You rarely run it by hand — ddflow adopt writes the launch entry into each agent's own config and leaves existing servers alone:

22 agents are supported. The full table, with what each one gets, is in Wiring it into your agent.

It also copies the driver to docs/ddflow/drivers/, and installs the pre-commit hook that enforces claim-before-you-edit.

What an agent sees the moment it connects, with no call to make:

  • Instructions, returned inside the initialize result itself — and state-aware: what is ready, what is in flight, which setup is missing, whether this project has history worth importing, whether an import was left unfinished.

  • Tools — one per CLI command.

  • Resources — ddflow://board, ddflow://brief, ddflow://lessons, ddflow://research.

  • Prompts — which a client turns into slash commands. Tools are things an agent calls; prompts are things you invoke.

Two tools exist so an agent can orient itself without being told: ddflow_help (what is this, what is the loop) and ddflow_workflow (what are the rules here).

What goes in AGENTS.md / CLAUDE.md

ddflow adopt writes it as a managed block between <!-- DDFLOW:BEGIN --> and <!-- DDFLOW:END -->. Your own prose around it is preserved; re-running updates only what is inside. If you write it by hand, four things have to be in it:

  1. Start every session with ddflow_brief (or ddflow brief in a shell).

  2. Claim before you edit — ddflow_next → ddflow_claim → work in the worktree it creates.

  3. The loop — ddflow_gate_status → satisfy each gate → ddflow_complete → ddflow_merge.

  4. The exit codes, and that 2 is not success.

Without that block an agent sees the tools and has no reason to reach for them before editing. The block is what makes the queue authoritative rather than optional — and it is 232 words, because an instruction file nobody finishes reading is one nobody follows.


The workflow, and changing it

$ ddflow workflow
# The workflow this project runs

   1. research      agent
   2. rules         agent
   3. implement     agent       (required)
   4. lint          command     (required, NOT proven able to fail)
      $ ruff check .
   ...

## The rules, and where each came from

  gates.require_outcome                  True              [default]
  gates.enforce_order                    block             [file]
  schedule.max_parallel_tasks            4                 [default]

One answer to "what are the rules here": every gate in order, which are commands and which you perform, which are required, which need evidence, which need a different-family reviewer, which have been proven able to fail — plus the completion rules, the caps, the reviewers, and where each value came from, so a deliberate choice is distinguishable from a default nobody touched.

Changing it

$ ddflow workflow gate lint --command "ruff check ." --into task --after implement --required
$ ddflow workflow pipeline task research,implement,lint,unit_tests,merge
$ ddflow workflow drop dedupe

All four reach MCP — ddflow_workflow, ddflow_workflow_pipeline, ddflow_workflow_gate, ddflow_workflow_drop — so an agent can change the workflow with the operator's agreement. Their descriptions say to ask first and offer dry_run, because a pipeline governs every future item, not the one in hand.

Nothing is written until it is checked, and the order is the point: compose the change, validate the result, then replace the file atomically.

  • A pipeline naming an undefined gate is refused, naming the near miss. That one is otherwise silent and permanent: the outcome folds to empty, completion refuses it forever, and gate record rejects the id as unknown — so the item can never be completed at all, and nothing says why.

  • An unknown section or knob is refused, with a suggestion. [gatez] is valid TOML and used to be written happily, breaking every later command — the write path validated the merged text for syntax and then validated the config already on disk, which is a writer checking the state it is replacing.

  • Dropping a gate takes it out of required too, or it becomes a requirement that quietly requires nothing.

ddflow workflow and ddflow doctor both re-run those checks against what is on disk. Everything is a file you can also edit by hand: gates in [gate.<id>], reviewers in [[reviewer]], companions in .ddflow/companions.toml, and every prompt — including the instructions your agent receives at connect — under .ddflow/prompts/.

One caveat with MCP: the connection instructions are computed once, when the server starts. A workflow changed mid-session is live for every tool call immediately, but the text the agent was handed is stale. Tell it to call ddflow_workflow, or restart.


Why it is built this way

The append-only event log is the source of truth; everything else is a projection that can be deleted and re-derived. The SQLite index, the markdown boards, the search index, the recovery bundle — all disposable, all rebuilt by ddflow rebuild.

That single inversion is what makes the four hard properties fall out for free rather than needing to be engineered:

You get

Because

Two agents on two branches never conflict

Each appends to its own file. Measured: a real two-branch merge resolves clean.

A crashed agent loses nothing

State is folded, never written. Nothing is half-updated.

The project rebuilds from the log

Operator prompts are events.

An edited history is detectable

Event ids are content addresses.

The design decisions, with the probes that settled each, are in docs/RESEARCH.md. The two that most shaped it:

  • SQLite-on-NFS is correct here but 36× slower than local (measured, 12 processes × 40 increments). So the log is authoritative and the database is a disposable cache — which also happens to be the choice that stays correct on filesystems where locking is broken.

  • An expired lease must never be reclaimed automatically. A crashed agent's worktree is sometimes irreplaceable work and sometimes a superseded draft, and nothing in the metadata distinguishes them. Recovery measures and advises; it never deletes.


Install into any project

One line in your agent's MCP config. Nothing else.

{ "mcpServers": { "ddflow": { "command": "uvx", "args": ["ddflow-mcp"] } } }

uvx fetches and runs the published package in an ephemeral environment on first use — no clone, no virtualenv, no PYTHONPATH, no install step for an operator to forget, and no vendored copy to drift from upstream. ddflow has zero runtime dependencies beyond python3 and git, which is what lets it install inside sandboxes, CI images and other tools' ephemeral containers.

Then, from the agent, with no shell at all:

Call

What it does

ddflow_setup

creates .ddflow/, writes the driver and the AGENTS.md section

ddflow_configure with toml: '[gate.unit_tests]\ncommand = "pytest -q"'

sets your test command

ddflow_reviewers_detect with write: true

finds a local model server and registers it as a cross-family reviewer

ddflow_phase_add, ddflow_task_add

fill the queue

ddflow_brief

start every session here

That is the whole adoption. The per-project instruction text is 232 words — a managed block in AGENTS.md, because the MCP tool descriptions already carry the how, and a second copy of that would drift from the one the model actually reads.

uv tool install ddflow-mcp        # or: pipx install ddflow-mcp
cd /path/to/your/project
ddflow adopt            # every supported agent
ddflow adopt --agents claude,cursor,vscode,kimi   # or name the ones you use

adopt is idempotent and writes managed blocks, so re-running after an upgrade updates them and leaves your own prose alone. It writes the MCP registration into each agent's own config location, merged with whatever servers are already there. From a source checkout it points the config at that checkout instead of the published package, so developing ddflow does not silently configure your project against the released version.

Docker — for operators with no Python toolchain

{ "mcpServers": { "ddflow": { "command": "docker", "args": [
    "run", "-i", "--rm",
    "-v", "${workspaceFolder}:/repo",
    "--add-host=host.docker.internal:host-gateway",
    "ghcr.io/delian/ddflow-mcp:latest" ] } } }

ddflow adopt --launch docker writes exactly that. The image is 107 MB (Alpine; ddflow is pure standard library, so there is no compiled dependency to worry musl about) and behaves identically on Linux, macOS and Windows.

Four things go wrong when a containerised tool touches a bind-mounted git repo. All four are silent, one of them loses work, and all four are handled:

Trap

What it looks like

Handled by

Worktrees land outside the mount

worktree.root defaults to ../.ddflow-worktrees, a sibling of the repo. In a container only the repo is mounted, so worktrees go to the ephemeral layer and are destroyed on exit with the agent's uncommitted work inside them.

container.default_worktree_root relocates a sibling root to .ddflow-worktrees inside the repo, and adopt gitignores it

Root-owned files

On a Linux bind mount the operator needs sudo to edit their own project afterwards

the entrypoint reads the mount's uid/gid and su-execs down to it

git refuses the mount

"detected dubious ownership", surfacing as an unexplained ddflow failure

safe.directory set in the entrypoint

No git identity

git commit fails with "Please tell me who you are"

entrypoint prefers GIT_AUTHOR_*, then the repo's own config, then a clearly-marked placeholder

And one that cannot be fully handled, so it is reported: 127.0.0.1 inside a container is the container. A model server on your own machine is not reachable from there. ddflow rewrites loopback reviewer URLs to host.docker.internal, and ddflow doctor tells you that on Linux you must also pass --add-host=host.docker.internal:host-gateway, because unlike Docker Desktop the Linux engine does not provide that name.

The related portability fix: worktree paths are stored in the event log relative to the repo root. The log is committed and shared, so an absolute path is true only on the machine that wrote it — false for a teammate who cloned elsewhere, for CI, and for a container where the repo is /repo. Pinned by test_the_event_log_carries_no_absolute_paths.

Extending it by writing text, not code

Every prompt is an external template, resolved config → project → shipped:

ddflow prompts list              # where each template currently comes from
ddflow prompts eject             # copy the shipped ones into .ddflow/prompts/
$EDITOR .ddflow/prompts/review_system.md

Adding a mode of your own: [[macro]]. Overriding a shipped workflow needs no code, and neither does adding one. A macro is a named, parameterised prompt — "enter debugger mode" — that appears everywhere the shipped workflows do: prompts/list and prompts/get over MCP, which is what a client turns into a slash command, and ddflow prompts list|show in a terminal.

# .ddflow/config.toml   (or .ddflow/macros.toml, if you prefer to split it out)
[[macro]]
name = "debugger"
title = "Enter debugger mode"
description = "Reproduce first, then bisect. No fix without a failing probe."
params = ["symptom"]                                  # required, not optional
tools  = ["ddflow_bug_found", "ddflow_gate_run", "ddflow_bug_fixed"]
prompt = """
You are debugging: {{ symptom }}

Reproduce it before you theorise. Paste the command and its output.
"""

Use prompt_file = "docs/modes/debugger.md" instead for anything long enough that TOML quoting gets in the way.

When to reach for a macro rather than a gate. A gate is a step every item passes through, recorded against that item and blocking its completion. A macro is a MODE an operator enters, belonging to no item and recorded nowhere — "audit this release", "handle this incident". If the thing should hold up a task until it is done, it is a gate; if it is a way of working you want to name and re-enter, it is a macro. Putting a mode in the pipeline makes every task wait for something that was never about that task.

tools is declarative, not a sandbox. It is rendered into the prompt as the ordered set the mode expects, so the agent is told what the mode is for and the next reader can tell what it was supposed to do. It does not restrict what the agent may call — MCP has no mechanism for that, and claiming a security property this cannot honour would be worse than not having it. This is the deliberate departure from dx-zero/mcpn, whose toolMode: situational lets the model pick freely from a bound set with no recorded ordering: a session you cannot replay is a session you cannot review, which is the property the event log exists to give you.

Three things a macro refuses, because each alternative fails quietly: a missing parameter (a prompt with a hole in it reads as a complete instruction), a name that belongs to a shipped command (silent shadowing leaves you editing a block that does nothing), and both prompt and prompt_file (two sources for one body means one is dead and looks live).

Including the one the agent actually reads first. mcp_instructions.md is the block an MCP client injects into the model's context on connect — the workflow, the reporting duties, and which companion tools to reach for. It is the file to edit when you want this project to work differently:

ddflow prompts eject mcp_instructions
$EDITOR .ddflow/prompts/mcp_instructions.md      # or [prompts] mcp_instructions = "..."

It renders against the live state — adopted, task_pipeline, setup_todo, companions, missing_companions, gate_gaps, recoverable, loops — so the instruction is the next concrete action rather than a fixed blurb the model learns to skip. A broken override says so in the instruction block itself instead of falling back to the default: this is the one surface where nobody would ever notice their edit was not live.

Templates render with Jinja2, which is ddflow's one runtime dependency, and with a strict standard-library renderer when it is absent — a stripped deployment with no reachable package index still starts. The shipped templates use the subset both engines agree on, and tests/test_template_engines.py walks the template REGISTRY, rendering every entry through both engines and asserting the outputs are byte-identical.

That test is iterated rather than hand-listed for a reason. Its predecessor named three templates in a dict, mcp_instructions.md was never added, and in 0.1.1 the largest and most important template rendered correctly under Jinja2 and failed under the fallback — so the entire MCP handshake for an unadopted repository, the first thing a new user ever sees, degraded to ddflow's instruction template could not be loaded. Jinja2 was not a declared dependency at the time, so developers had it and the project venv did not: python -m pytest was green and uv run pytest was red on the same commit.

The fallback now raises on any construct it does not implement rather than copying it through. The old regex engine emitted what it could not parse, so a condition as ordinary as {% if a or b %} — which its single-name pattern never matched — reached the client as literal template source.

Both renderers are strict about undefined variables: a prompt silently missing the diff it was supposed to carry is the vacuous review in template form — the model dutifully reviews nothing and reports no findings.

The rest is TOML: gates and their pipelines ([gate.*], gates.task_pipeline), reviewers ([[reviewer]]), companions ([[companion]]), enforcement ([enforce]), cadences, and the rest of the 75 knobs. ddflow config --set <key> <value> edits one key in place, preserving comments.

Publishing and registry

Nobody should have to paste JSON into an IDE to use this. server.json is the MCP registry manifest (io.github.delian/ddflow-mcp), and publishing it is what makes ddflow findable in the VS Code and Cursor marketplaces rather than something you configure by hand. It offers three ways to run the same server, so a client picks whichever it supports:

Package

Identifier

For

pypi

ddflow-mcp, runtimeHint: uvx

Anything with uv — no clone, no install step

oci

docker.io/delian/ddflow-mcp:<version>

Operators with no Python toolchain

oci

ghcr.io/delian/ddflow-mcp:<version>

The same image, no Docker Hub account needed

The image is built for amd64 and arm64, because an Apple-silicon operator running it under emulation pays that cost on every tool call, and tool calls are all this server does.

How CI authenticates — four mechanisms, one stored secret:

Target

Mechanism

Stored secret?

Setup

PyPI

OIDC trusted publishing (id-token: write)

No

Add a trusted publisher on PyPI, once

ghcr.io

GITHUB_TOKEN, injected per run, expires with the job

No

none

Docker Hub

DOCKERHUB_USERNAME + DOCKERHUB_TOKEN

Yes

Create an access token, add both secrets

MCP registry

GitHub OIDC — proves control of the account that owns the io.github.delian/* namespace

No

none

tag + release

GITHUB_TOKEN (contents: write)

No

none

Docker Hub is the only one that needs a long-lived credential, because it has no OIDC equivalent. Use an access token scoped to this repository, never an account password. If that is one secret too many, delete the Docker Hub login and its two tags — ghcr.io alone satisfies the OCI entries a marketplace needs, and server.json lists both so a client picks whichever resolves.

environment: release on the publishing jobs is a control worth knowing about: point it at a GitHub environment with required reviewers and every release waits for a human, with no change to the workflow.

Order matters and the workflow encodes it. mcp-publisher validates that every package named in the manifest exists, so the registry step runs after both PyPI and Docker — publishing the manifest first would advertise a version nobody can fetch.

Four things gate a release, and each exists because the failure it catches is public and irreversible:

  • the tag, pyproject.toml, server.json's version and every OCI identifier's tag must agree — a :0.1.0 left behind while version moved on publishes a manifest pointing at the previous image, installable and wrong;

  • the full suite, plus the slow end-to-end scenarios, which -m 'not slow' otherwise excludes from every ordinary run;

  • the wheel must install into a clean venv and run, and carry its templates — uv build succeeding proves the metadata parses, not that ddflow help works;

  • the image must answer initialize over stdio. A built image that cannot is a broken release every marketplace will happily offer.

Cutting a release

$ scripts/bump.sh patch          # 0.1.0 -> 0.1.1, in all FIVE places that declare it
$ scripts/release.sh             # build + verify everything locally; publishes nothing
$ git commit -am 'release 0.1.1' && git push origin main

That push is the whole release. CI publishes PyPI, Docker Hub, ghcr.io and the MCP registry, then creates v0.1.1 and a GitHub release — last, and only once every publish succeeded, because a tag pointing at a half-release is worse than no tag: it looks authoritative.

The version bump is the release decision, and it is deliberate on purpose. A push to main publishes exactly when that number changes. Publishing on every push is arithmetic that does not work — PyPI refuses to re-upload a version, so the second push fails and every one after it — and deriving a unique version per commit instead would mean an irreversible release for a README typo. So one reviewable line in a diff decides, and everything after it is automatic. workflow_dispatch with force: true is there for the case where you need to republish deliberately.

The version lives in five places — pyproject.toml, server.json's version, its per-package version, the tag inside every OCI identifier, and SERVER_INFO, which is what the server tells every client it is. scripts/bump.sh moves all five and then re-reads them to check it did; tests/test_packaging.py fails if they ever drift. (That test caught the bump script missing SERVER_INFO on its first run.)

scripts/release.sh runs all of that locally and publishes nothing. It is dry by default, needs no credentials, and exists because a tag is not reversible: PyPI refuses a re-upload, :latest is on someone's disk before you notice, and a registry manifest is what an IDE offers people. If it fails on your laptop, the tag was going to fail an hour later in public. --publish is the escape hatch for when CI is unavailable, and it makes you type the version to confirm.

What CI checks

.github/workflows/ci.yml runs on every push and pull request, in four jobs that fail for different reasons so you can tell at a glance which:

Job

Checks

quality

ruff check + format --check; the wheel installs into a clean venv, runs, and carries its templates; gitleaks over full history; bandit over the package; a dependency audit that also asserts the runtime dependency list is still empty

tests

The suite on Python 3.11 and 3.13 — the floor and the current release, because a version-specific break is a break for somebody

codeql

GitHub's security-and-quality queries, landing in the Security tab rather than a log

scenarios

The slow end-to-end runs, and the concurrency/load suite, each as its own step with if: always()

Two of those exist because of specific failures. The wheel check is there because uv build succeeding proves the metadata parses, not that ddflow help works — a wheel missing its templates fails on the user's machine. And gitleaks is there because this project has already committed a live API key: a secret in git is a leaked secret, rotation is the only remedy, so the check that matters is the one that runs before every push.

bandit deliberately skips tests/, which use subprocess and temporary paths constantly and by design. A scanner that cries wolf on every fixture is a scanner nobody reads.

Any LLM as a reviewer — local, remote, SaaS, or a CLI

The critic and rubber_duck gates are run by ddflow, not claimed by the agent. Point them at whatever you have:

ddflow reviewers presets            # 19 ready-made provider settings
ddflow reviewers add --preset ollama --model qwen3:8b
ddflow reviewers detect --write     # probe local ports and register what is serving
ddflow reviewers test               # send a known-buggy diff, check the reply

Four backends, because "any LLM" means four wire formats in practice:

kind

Reaches

Examples

openai (default)

anything OpenAI-compatible — which is most things

ollama, vLLM, LM Studio, llama.cpp, sglang, LiteLLM, OpenAI, DeepSeek, Groq, Together, Fireworks, Mistral, OpenRouter, xAI

anthropic

the Messages API (system is a top-level field, not a message)

Claude

gemini

generateContent (key in the query string, not a header)

Gemini

command

anything at all — a CLI that reads a prompt on stdin and writes the reply to stdout

claude -p, gemini -p, codex exec, llm -m, your own script

command is the escape hatch that makes the answer to "can it use X?" always yes: a model with no HTTP API, behind a corporate gateway, or wrapped in an in-house tool is still usable, with no SDK and no dependency.

[[reviewer]]
name   = "local-qwen"
kind   = "openai"
base_url = "http://127.0.0.1:11434/v1"
model  = "qwen3:8b"
family = "alibaba"              # must differ from the author's family
gates  = ["critic"]
# Optional: start it if it is not already running.
launch = { command = "ollama serve", ready_url = "http://127.0.0.1:11434/v1/models" }

[[reviewer]]
name    = "claude-via-cli"
kind    = "command"
command = "claude -p --model {model}"
model   = "claude-sonnet-5"
family  = "anthropic"
gates   = ["rubber_duck"]

Auto-launch is opt-in per reviewer — starting a multi-gigabyte model server as a side effect of asking for a code review is a surprise nobody wants by default. When it fails it never leaves a half-started process behind, because a reviewer stuck "starting" forever is indistinguishable from one that is down except that it also holds a process.

Keys are never written to the config. Only api_key_env, the name of an environment variable — the config file is committed, and a key in git is a leaked key.

Every way of not reviewing is reported distinctly, with its remedy: no key names the variable, a missing CLI names the binary, a dead server names the launch block you could add, a non-zero exit shows stderr, and empty output on exit 0 is UNAVAILABLE rather than "no findings" — the vacuous pass arriving by the most innocent-looking path there is.

Reasoning models need a large max_tokens. Default 32000, measured not guessed: on Qwen3.8-Flash-Next over a 30 KB diff, a 6000-token budget produced zero characters of content — the whole budget went to reasoning and the reply was truncated. That case is reported as TRUNCATED with the remedy named, never as an empty completion and never as a clean review.

Companion tools

ddflow imposes the order and demands the evidence. It does not perform the judgement inside most of its gates: standards wants an automated standards review, research wants documentation to check a claim against, rules wants memory of the last time somebody hit this. A project that installs ddflow and stops has those gates wired to nothing — and because an agent gate passes on an assertion, that gap is invisible in exactly the way the rest of this design exists to prevent.

So the gap is named:

$ ddflow companions
Companion tools

  [x] context7   Current library documentation
       gates: research, standards
       registered for: claude, cursor
  [x] roborev    Automated second-opinion code review
       gates: standards, bug_hunt, dedupe
       installed (roborev 0.9.1). A cli tool — the agent shells out to it, so
       there is nothing to register.
  [+] codeguide  Language and framework coding standards
       gates: standards
       installed (…) but no agent is configured to launch it.
       -> ddflow companions add --id codeguide
  [ ] sequential Structured step-by-step reasoning
       gates: research, rubber_duck, bug_hunt
       not here: `npx --no-install @modelcontextprotocol/server-sequential-thinking` exited 1
       -> ask the operator, then: npx -y @modelcontextprotocol/server-sequential-thinking

Gates in this project's task pipeline with no companion behind them:
  rules, implement, rubber_duck, critic, unit_tests, bug_hunt, dedupe, merge

Three states, reported separately because the remedies differ: registered, installed but not wired up (one command away), not installed (with the command and the URL). ddflow adopt prints the same summary, so the gap is visible at adoption rather than discovered six tasks later. Exit 2 when a default companion is missing — "no data", never collapsed into "no problem".

Serves

Why

roborev (cli)

standards, bug_hunt, dedupe

Cross-file duplication analysis, which is the failure mode of agent-written code specifically: an agent changing replicated logic reliably updates one copy and misses the rest

codeguide

standards

Checks against a written standard instead of the reviewer's taste

context7

research, standards

A model's memory of a library's API is exactly the kind of claim that is cheap to check and often wrong

memory

rules

Operational facts about this machine — ddflow's own recall covers the project's memory, which is a different thing and belongs in the committed log

sequential

research, rubber_duck, bug_hunt

The three gates that are reasoning, not tool-running. A thought can be marked a revision or a branch instead of being appended to a transcript that only grows — so a retracted hypothesis reads as retracted, and what a bug hunt ruled out stays visible

optmem (cli)

rules

Append-only cross-session memory that compresses as it grows. The other half of memory: recall answers "what did this project decide and learn", OptMem answers "what does this environment do"

Servers and command-line tools are different things, and the registry says which: kind = "mcp" is registrable into an agent's config, kind = "cli" is a tool the agent shells out to. OptMem is the live example — a real tool with no MCP mode, so companions add refuses it and says why instead of writing a launch entry that would fail its first handshake. A cli companion counts toward its gate's coverage once it is installed; registered is a state it cannot reach.

Your stack needs servers this registry cannot know about. ddflow prompts show research-companions walks an agent from the pipeline's uncovered gates, through the repository's actual manifests, to candidates checked against their primary sources — provenance, maintenance, what they execute, what credential they want — and produces [[companion]] blocks you can read and delete. It proposes; you install. A rejection is part of its report, so the next session does not re-research it.

Registering is previewable. ddflow companions add --dry-run (and ddflow_companions_add with dry_run=true) reports the exact config entry it would write and writes nothing — not the file, not even its parent directory. The handshake tells an agent to dry-run first and show the operator the actual entry rather than a description of it, because registering changes which processes their agent launches. The preview is asserted to match what the real write produces; a preview that drifts from the write is worse than none, since the operator has now signed off on it.

ddflow never installs anything itself — running an install command on someone's machine is the operator's decision. What it does instead is instruct the agent to ask: the MCP instruction block lists each missing companion with the gates it serves and the exact command that would install it, and tells the agent to put that to the operator early, install it if they agree, and record the affected gates unavailable if they decline. Never on its own word.

companions add also refuses to register a server that is not present: that writes a launch command which fails mid-task, at the moment a gate told the agent to reach for it. Detection is read-only and bounded — and when it has not run, the state is reported as unknown, not as absent. ddflow companions probes; the MCP handshake does not, because making an agent wait on npx before it can do anything is the wrong trade.

Adding a fifth is a TOML block in .ddflow/companions.toml, not a patch:

[[companion]]
id      = "my-linter"
title   = "House linter"
gates   = ["standards"]
detect  = ["my-linter", "--version"]
command = "my-linter"
args    = ["mcp"]
install = "cargo install my-linter"

Wiring it into your agent

ddflow adopt --agents claude,cursor,codex writes everything below. This table is what it writes, so you can check it or do it by hand.

Agent

--agents

MCP config it writes

Rules

Claude Code

claude

.mcp.json

CLAUDE.md + AGENTS.md

Gemini CLI

gemini

.gemini/settings.json

AGENTS.md

Codex CLI

codex

.codex/config.toml

AGENTS.md

GitHub Copilot (CLI + cloud)

copilot

.github/mcp.json

AGENTS.md

VS Code (any agent)

vscode

.vscode/mcp.json

AGENTS.md

Kilo Code / Roo

kilo

.kilo/kilo.json

AGENTS.md

Cursor

cursor

.cursor/mcp.json

.cursor/rules/ddflow.mdc + AGENTS.md

Kimi Code CLI

kimi

.kimi-code/mcp.json

AGENTS.md

opencode

opencode

opencode.json

AGENTS.md

ZCode (GLM / Zhipu)

glm

.zcode/config.json

AGENTS.md

Qwen Code CLI

qwen

.qwen/settings.json

AGENTS.md + pointer in QWEN.md

Google Antigravity

antigravity

.agents/mcp_config.json

AGENTS.md

Devin CLI

devin

.devin/mcp_config.json

AGENTS.md

Qodo Command

qodo

mcp.json

AGENTS.md

Tabnine

tabnine

.tabnine/agent/settings.json

AGENTS.md + pointer in .tabnine/guidelines/

7 more are supported with no MCP file to write — a verified absence, not an unresearched gap. adopt writes the delta doc and the AGENTS.md block and names the one manual step. Inventing a path would be worse: ddflow would write a file the agent never reads, and you would believe it was wired up.

Agent

--agents

Add the server here by hand

Rules

Aider

aider

no MCP client support at all — drive it from the CLI

AGENTS.md, loaded via read: in .aider.conf.yml

Cline

cline

global settings only; add via its MCP Servers panel

AGENTS.md + pointer in .clinerules/

Windsurf / Cascade

windsurf

global ~/.config/devin/mcp_config.json

AGENTS.md

Replit Agent

replit

web UI only, remote servers by URL — use the CLI here

AGENTS.md + pointer in replit.md

OpenHands

openhands

Settings → MCP (its config.toml form is dev-only)

AGENTS.md

Goose

goose

user YAML ~/.config/goose/config.yaml, under extensions:

AGENTS.md

Sourcegraph Cody

cody

the editor's settings.json, key cody.mcpServers

AGENTS.md ⚠ its own convention is undocumented

One set of rules, every agent

AGENTS.md is the cross-agent convention and most of the 22 read it. Seven do not read it first, or at all, so adopt writes the same managed block into their own surface too:

Agent

Its own surface

Why AGENTS.md alone is not enough

Cursor

.cursor/rules/ddflow.mdc

project rules outrank AGENTS.md

Qwen Code

QWEN.md

QWEN.md is its DEFAULT context file

Cline

.clinerules/ddflow.md

reads .clinerules/, not AGENTS.md

Tabnine

.tabnine/guidelines/ddflow.md

reads .tabnine/guidelines/*.md

Replit

replit.md

its own root-level convention

Goose

.goosehints

CONTEXT_FILE_NAMES is configurable

Aider

.aider.conf.yml read:

discovers nothing automatically

The rules are inlined, not pointed at. A one-line "see AGENTS.md" stub was the obvious design, and ddflow's own notes had already refuted it: a link is only followed if the agent chooses to follow it. A rule that binds only when the model feels like opening a file is not an enforced rule.

That means several copies of one text, and the answer is that a check owns them: one generator, a managed DDFLOW:BEGIN/END block in each, and ddflow doctor comparing every copy against the generator. Five kinds of break are reported and each fails doctor:

$ ddflow doctor
note:    QWEN.md's ddflow section is from an older version and has drifted
note:    .clinerules/ddflow.md exists but its ddflow section was removed
PROBLEM: .goosehints does not exist — the agent has no project rules at all
PROBLEM: .cursor/rules/ddflow.mdc exists but does not bind: `alwaysApply` is not
         true, so the agent may never load it
PROBLEM: .aider.conf.yml exists but does not bind: it does not list `AGENTS.md`
         under `read:`, and Aider loads no instruction file it was not told to load

Files the project already owns — QWEN.md, replit.md, .goosehints — get a block merged into them; your own content stays. Aider's read: list is extended, not replaced. Adopting is idempotent: re-running never appends a second block.

Every other agent reads AGENTS.md directly, which is the point of it being canonical.

Cursor gets its own rules file because its precedence puts project rules above AGENTS.md — writing only AGENTS.md there would be writing to a file the agent outranks. adopt merges into these files rather than overwriting: they hold your other servers and your other rules, and a tool that stomps them is a tool you run once.

The MCP entry is one line in any of them:

{ "mcpServers": { "ddflow": { "command": "uvx", "args": ["ddflow-mcp"] } } }

uvx fetches and runs it in an ephemeral environment on first use — no clone, no PYTHONPATH, no install step to forget. Prefer Docker? docker run -i --rm -v "$PWD:/repo" ghcr.io/delian/ddflow-mcp, which needs the repo bind-mounted because ddflow operates on your actual git checkout.

Standalone, with no MCP at all, is a first-class mode rather than a fallback. Add to AGENTS.md / CLAUDE.md:

This project's work is a queue managed by ddflow. Before doing anything, run
`ddflow brief`. Claim before you edit (`ddflow claim <id>`), satisfy every gate
(`ddflow gate status <id>`), then `ddflow merge` and `ddflow complete`.
Never pass a gate you did not perform — record `unavailable` with the reason instead.

That is the whole integration. An agent with nothing but a shell can drive the entire workflow, which is why MCP is a convenience layer here and never a requirement.


From plan mode to the queue

Agents plan well and forget reliably. A plan that lives in a chat transcript is gone at the next session; a plan in the queue survives, fans out to parallel agents, and carries its own gates.

Tell the agent, at the end of planning:

Put that plan in ddflow before you build any of it. One phase for the whole plan, one
task per independently-shippable step. Declare each task's globs — the files it will
write — and its needs, the tasks that must finish first. Then show me `ddflow next`.

What the agent does with that:

$ ddflow phase add P3 --title "Rate limiting"
$ ddflow task add P3.T1 --phase P3 --globs 'limiter/**'      --title "token bucket"
$ ddflow task add P3.T2 --phase P3 --globs 'api/middleware/**' \
      --needs P3.T1 --title "wire it into the request path"
$ ddflow task add P3.T3 --phase P3 --globs 'docs/**' --needs P3.T2 --title "document it"
$ ddflow next
Ready (1 ready, 0 running, 2 blocked):
  P3.T1  token bucket
      writes: limiter/**
  (blocked) P3.T2: deps — P3.T1 is open
  (blocked) P3.T3: deps — P3.T2 is open

The two fields that do the work are --globs and --needs. Globs are how two agents are stopped from editing the same file: claim refuses an item whose writes overlap one already held, and names what to take instead. Needs are how ordering is enforced without anyone remembering it. A plan whose tasks declare neither is a list, not a queue — it will look parallel and then two agents will fight over one file.

ddflow split <id> --into a,b,c exists for when a task turns out to be three, which is the normal case rather than a failure of planning.


When a companion is missing

ddflow imposes the order and demands the evidence. It does not perform the judgement inside most gates — that is what the companion tools are for. So the honest question is what happens when one is absent, and the answer is deliberately never "the gate passes".

Companion

Serves

If it is missing

roborev (cli)

standards, bug_hunt, dedupe

Record the gate unavailable with the reason. A second opinion is missing and the log says so.

codeguide

standards

The standards gate falls back to the reviewer's taste. Still recordable — but say which it was.

context7

research, standards

Claims about a library's API rest on the model's memory, which is exactly the claim that is cheap to check and often wrong.

sequential-thinking

research, rubber_duck, bug_hunt

A retracted hypothesis becomes one more assertion in a linear transcript, and what you ruled out disappears.

OptMem (cli)

rules

ddflow recall still covers the project's memory — decisions, lessons, research, bugs. What is lost is memory of this machine.

The rule, and it is enforced: a gate whose tool could not run is recorded unavailable with the reason, never passed. ddflow complete reports those as a coverage gap on the completion event, so a finished item never silently implies that a check happened. Set [gates].unavailable_is_failure = true and a gap blocks completion outright.

ddflow companions reports four states, and the difference between the last two is the whole point: registered, installed but not wired up (one command away), missing (with the install command and the URL), and not checked — because the MCP handshake does not probe, and "nobody looked" must never render as "not there".

ddflow never installs anything. Detection is read-only and the report is advice. ddflow companions add --dry-run shows the exact config entry it would write, so an agent can show you the change before making it.


Adopting a project that already has history

A queue that starts empty tells the next agent "nothing is in flight" about a repository with three branches in flight and forty open items in a todo file — and the agent believes it, because the tool said so. That is worse than having no tool at all.

$ ddflow import                      # looks; writes nothing
What this project already has (nothing written yet):

  314 phase(s):
    [ ] 142.A     the scaling-law advisor is wrong (P0; CONFIRMED)   docs/todo.md:26517
  1170 task(s):
  442 lesson(s):
  ...
  47 memory(s):
    [ ] M-0002    Hardware: 8x H200 GPUs on this box, usually idle.  .agent_memory/LOG.txt:3

  NOTE: 3631 already-ticked task(s) were NOT imported. They are history, not a queue.
  NOTE: 32 phase heading(s) say the work is finished while their checkboxes are still
        unticked: 99 (4 open), 103 (3 open), ... Ask the operator which is stale.

$ ddflow import --apply              # writes them, each recording its source line

Seven sources, all optional, all in the places projects actually keep them:

Source

Read from

Becomes

Todo checklists

docs/todo.md, docs/todo/open/*.md, tasks/todo.md, TODO.md, docs/plan.md, ROADMAP.md

phases and tasks, with declared Needs:/Globs:

Lessons

docs/lessons.md, LESSONS.md, docs/retrospectives/*.md

lessons, searchable by ddflow recall

Decisions

docs/adr/*.md, docs/decisions/*.md

decisions, Superseded preserved as superseded

Research

docs/RESEARCH.md

research notes, CONFIRMED/REFUTED/THEORETICAL carried across

Journal

docs/log/*.md, CHANGELOG.md, docs/journal/*.md

session notes, dated by when they happened

Cross-session memory

.agent_memory/LOG.txt (OptMem), .memo/, .optmem/

session notes, with each record's own date

In-flight work

branches with commits not on the base

tasks, named with how far ahead they are

What it will and will not decide for you

Mechanical, and verifiable: a ticked checkbox is a fact, a ## heading is a section, a branch with unmerged commits is work. The id in ### 142.A — … or - [ ] **WFOPT.4.6** — … is read, not invented, so the imported queue uses the ids the project has been writing in commit trailers for months.

Judgement, and yours: which open items are actually live, what each task writes, what depends on what. The /import-existing-project prompt walks an agent through that with the operator. It is not automatable, and a confident guess produces a wrong queue the scheduler then hands out.

Four guard rails, each of which exists because the alternative is silent:

  • Dry run by default. --apply writes. Looking is free and never a side effect.

  • Finished work stays out — it is history, not a queue — except a completed item that open work depends on, which comes along as done so the open item is not stranded on an id the queue has never heard of.

  • [importer] max_tasks (default 200) refuses a whole history. An import writes events into a log that is committed to git; one real repository yielded 4,799 checkboxes. Over the cap it proposes none and says so — the phases are withheld with them, because a queue of empty phases is not a smaller import, it is a misleading one.

  • Idempotent. Ids derive from the source, so re-running after you edit the todo adds what is new and leaves the rest alone. A second run over an unchanged project exits 2.

Verifying an import, at any time

The import's weak spot was never the parsing. It is everything after --apply: 1,170 tasks arrived in the real-corpus run, and the workflow prompt tells an agent to give each one globs and declare its dependencies. Nothing checked whether that ever happened — and an imported queue nobody finished misrepresents the project exactly as an empty one does, believed harder because a tool produced it.

$ ddflow import --verify
Imported between 2026-09-25 and 2026-09-25:

       4 decision(s)
    1727 journal(s)
     442 lesson(s)
      47 memory(s)
     314 phase(s)
      72 research(s)
    1170 task(s)

Left to decide or fix:
  - 1078 imported task(s) declare no globs, so the conflict detector cannot protect
    them and two agents can be handed the same file: OPIK.1b, OPIK.2, ...
  - 27 phase(s) say the work is finished while a task under them is still open: 99,
    103, 115.D.2, ... Ask the operator which is stale before anyone claims from them.

Three answers, three exit codes, because collapsing them loses the one that matters:

Exit

Meaning

0

imported, still matches the sources, and every imported task says what it writes

1

imported — and here is what a human still has to decide

2

nothing was ever imported. An answer, not a failure

It reports status (what is imported, per kind, and when), whether it is still true (what a re-run would add, which sources yielded nothing, which source files have since vanished), and whether anyone finished it (tasks with no globs; phases whose heading claims SHIPPED over an open task).

It deliberately does not repeat ddflow doctor, which already reports unresolved dependencies, duplicate globs and cycles. Two commands reporting one defect in different words is how an operator learns to read neither.

Provenance is a field, not prose. Item.source is docs/todo.md:41; the body still says "Imported from docs/todo.md:41." for a human reading ddflow show. Answering "which items came from the import" by regexing that sentence would mean the day someone rewords it, the count silently becomes zero and the verification passes.

The connection handshake follows through. The offer to import stops once the queue has anything in it — but if imported work is still missing globs, or a phase still claims SHIPPED over open tasks, the MCP instructions say so and tell the agent to run ddflow_import_verify before handing any of it out. That check is computed from the already-folded queue, so it costs nothing; the source re-scan (~0.65 s) stays out of every session start and happens only when someone asks for it.

Re-running is a first-class path. /import-existing-project opens by checking what is already imported and switches to finishing and refreshing rather than repeating — fix the globs it names, ask the operator about the SHIPPED drift, re-run ddflow import for sections added since.

What it reports rather than fixes

Three kinds of drift it can see and must not resolve on its own, because either answer could be the wrong one:

  • A phase heading that says SHIPPED over unticked checkboxes (32 of them in the repository this was measured against). One-sided risk: if the heading is right, the queue is about to hand out work that is already done.

  • A dependency on an id nothing produced. Kept and treated as unmet — deliberately, so a typo surfaces as blocked work rather than as work that starts early — but named, because "never offered" otherwise looks exactly like "nobody has got to it yet".

  • A file that matched a source pattern and yielded nothing, which usually means an unusual format rather than an empty file.


The model: phases, tasks, dependencies, globs

Plan ──► Phase ──► Task

A phase is a unit of review: its own research, its own whole-suite test pass, its own live smoke run, merged as one coherent feature. A task is a unit of execution: one agent, one worktree, one pipeline, one merge. Both carry needs (dependencies, which may cross phases) and globs (the files they will write).

ddflow phase add P2 --title "Billing" --needs P1
ddflow task add P2.T1 --phase P2 --title "invoice model"  --globs "src/billing/invoice.py"
ddflow task add P2.T2 --phase P2 --title "tax rules"      --globs "src/billing/tax.py"
ddflow task add P2.T3 --phase P2 --title "checkout wiring" --needs "P2.T1,P2.T2" \
                                                            --globs "src/checkout/*"

Declare globs. They are what lets two agents work at once safely. A task with no declared globs is a task the conflict detector cannot protect.

Dependencies are inherited. A phase is never claimed — only its tasks are — so P2 needs P1 has to govern everything inside P2, or it governs nothing that anyone picks up. The readiness rule therefore consults an item's ancestors as well as itself:

$ ddflow next
Ready (1 ready, 0 running, 1 blocked):
  P1.T1  money
  (blocked) P2.T1: deps — phase P1 has 3 open task(s) (inherited from P2)

The refusal names where the dependency came from, because an operator told only "P2.T1 needs P1" goes looking for a declaration that is not written there. The one dependency not inherited is one pointing into your own subtree: an umbrella that declares a dependency on its own child would otherwise make the child wait for itself, turning a plan typo into a permanent hang.

ddflow claim asks the same predicate ddflow next does. They used to disagree — next withheld a task on its dependencies and claim handed out a worktree for it a second later — so an agent picking work by id rather than by asking bypassed the dependency graph entirely.

There is deliberately no third level of kind: a sub-task is a task whose parent is a task, so depth is unlimited while the rules stay one set.


Work that changes shape while you do it

Tasks can be added at any time, including while their parent is being worked — mid-task discovery is the normal case, not an exception, and a queue that cannot absorb it pushes the work into someone's head.

Sub-tasks are just tasks whose parent is a task. Not a separate concept with its own rules: a sub-task declares its own globs, carries its own dependencies, is claimed by its own agent, and runs in parallel with its siblings when nothing links them — exactly like any other task.

ddflow task add P1.T1a --parent P1.T1 --globs "src/parse.py"
ddflow split P1.T1 --into "P1.T1a=parse input" --into "P1.T1b=write records"

split works in place: the original keeps its id, its lease history and everything recorded against it, and becomes an umbrella that completes when its children do. Closing it and opening two new ones instead would lose the thread between what was planned and what happened — which is exactly what ddflow replay needs.

An umbrella is never offered as ready (its children are), and cannot complete while any descendant at any depth is unfinished. An abandoned child counts as settled, so a sub-task you decide against does not hold its parent open forever.

Becoming an umbrella releases the lease, however you get there — by split, or by adding the first sub-task to a task you are already working. An umbrella holding a live claim on globs that overlap every child's means a second agent cannot take one of those children, and crash recovery points at a worktree where nothing further will happen. split already did this; task add --parent did not, which is the shape of bug worth naming: one transition, two ways in, guarded on one.

Architectural decisions

The code shows what was built and never why, nor what was rejected on the way. So decisions are recorded as events, and reach the person writing the code:

ddflow decision add --title "Storage is SQLite with WAL" \
  --decision "One file, WAL mode, BEGIN IMMEDIATE for writes." \
  --context "Three call sites were each opening their own connection." \
  --alternatives "Postgres — rejected: no server allowed in this deployment." \
  --globs "src/storage/*" --by operator

--globs is what makes a decision consulted rather than merely filed. ddflow brief and ddflow decision applicable <item> surface the decisions governing an item's declared files automatically — the agent does not have to suspect they exist.

Decisions are never edited or deleted. A reversal is a new decision naming the old one (--supersedes), so the history of how the architecture got here survives, and a superseded decision is shown with a pointer to its replacement rather than silently withheld.

Recall — "have we been here before?"

ddflow recall "how should durations be represented"

One search across everything the project remembers: architectural decisions, lessons, research verdicts, past bugs, similar tasks, and the operator's own earlier prompts. Results are labelled by kind, because a binding decision, a transferable lesson and a prompt from three weeks ago should change what you do in different ways.

It exists so the operator does not have to say the same thing twice and the agent does not have to learn the same thing twice. Both failures are invisible in the moment and obvious in the log.

Status, progress, and loops

ddflow status      # what is done, in flight, ready, blocked — one answer
ddflow progress    # attempts, hours held, gate runs, commits, per item
ddflow loops       # circular references and runtime loops (exit 2 = none)

Dependency cycles are the easy case. The expensive ones are runtime loops, where the graph is perfectly acyclic and the work still never finishes:

Detector

Catches

dependency_cycle

A needs B needs C needs A — always blocking

repeat_claims

claimed and given up N times without completing (crash-expiries excluded: that is a different problem)

gate_flapping

a gate whose verdict keeps flipping — flaky, or measuring a moving target

reopened

work that will not stay done, usually because the acceptance criteria are not in the item

duplicate_work

two live items declaring the same files

no_progress

N recent events with no completion, no gate pass, no merge

Every threshold is a [loops] knob, and on_detect = "block" makes ddflow claim refuse an item that is already looping — a warning is read by a human later, a refused claim is read by the agent now.

The task pipeline

Ten gates, in order, configurable per project:

#

Gate

Run by

Purpose

1

research

agent

State a falsifiable claim; probe it before building on it

2

rules

agent

Load project rules + the lessons relevant to this task

3

implement

agent

Write the change, in its own worktree

4

rubber_duck

different-family model

Try to refute the change

5

critic

different-family critic

Where does the diff disagree with the intent?

6

standards

tooling

Linters, architecture review, coding-standards MCP

7

unit_tests

tooling

The project's suite, actually executed

8

bug_hunt

agent

Hunt the recurring classes across everything touched

9

dedupe

agent

Did this re-implement something already present?

10

merge

ddflow

Land it, from the primary checkout, with no checkout

Four things are enforced rather than requested:

Silence is not a pass. Every gate in the pipeline must carry some outcome before an item completes — passed, failed, unavailable, partial, or an explicit ddflow gate skip <id> <gate> --reason "...". Without this, gates.required held only implement, unit_tests and merge, so six of the ten steps could be omitted with no trace at all. gates.require_outcome = false makes the pipeline advisory again; gates.enforce_order ("warn" by default, or "block") reports a gate recorded before an earlier one has run, because a rubber-duck review recorded before implement reviewed an empty diff.

UNAVAILABLE is never a pass. A reviewer whose endpoint was down approved nothing; a linter that is not installed found nothing. Each gets its own outcome and shows as a coverage gap. (The inverse matters too: this codebase's first version classified a missing binary — shell exit 127 — as failed, so an uninstalled linter looked like a linter reporting problems. Fixed, with a mutation-verified regression test.)

Evidence or it did not happen. Gates in gates.evidence_required reject a bare pass; they want the command, its exit code and its output digest.

And evidence says WHICH tree and HOW MUCH. Every gate that produces an OUTCOME — command gates, and agent gates recorded with gate record — carries a tree_sha: a fingerprint of the working tree it ran against, covering committed state, uncommitted changes to tracked files, and the content of untracked ones (a new module is untracked until its first commit, which is the ordinary state of agent work). ddflow's own .ddflow/ is excluded, or recording a gate's outcome would invalidate the gate that just recorded it. If the tree moves afterwards, complete warns that the pass describes source nobody is shipping — a warning, not a block, because refusing on a comment-sized change is how a check gets switched off.

Beside it, diff_stat records files, insertions, deletions and untracked count — including the lines in untracked files, because a new module is untracked until its first commit and a task that is entirely new files would otherwise report zero insertions. The fingerprint answers which tree and is opaque; this answers how big, and that is what makes a pass auditable later — a review gate that passed over 4,000 changed lines in two minutes is a different claim from one that passed over 12.

Neither is recorded for a skip (nothing was reviewed, so a magnitude would imply an inspection that did not happen) nor for an unavailable gate that never ran.

Hashing untracked content is capped by MAX_UNTRACKED_HASHED (512). Above it the fingerprint falls back to file names and says so inside the digest, because a check that quietly stopped covering content would go silent for exactly the repositories that need it most.

Reviewer independence is checked, and an unidentified reviewer establishes nothing. Same-family reviewers share the author's blind spots, so their agreement measures shared priors rather than correctness. complete refuses unless one reviewer came from a different pretraining family — and a reviewer whose model is not in [agent].families counts as unknown, never as different. (It used to count as different: gate record defaults the reviewer to the agent id, so a standards gate recorded with no --model arrived as family "host-12345", compared unequal to "anthropic", and satisfied the independence requirement on its own.)

$ ddflow complete P1.T1 --model claude-opus-5
cannot complete P1.T1 — 1 unmet condition(s):
  - reviewer independence not satisfied: every reviewer (rubber_duck) was family
    'anthropic', the same as the author. Same-family agreement is not independent evidence.

Every unmet condition is listed at once — a refusal that reveals one problem at a time trains an agent to reach for --force.


Proving a gate can fail at all

$ ddflow gate verify T1 unit_tests
  OK   src/calc.py: detected

unit_tests CAN fail: every registered mutation was caught.

A gate that cannot go red is worse than no gate — it reports success on every change, and everyone downstream reads that as evidence. gate verify breaks what the gate guards, using the mutations registered beside it, and requires the gate to notice:

[gate.unit_tests]
command = "python -m pytest -q"
cwd = "repo"
mutations = [ { file = "src/calc.py", old = "return a + b", new = "return a - b" } ]

Four things make it honest rather than ceremonial, and the last one is this feature's own bug, found by a cross-family review of it:

  • A mutation that did not apply is a FAILURE, not a skip. If old is absent — or present twice, so the edit is ambiguous — the check fails. Skipping turns "the mutation never happened" into a green run, which reads as the opposite of the truth.

  • The source is restored whatever happens, including on exception, or a failed verification leaves the tree broken and the next gate reports the verifier's fault.

  • A gate with no registered mutations is reported as unproven. Declaring a check nobody has shown can fail is what this exists to catch. An agent gate says plainly that it has no command to mutate and rests on its evidence contract instead.

  • A green baseline is required first. A gate already red for an unrelated reason — one pre-existing failing test, a tool that stopped being installed, a flake — reports failed for every mutation, so every mutation reads as detected and the gate is certified as able to fail when nothing has shown any such thing. The check written to catch the vacuous-pass class contained it. It now runs unmutated first and refuses without a pass.


The phase pipeline

research → [ task, task, task … ] → unit_tests → bug_hunt → dedupe
         → live_test → corrections → merge

live_test is the one most often skipped and the one most worth keeping: a green unit suite and a working feature are different claims. Run the real thing on a small input and paste what it printed.


Human approval: a gate the agent cannot clear

Every other gate here is satisfied by the agent — it runs a command, or it asserts it did the thinking. That is right for work whose correctness is checkable afterwards, and wrong for a plan: by the time an agent has built the wrong thing, the cost is already paid.

A gate marked human = true is where the operator says yes, build that before the compute is spent.

# .ddflow/gates.toml
[gate.plan_approved]
title  = "Operator approves the plan"
human  = true
prompt = "Show the operator what you intend to build, then ask."
$ ddflow gate run T1 plan_approved
gate 'plan_approved' is a HUMAN-APPROVAL gate. It is not something you can run or
record — it is where the operator decides whether this work should proceed.
  ddflow approve T1 plan_approved
  ddflow approve T1 plan_approved --reject --reason '...'

$ ddflow gate record T1 plan_approved --outcome passed --evidence "looks fine"
'plan_approved' is a human-approval gate: it is cleared by a person, not by an agent
recording that it happened.                                            # exit 3

$ ddflow approve T1 plan_approved --note "read the plan, ship it"
T1.plan_approved approved by delian — read the plan, ship it

A rejection is a first-class outcome, not the absence of an approval: "the operator looked and said no" and "nobody has looked yet" are different states, and an item sitting in the second forever is how a checkpoint becomes a silent stall. --reject requires --reason.

There is deliberately no MCP tool for this, and tests/test_mcp_parity.py records the exemption with that reason. A human checkpoint reachable from the MCP surface is not a human checkpoint — it is a second gate record with a longer name. gate skip is refused too: "the operator does not need to approve this" is not the agent's call.

What this is, precisely. An audit trail and a speed bump, not a security boundary. An agent with shell access can run ddflow approve itself, and no design here changes that — the tool does not control the machine.

The guarantee, as narrowly as it holds: no MCP tool records a human outcome, and a clearance carries the OS user and a human flag, so a forged one is visible in the log rather than indistinguishable from a real one. (gate.<id>.human is also refused by the config writer, because two MCP calls — flip the flag, then record — used to clear the gate with no shell involved. Declare human gates in .ddflow/gates.toml, which no tool writes.)

Opt-in: the shipped pipeline has no human gate, and a test keeps it that way.


Parallelism and coordination

$ ddflow next --phase P1
Ready (3 ready, 0 running, 1 blocked):
  P1.T1  persistent store
      writes: shortener/store.py, tests/test_store.py
  P1.T2  base62 encoder
      writes: shortener/encode.py, tests/test_encode.py

These are independent — run them in parallel worktrees.
  (blocked) P1.T3: deps — P1.T1 is open; P1.T2 is open

ddflow claim <ID> leases the item and binds it to a worktree.

If you are already in one, it adopts that one. Agent harnesses — Claude Code, Cursor — often isolate the agent themselves. Claiming from inside a linked worktree binds the item to that tree and branch rather than building a rival and telling you to leave the one holding your uncommitted work:

$ ddflow claim T1            # run from inside the harness's own worktree
claimed T1 (lease 1800s, renew every 300s)
  worktree: /work/agent-tree  (adopted — you were already in it)
  branch:   agent-work
  Carry on where you are.

ddflow never needed to have created the tree — it needs to know which tree an item is worked in, so recover can find stranded work and merge knows what to merge. An adopted tree is recorded as adopted, not created, so remove_on_merge will never delete something ddflow did not make. A tree already bound to another open item is refused: two items in one tree cannot be merged or recovered separately. worktree.adopt_existing = false restores the old behaviour; --no-worktree skips binding entirely.

A second agent is refused, and told what to take instead:

$ ddflow claim P1.T4 --agent gamma
P1.T4 writes 'shortener/store*.py' which overlaps 'shortener/store.py' held by alpha on P1.T1

You could take instead: P1.T2, P1.T5

Exit codes are the contract, and agents branch on them:

Code

Meaning

0

healthy

1

real failure

2

could not run / nothing to do — never collapsed into 0

3

coordination refused

"Nothing is ready" and "everything is fine" are different facts. An agent that cannot tell them apart invents work.

The critical path is reported, because it, not the task count, sets the wall-clock floor — adding a fifth agent to a phase whose runtime is a four-deep chain buys nothing.


Many agents, one server: identity, state and sharing

Several agents and subagents sharing one queue is the case this tool is for. Here is exactly how that works, because each of these has a wrong answer that looks right.

Is it stateless?

The queue is. The connection is not, in exactly one respect.

The append-only event log is the sole source of truth, and every read re-derives state from it — fold(read_all()), from scratch, on every call. Nothing is cached between requests, so there is no stale projection, no invalidation, and no divergence between two agents' views. Restart the server mid-task and nothing is lost: it never had anything the log did not.

The one piece of per-connection state is who you are (below). It is deliberately not in the log, because it is a property of the caller, not of the work.

Who is calling?

By default, identity is derived from the working tree. That is right for one agent per worktree, and silently wrong for several agents in one tree — they all resolve the same path to the same name, their events merge into one stream, brief answers with a sibling's task, and reviewer-independence compares an agent with itself and passes. Nothing errors. There is no signal that can tell them apart, so identity is declared:

How

When

ddflow_identify (MCP)

An agent or subagent announcing itself on its connection. Call it first.

DDFLOW_AGENT env var

A harness that spawns agents and knows their names. Process-wide.

--agent (CLI)

Scripts and one-off commands.

tree-derived default

One agent per worktree. Reported as undeclared, so you can see it.

Innermost wins. ddflow_identify is idempotent, persists for the connection, and refuses a name that could not be a log filename — it becomes one, and refusing at declaration time means the caller reads the reason rather than discovering it at the first write.

If more than one agent works one tree at once, declare identity. Everything that attributes work depends on it.

Can one server serve several projects?

No — one server process serves one repository, fixed at start from --repo, DDFLOW_REPO, or the working directory. No tool takes a repo argument, and a test asserts none ever does. Point a second agent at a second project by running a second server; they are cheap, and the isolation is the point.

One project shared by many agents is the supported case — and the one that needs no special setup beyond declaring identity:

  • Writes never conflict, and they never block readers. Each agent appends to its own log shard, so there is no shared file to overwrite and no merge conflict to resolve. Writers do serialise briefly: one repo-wide lock is held across the clock allocation, the append and its fsync. Short, but not nothing — per-agent shards remove file contention, not lock contention.

  • Reads take no lock at all, so a read-heavy agent cannot be starved by a write-heavy one, and a reader can never block a writer.

  • File ownership is coordinated by globs. claim refuses an item whose writes overlap one already held, and names what to take instead — exit 3, not a failure.

  • Lessons, decisions, research and bug history are shared by construction: they are events in the same log, so one agent's finding is immediately visible to every other.

Locking, contention and measured cost

Measured on this machine, single process, full read_all() + fold():

Events

read + fold

per event

500

11.5 ms

22.9 µs

2,000

32.6 ms

16.3 µs

5,000

46.0 ms

9.2 µs

10,000

89.2 ms

8.9 µs

20,000

174.7 ms

8.7 µs

Linear, converging on ~8.7 µs/event; the higher figure at small sizes is fixed per-call overhead, not the fold. A project with 20,000 events pays ~175 ms for a state-reading call. Search and recall do not pay this — they run off a SQLite projection rebuilt only when the log's head moves.

tests/test_mcp_load.py runs 12 concurrent agents through the real MCP surface and asserts no deadlock, no lost append, no repeated Lamport value within an agent, and correct attribution for every event — not for a sample. Its thresholds are environment variables (DDFLOW_LOAD_AGENTS, DDFLOW_WRITE_LATENCY_BUDGET_S, DDFLOW_GROWTH_TOLERANCE, …) because a load test with a hardcoded budget either flakes on a shared runner or is too loose to fail.

The deadlock bound is a hard timeout: a wedged lock does not fail, it hangs, and an unbounded hang reads as a broken CI runner rather than as a bug.


Crash recovery

An agent is killed. Nothing is cleaned up, because in a real crash nothing runs.

$ ddflow recover
1 recoverable situation(s); 1 may contain work:

!! P1.T1  [expired_lease]  was: delta
     worktree /repo/../.ddflow-worktrees/P1.T1
     INSPECT FIRST — 1 uncommitted file(s), 1 unmerged commit(s).
     `git -C .../P1.T1 diff main` then salvage,
     then `ddflow release P1.T1 --note salvaged`.

Four behaviours, each chosen against a specific way this goes wrong:

  • While the lease is live, nothing happens. A dead agent is indistinguishable from a slow one until the lease expires, and guessing is how two agents end up in one tree.

  • Recovery measures the tree — uncommitted files, unmerged commits — rather than trusting the recorded state. "Is there work in here?" is the only question that decides the remedy.

  • An expired lease is never stolen silently, and recover --apply expires only trees it measured as empty.

  • Adoption, not duplication: an agent resuming a recovered item gets the existing worktree back, not a second one beside it.


Reconstruction from logs alone

$ ddflow replay --out ./recovery-kit
wrote:
  recovery-kit/RECONSTRUCTION.md
  recovery-kit/QUEUE.md
  recovery-kit/LESSONS.md

RECONSTRUCTION.md is written as instructions to a fresh agent, not as a report about the past: every operator prompt in order, every research verdict, every lesson, the queue's shape — and every approach already tried and rejected, with the measurement that killed it.

It states its own limit, in the document: it reproduces the decisions, not the bytes. Model outputs are not deterministic, so replaying prompts will not recreate the original source. What it recreates is every input that produced it, which no other artefact holds.

The demo destroys an entire repository and rebuilds from 3.9 KB of JSONL, then checks nine specific fragments are present — including the operator's stated reason for a constraint, and the probe output behind a rejected design.

Secrets are redacted on the way in, not on the way out — the log is committed, so a scrub at read time is a scrub that git show walks straight past.


Lessons that check themselves

A lesson can name the mistake in code, not just in prose:

$ ddflow lesson add --title "Never swallow a bare OSError" \
      --rule "Catch the specific error; a broad except turns a loud failure into a silent one" \
      --pattern "except OSError" --globs "*.py"
lesson La0c745cc recorded — inventory: 2 site(s) now

$ ddflow lesson verify          # later, after somebody adds a third
La0c745cc: 1 NEW site(s): c.py: except OSError:

Exit 1 names the file. That is the whole design, and it comes from a failure worth repeating: on the project ddflow was extracted from, a count-based clone ratchet sat red for ~350 commits. It was advisory so it never blocked, it reported a number so every reader learned to skip it, and 24 new clones arrived through that gap.

A count says "worse" and never "which".

A number cannot be acted on or reviewed. A list can: a new entry is a line somebody opens, and a disappeared entry is progress — reported, and never a failure, because the inventory may only shrink.

Three details that decide whether a ratchet survives contact with a real repository:

  • A site is <path>: <matched text>, not path:line. Line numbers churn on every edit above a site, which would invent a matching pair of "new site" and "fixed site" findings out of an unrelated change — and a ratchet that cries wolf is one that gets switched off.

  • An uncompilable pattern is refused, not stored. An empty inventory reads exactly like a clean repository, and would ratchet every real occurrence away the first time it ran.

  • Exit 2 when no lesson declares a pattern. Not a pass. A corpus with zero ratchets should not be able to report "all clear".

Vendored and untracked files are never sites — matches in code nobody owns are findings nobody will act on. Most lessons stay prose, and a prose lesson produces no findings at all.

Checking that the checks are working

Three questions ddflow asks about itself, all derived from the log and all reported by ddflow doctor. They exist because an unmeasured mechanism is indistinguishable from a missing one.

Can the queue's work actually be picked up? Every other check counts the items that are present; this one asks whether any of them can be started. An open phase with no task under it is work ddflow next will never offer — a note by default (schedule.empty_phase, since a project that files phases before breaking them down lives there on purpose) — and a phase whose tasks are all finished while the phase stays open is always a problem, because that is a queue held open by an item nobody can act on.

Tasks cannot go missing here, and that is a property rather than an untested gap: two hypotheses about how one could were probed and both refuted, and a test now pins the invariant so a future filter cannot quietly reintroduce it.

Does a gate ever say yes? A gate that fails on everything is worse than no gate: it trains the next reader to skip it. A gate at or above gates.rate_max_fail once it has gates.rate_min_runs decisive runs is reported as flaky or as measuring a moving target — re-running it will not converge. A skipped gate is not a run, because counting skips as failures would make an unconfigured gate look like a broken one.

Did the periodic passes ever fire? The mechanism you did not measure is the one that is not running. Because a cadence here counts completions rather than wall-clock, this is exact rather than estimated: since is the completions elapsed since the pass last fired, which is the same quantity ddflow cadence uses to decide due-ness — deliberately, because two measures of "is this behind" that can disagree is a situation nobody can reason about. A pass more than cadence.max_missed scheduled runs behind is reported. Being merely due is not a finding (ddflow cadence already says that), and running early is not one either.

All three are notes, not problems: a defect in the machinery that checks the work must not block the work.

Reading the log, and why it is never compacted

The log only grows, so every state-reading call used to re-read and re-parse all of it. Measured at 20,000 events, that read costs 115 ms — and the breakdown is the whole design argument:

stage

cost

share

Event.from_json

97 ms

84%

fold into state

9 ms

7%

sort by Lamport key

3.8 ms

3%

read the bytes off disk

3.7 ms

3%

de-duplicate by content address

0.7 ms

<1%

Parsing dominates, and an append-only file guarantees the bytes already parsed have not changed. So EventLog.read_all re-parses only the appended tail, and re-hashes the bytes it is re-using to prove they are still the same bytes:

events

read, uncached

warm read

20,000

121.1 ms

11.9 ms

10.2×

100,000

623.8 ms

62.8 ms

9.9×

A command like ddflow doctor — which reads four times — pays the full cost once instead of four times.

Two knobs, [log]:

knob

default

what it trades

reuse_parsed

true

Off = always re-parse from scratch. Slower, and worth it only if a shard is being rewritten in place under a running process.

max_cached_events

100000

Memory ceiling, in events. ~736 bytes per parsed event, so the default holds ~74 MB in a long-lived MCP server. Over the ceiling the cache is dropped and reads cost what they always did.

The validity check is a content check, and that is the whole design. The consumed prefix is re-hashed on every read — 1.1 ms to read plus 4.3 ms to digest, against the 97 ms of parsing it avoids. The first version used st_ino instead, on the reasoning that "a git merge writes a temp file and renames, so the inode changes". That is false:

$ git checkout -q other && stat -c %i .ddflow/events/a1.jsonl
218500670
$ git checkout -q main  && stat -c %i .ddflow/events/a1.jsonl
218500670

Git rewrites tracked files in place. So switching between two branches that had diverged left a warm server serving events from the branch you left, silently losing the ones actually on disk, with the tail read starting mid-line — and because Store.rebuild takes its fingerprint from the real file while taking its events from the cache, that wrong state was written into the SQLite index stamped as current, which a fresh process would not rebuild away. A content digest makes a rewrite, a truncation, a git checkout, a git merge, a delete-and-recreate and a torn tail all one case, so there is no list of mechanisms to keep current.

A torn final line from an append that died mid-write is reported by ddflow doctor and re-read until the writer completes it, never marked consumed. Every guarantee here is mutation-verified in tests/test_log_read_cache.py — including that the digest covers the whole prefix rather than a trailing window of it, which a smaller fixture cannot tell apart.

The compaction that was declined

An event kind log.compacted was reserved for a retention pass that shrank the log. It has been removed, because the recipe it was reserved for cannot be implemented without breaking two shipped commands. Three probes:

  1. A compaction survives a merge=union merge. One branch compacts, the other appends; the deletions stick. So union is not the obstacle.

  2. Two divergent compactions merge to neither side's result, and out of Lamport order — the case union cannot resolve.

  3. The decisive one. ddflow progress and ddflow loops read raw events, not folded state: progress.work pairs each lease.acquired with the next release across the whole history. Leases and gate outcomes are not PROVENANCE_KINDS, so keeping "the last state-bearing event per subject" leaves a lease.released with no acquire to pair with. A queue whose loop detector fires repeat_claims before compaction reports nothing after it, and six attempts become zero — and under [loops] on_detect = "block" that is a behaviour change, not just a lost report.

Growth is addressed by making the read cheap rather than the log short, which keeps it append-only and auditable. Beyond ~100k events the right answer is an on-disk state snapshot, not a shorter history.


Lessons, research and bugs

ddflow lesson add --title "Truncating a slug can leave a trailing separator" \
                   --rule "Strip separators AFTER slicing to length, not before."
ddflow lesson search "cutting a url short leaves a dangling hyphen"

Retrieval is BM25 over FTS5 and finds that entry despite no shared keyword. Probed against embeddings and found sufficient at lesson-corpus scale (R4); lessons.search_backend exists for when that stops being true.

Research entries must carry a verdict, and CONFIRMED/REFUTED are refused without a probe:

$ ddflow research --question "is it fast?" --verdict CONFIRMED
CONFIRMED requires a --probe (and ideally --probe-output): a verdict with no probe behind
it is an opinion. Use THEORETICAL and say why no probe was possible.

And a bug cannot be closed without the test that would catch it again:

$ ddflow bug fixed B1
a bug may not be closed without --regression-test naming the test that would catch it
again. Write the test, watch it FAIL against the unfixed code, then close.

That refusal is the whole mechanism by which the same bug does not ship twice.


Cadences

Periodic whole-repo passes a per-task gate structurally cannot do. Due-ness is derived from completed work, so there is no state file to drift:

$ ddflow cadence
DUE: integration_tests — 5 tasks since last (every 5)
DUE: mutation_tests — 3 phases since last (every 3)

Record one with: ddflow cadence --ran <name>

Configurable: integration tests, architecture review, mutation testing, duplication sweep, lessons compression.


Keeping AGENTS.md true

ddflow adopt writes a managed block into AGENTS.md (and CLAUDE.md, and each agent's native rules file). That block is what tells an agent it must claim an item before editing — and every coordination guarantee here rests on that, because an agent that does not claim has its work destroyed by a parallel one.

Nothing used to check it again. Adoption is judged by .ddflow/config.toml existing, so a deleted AGENTS.md, a block someone stripped, or a block written by an older ddflow all left the agent reading rules that were absent or wrong while every surface reported the project as adopted. Adoption is a config file; the instructions are a separate fact.

Cursor does not really follow AGENTS.md, and it is not alone. Its precedence is Team Rules > Project Rules > User Rules > .cursorrules > AGENTS.md, so .cursor/rules/ddflow.mdc is what actually binds — which is why adopt writes it. That file is checked too, for every agent the project was adopted for (read from the driver deltas on disk, so a Claude-only project is never asked for a Cursor rule).

It carries the same block with binding frontmatter, and alwaysApply: true is part of what is verified: a rule with alwaysApply: false exists, reads perfectly, and may never be loaded — which for claim-before-you-edit is the same as not having it, and strictly worse than drifted text. It is reported at the severity of missing, not of stale.

Five states are detected — current, stale (drifted from what this version writes), no_block (file there, block gone), not_binding (native rule that will not apply), missing — and reported on three surfaces:

Surface

What it does

ddflow doctor

missing and not_binding are PROBLEMS (exit 1) — the agent has no rules, or has them and will not load them. stale is a note, so an upgrade does not turn the health check red.

The MCP handshake

A block naming the file, what is wrong, and ask the operator first.

The footer on tool results

Reports it mid-session, because the handshake fires once.

The two surfaces repair it differently, on purpose.

  • From a shell, the operator is right there: ddflow adopt rewrites the block. It replaces only what is between the DDFLOW:BEGIN/DDFLOW:END markers and leaves the rest of your file alone, and re-running it is a no-op. ddflow init reports the problem and does not write — writing prose into your AGENTS.md is not what init was asked to do.

  • Over MCP, ddflow does not touch it. The handshake tells the agent to show the operator what is wrong and call ddflow_setup only if they agree. It is a file in their repository, usually with their own prose around the block, and rewriting it is not a decision a tool gets to make on their behalf — the same rule as companions ("propose; never install") and the human-approval gate.


Surviving a compaction

The instruction block reaches the model once, at connect. After a context compaction it may retain none of it, and MCP has no server-to-client primitive for injecting context — the three that exist (roots/list, sampling/createMessage, elicitation/create) all go the other way or ask a question. Three things already survive:

  • the AGENTS.md / CLAUDE.md sections ddflow setup writes, plus each agent's native rules file — the client re-reads its own rules, so this is the durable channel;

  • the commit hook, which refuses a commit with no item trailer and says what to add. Enforcement at the moment of the act needs no context at all;

  • ddflow help <topic>, which the agent can ask for — if it thinks to.

What none of those do is speak up unprompted. A footer on tool results is the only channel that is guaranteed to be heard again, because an agent driving ddflow calls tools continuously:

ddflow: left undone in this project —
  · 1 bug(s) still open: B1 — close with `ddflow_bug_fixed` (it requires the regression test) or say why not
  · 2 gate(s) skipped, not run: T4.critic, T4.standards — run them, or leave the skip on the record deliberately

It is not a banner, and the difference is the whole design. A fixed reminder appended to 63 tools is trained out inside a session and costs tokens on every call. This one:

  • names what happened, never restates a rule — an id, a count, and the call that discharges it;

  • stops once the thing is dealt with, so it cannot be trained out by repetition;

  • says nothing at all when the project has nothing outstanding — not a cheerful "all clear", which is the same thing readers learn to skip;

  • is cadenced: at most once every every_calls calls and every_seconds seconds, so a burst of calls is not a burst of footers;

  • cannot break the call it rides on. It is a courtesy on top of an answer, appended after the body, and a failure inside it is swallowed. content[0] is still the structured result.

[reinstruct]
enabled      = true   # false silences it entirely
every_calls  = 12
every_seconds = 240
max_items    = 3

What it currently notices: bugs found and never closed, gates skipped and never revisited, and work finishing with no lesson ever recorded (after [lessons] reflect_after_items, so one task is not reported — the pattern is, and the threshold is a knob because where the line sits is a judgement).


Keeping session-start cost flat

ddflow brief --phase P2

Returns, inside session.brief_max_tokens (default 1200): recoverable work first, then the current item and its remaining gates, then what is ready, then why everything else is blocked, then the handful of past lessons ranked against this task's text.

This replaces reading the project's rule and lesson corpora. The budget is enforced by truncating from the bottom, so the safety-critical head survives a squeeze — and a project's opening cost stays roughly constant as its lesson corpus grows.


Agent portability

One canonical driver, templates/drivers/implement-phase.md, plus a delta per agent covering only what genuinely differs: how iteration continues, how to ask the operator, how to spawn a subagent, file-reference syntax.

Deltas rather than copies, for a measured reason: on the project this was extracted from, a reworded per-agent duplicate of the driver silently accumulated three instructions that were false at the time of writing while missing four gates the canonical file had gained. A delta removes the surface that can drift instead of policing it.

Both surfaces are one implementation — the MCP server maps each tool onto the same cli.main() call in-process, and two tests plus a demo step assert they cannot diverge.

Agent

Reads

MCP config written by adopt

Claude Code

CLAUDE.md → driver

.mcp.json

Gemini CLI

AGENTS.md

.gemini/settings.json

Codex CLI

AGENTS.md

.codex/config.toml

GitHub Copilot

.github/copilot-instructions.md, AGENTS.md

.vscode/mcp.json

Kilo / Cline

AGENTS.md

.kilo/kilo.json

CI / Make / human

—

none; the CLI is complete on its own


Keeping the two surfaces honest

Every CLI command is reachable over MCP — that is the point of the tool list, and it is the requirement that an operator in a chat window, possibly driving a remote agent, can do everything a shell can. Three ratchets keep it true, and each one was added after the previous one turned out to be too shallow:

Ratchet

What it caught on its first run

every CLI command has a tool

the original check

every CLI subcommand has a tool

ddflow gate skip and bug found had none — gate counted as "covered" by gate run, and a parent's coverage says nothing about its children

every CLI flag is reachable from its tool

27 divergences — 16 on its first run, and 11 more the moment it derived its own coverage instead of using a hand-written list. Including phase add --globs: over MCP a phase could not declare what it writes, so the conflict detector had nothing to compare at phase level

The flag ratchet derives its own input from the parser rather than a hand-written list — its first version carried eleven tools and was blind to remove --force for exactly that reason. Omissions are allowed, but each must be an entry in FLAG_EXEMPTIONS with its reason, so "we chose not to expose this" and "nobody noticed" stop looking alike.

A fourth pins something subtler: whether a tool returns JSON or prose is a decision, not an accident. Some tools deliberately return prose — brief, gate status and replay exist to hand the model an instruction or a narrative, and JSON-encoding a paragraph so the client can decode it again helps nobody. But decision add returned JSON while task add returned prose for no reason either could state. Each prose tool now carries its justification in PROSE_TOOLS.


Command reference

ddflow adopt [--agents ...]     install into a project, for one or more agents
ddflow init                     create .ddflow/ only

ddflow phase add <id> [...]     add a phase
ddflow task add <id> --phase .. add a task
ddflow update <id> [...]        change title/body/needs/globs/tags/priority

ddflow next [--phase P]         what may start now       (2 = nothing actionable)
ddflow claim <id> [--globs ..]  lease + create worktree  (3 = refused)
ddflow heartbeat <id>           renew a lease
ddflow release <id>             give it up

ddflow gate status <id>         pipeline position + the next gate's instruction
ddflow gate run <id> <gate>     execute a command gate, record its evidence
ddflow gate record <id> <gate>  record an agent gate    (--outcome, --reason, --model)
ddflow gate skip <id> <gate>    skip, with a mandatory reason
ddflow approve <id> <gate>      a PERSON clears a human gate  (no MCP equivalent)
ddflow approve .. --reject      ...or refuses it, with --reason
ddflow gate verify <id> <gate>  prove the gate CAN fail  (1 = it cannot)

ddflow merge <id>               merge from the primary checkout, no checkout
ddflow complete <id>            finish        (3 = unmet conditions, all listed)
ddflow block <id> --reason ..   mark blocked

ddflow brief [--item|--phase]   budgeted session-start pack
ddflow board / show <id>        human views
ddflow render                   regenerate docs/ddflow/*.md

ddflow help [topic]             what this is, what it can do, the workflow
ddflow workflow                 the rules this project runs by  (1 = incoherent)
ddflow workflow pipeline ...    set the gates a task or phase passes
ddflow workflow gate ...        define or change one gate
ddflow workflow drop <id>       take a gate out of the pipelines
ddflow import [--apply]         propose an existing project's work  (2 = nothing)
ddflow import --verify          is the import still true, and did anyone finish it?
ddflow history [--item|--kind]  one timeline of everything that happened (2 = nothing)

ddflow lesson add|search        capture and retrieve lessons
ddflow research --verdict ..    record a finding (probe required for CONFIRMED/REFUTED)
ddflow bug found|fixed          regression test required to close

ddflow session start|prompt|note|end     provenance logging
ddflow replay [--out DIR] [--verify]     reconstruct from the log

ddflow recover [--apply]        find crashed agents' work   (2 = nothing)
ddflow doctor                   integrity + health
ddflow rebuild                  re-derive the index
ddflow cadence [--ran NAME]     which periodic passes are due  (2 = none)
ddflow config --explain         every knob, its value, its source and its docs
ddflow config --append-toml ..  add config without a shell editor (validated first)
ddflow reviewers detect|list|test   find and check cross-family review endpoints
ddflow review <id> --gate ..    run the configured reviewer, record the evidence
ddflow mcp                      run the MCP stdio server

Every one of these is reachable over MCP, and a test enforces it. One tool goes the other way and has no CLI equivalent, because it has nothing to mean there:

ddflow_identify(agent=...)      declare who you are ON THIS CONNECTION (MCP only)

A CLI invocation is one process that exits, so it says who it is with --agent and the question does not outlive the command. An MCP connection is a session, so identity is declared once and persists — see Who is calling?.


Configuration

75 knobs across 15 sections, every one documented in place:

$ ddflow config --explain --filter lease
lease.ttl_s = 1800   [default]
    Seconds a lease stays valid without a heartbeat. After this it is EXPIRED and
    reclaimable. Longer = fewer false expiries when an agent is deep in a slow gate;
    shorter = faster recovery after a crash.

Resolution: dataclass defaults → .ddflow/config.toml → DDFLOW_<SECTION>_<KNOB> env. An unknown knob is an error, never a silent drop. A test asserts every knob carries documentation, so the reference cannot rot.


What is automated, and what is not

The honest split, because a tool that claims to automate judgement is lying about the part that matters.

Automated — happens without anyone remembering it:

  • The handshake briefs the agent. On connect, the MCP server injects the live state: the pipeline every task must pass, work recoverable after a crash, what is ready, which companions are missing, and what to do about each. It is a template (ddflow prompts eject mcp_instructions), so the workflow is text you edit, not code you fork.

  • Gates are enforced, not suggested. ddflow complete refuses on a required gate that has not passed, on open sub-tasks, on a silent gate under require_outcome, on a requirement no pipeline runs, and on a reviewer from the author's own family. Refusals list every unmet condition, not the first — an agent that cannot see how many more are coming reaches for --force.

  • Conflicts are refused at claim time, by glob overlap, with an alternative named.

  • Dependencies gate readiness. ddflow next withholds a task whose needs are open and says which.

  • Gate evidence records which tree and how much — a working-tree fingerprint plus files/lines changed — so a pass names what it passed on. If the tree moves afterwards, complete warns that the evidence describes source nobody is shipping.

  • Crash recovery: ddflow recover finds worktrees whose lease expired, so an interrupted agent's work is found rather than lost.

  • Cadences (ddflow cadence) tell you which periodic passes are due — bug hunts, dedupe, lesson compression — from the log rather than a calendar.

  • A commit hook (ddflow hooks install) can refuse an unclaimed edit outright, and refuses a staged ddflow render view that the log no longer regenerates byte-for-byte — hand-edited, or stale ([enforce] generated_views).

Not automated, on purpose:

  • Installing anything. Detection is read-only; the report is advice.

  • The judgement inside an agent gate. ddflow records that you claim to have hunted bugs; it cannot check that you did. What it can do — and does — is make silence visible: a gate never run and never skipped blocks completion, so the failure mode is a refusal rather than a quiet omission.

  • Deciding whether a plan is right. That is what a human = true gate is for.

  • Pushing, releasing, or anything outward-facing.

The design assumption is that an agent's honesty cannot be verified, so the system is built to make an unverifiable claim expensive to make and easy to see: evidence contracts, mutation-verified gates, coverage gaps recorded on completion, and an exit code that distinguishes "could not" from "did not need to".


Testing

python3 -m pytest tests/ -q          # 423 unit/integration tests
python3 demos/run_all.py             # 6 end-to-end scenarios, 219 assertions

The demos invent whole projects and drive them for real — real git worktrees, real pytest and npm test runs, real merges, real concurrent processes:

Scenario

What it proves

parallel-phase

Two agents build a URL shortener in parallel; a third is refused on a file conflict; a dependent task unblocks automatically when its last dependency lands

crash-recovery

An agent is killed holding uncommitted work; it is found, measured, never stolen, and adopted intact on resume

reconstruct-from-log

The entire repository is deleted; everything rebuilds from 3.9 KB of JSONL, and nine specific facts are checked present

mcp-polyglot

A Node.js project driven end-to-end over real MCP JSON-RPC, with both surfaces asserted to agree

mcp-orchestration

A whole two-phase Python library built by two agents entirely over MCP — bootstrap, configure, discover a reviewer, fan out, get refused by the hook, real pytest, a real cross-family review, merge, close both phases, reconstruct. 24 steps, 57 assertions.

full-lifecycle

26 steps, 89 assertions — the whole arc, from an operator's first sentence to a rebuild from the log. A double-entry bookkeeping library across two dependent phases with sub-tasks: the operator states requirements in English, a decision is recorded and scoped to the files it governs, two agents fan out, a task turns out to be two concerns and grows sub-tasks, a bug is found and may not be closed without its regression test, a phase closes on its own pipeline, a new requirement arrives while a task is in flight, that task is split in place, the guardrails are tested by trying to break them, and finally every .py file is deleted and the project is reconstructed from the log alone.

The scenarios and the stress test have found most of the bugs this project fixed; the unit tests found few of them. The full-lifecycle scenario was written to exercise the requirements rather than the code, and found four defects before it passed once — every one of them a CLI/MCP divergence that command-level parity could not see. The composed MCP run alone found eight that 213 unit tests and four other scenarios missed — including two that made core features useless out of the box. They all lived in seams: between two processes, between a read and a write, between two output surfaces, between a declared vocabulary and its callers, including one that does not reproduce below ~6 concurrent processes. They are catalogued with their regression tests in R6.


Documentation index

Document

Contents

docs/ARCHITECTURE.md

The event-log inversion, ordering, concurrency, module map, what is deliberately absent

docs/RESEARCH.md

Twelve research questions with probes, measured output and verdicts; the self-found bug catalogue, the 2026-09-24 review pass (R10), the importer against a real 400-day corpus (R11), and what the MCP spec is worth for a mutating tool (R12)

docs/RECOVERY.md

Operator runbook: crashes, corruption, divergence, full reconstruction

templates/drivers/implement-phase.md

The canonical agent-agnostic driver

templates/drivers/deltas/

Per-agent deltas: Claude, Gemini, Codex, Copilot, Kilo, Cursor

probes/

Runnable probes behind the research verdicts

Available Tools

63 tools
ddflow_abandonA

Stop work on an item without completing it, with a reason. Use when a task turns out to be unnecessary or impossible. DIFFERENT from blocking: a blocked item is waiting and will resume; an abandoned one will not, and so it stops holding its phase open — which an unfinished task otherwise does forever, since nothing can ever finish it.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesItem id.
forceNoAbandon although a sub-task is still open. Those sub-tasks do NOT become abandoned with it — decide about each, or they sit in the queue under a parent nobody will finish.
reasonYesWhy it is being dropped.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the disclosure burden. It reveals the key consequence that an abandoned item will not resume and stops holding its phase open, which prevents an unfinished task from blocking a phase forever. It does not discuss permissions or reversibility, but the central behavioral difference is clearly disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is about three sentences and front-loads the purpose before the usage condition and blocking distinction. Every sentence contributes useful information, and the blocking comparison earns its place because it prevents a likely invocation error.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a three-parameter state-change tool with no output schema, the description plus schema explains when to call, what the phase implication is, and how force affects sub-tasks. A minor gap is that it never explicitly says the item remains in the system rather than being deleted, but this is mostly inferable from the 'phase open' phrasing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The schema already documents id, force, and reason with meaningful descriptions; the tool description itself adds no parameter-specific semantics beyond what the schema provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action ('Stop work') on a specific resource ('an item') and clarifies that it is not completion. The explicit comparison to blocking distinguishes it from a key sibling tool, ddflow_block, without requiring the agent to open schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides an explicit trigger condition: 'Use when a task turns out to be unnecessary or impossible.' It also gives a clear when-not-to-use rule by contrasting with blocking: blocked items wait and resume, abandoned items do not. This is actionable routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_blockA

Mark an item blocked on something outside the queue — a missing decision, an upstream outage, a question for the operator. Better than silently leaving it claimed: a blocked item states its reason, while a claimed one that nobody is working just looks busy until the lease expires.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesItem id.
reasonYesWhat it is waiting on.

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden and it does clarify the state transition and rationale ('a blocked item states its reason'). Still, it does not say whether the claim is released, whether blocking is reversible, how the lease is affected, or what the tool returns.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the core behavior and followed by a crisp rationale. No filler or repetition of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The definition is adequate for a simple two-parameter tool, and it explains the purpose and reason format well. However, it omits the practical consequences of blocking (e.g., relation to claim/lease, reversibility, response), which the description itself draws attention to by mentioning lease expiration.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3; the description adds value by giving concrete examples of what 'reason' should contain ('a missing decision, an upstream outage, a question for the operator'). This helps the agent populate the required 'reason' parameter accurately.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Mark an item blocked on something outside the queue'. It also distinguishes this from the 'claimed' state, which is a key sibling alternative, so an agent can tell what this tool does and how it differs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear decision context: use when blocked on something external like 'a missing decision, an upstream outage, a question for the operator'. It also implies when not to use it: rather than silently leaving an item claimed. However, it does not explicitly name a sibling tool or list additional exclusion conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_boardB

The whole work queue as a readable board, with the critical path.

ParametersJSON Schema
NameRequiredDescriptionDefault
phaseNoRestrict to one phase.

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of disclosing behavioral traits. The phrase 'readable board' hints at a read-only presentation, but the description never explicitly states that the tool does not modify the work queue, what it returns, or whether it performs any computation beyond presenting existing data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, efficient sentence with no filler or repetition. It front-loads the core concept ('whole work queue') and adds the distinguishing critical-path detail immediately, making it easy to parse quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema and no annotations, the description is too thin to be fully actionable. It explains what the board shows at a high level but does not clarify the output format, confirm that the operation is side-effect-free, or provide enough context to distinguish this tool from the many related workflow-visualization siblings.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already documents the single 'phase' parameter with a clear description ('Restrict to one phase'), and schema description coverage is 100%. The tool description adds no parameter-level meaning beyond what the schema provides, so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies the resource as the 'whole work queue' and the deliverable as a 'readable board' with the critical path, which is specific enough to convey the tool's purpose. However, it lacks an explicit action verb like 'display' or 'show' and does not name or distinguish itself from visually or status-oriented siblings such as ddflow_status, ddflow_workflow, or ddflow_show.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied by the description: use this when you want the entire work queue presented as a board with the critical path. It provides no explicit when-to-use guidance, no exclusions, and no comparison to alternatives, which is a noticeable gap given the large sibling tool list.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_briefA

START HERE every session. Returns a budgeted pack: work recoverable after a crash, the current item, what is ready to start now, why everything else is blocked, and the past lessons ranked as relevant to this task. Use this INSTEAD of reading the project's lesson or rule files — it is the same information retrieved for the task at hand, at a fraction of the tokens.

ParametersJSON Schema
NameRequiredDescriptionDefault
itemNoFocus on this phase or task id (optional).
phaseNoRestrict the ready set to this phase (optional).
check_recoveryNoAlso scan for crashed agents' worktrees and lead with them. Worth it at session start: unclaimed work left by a dead process is the one thing to know BEFORE picking up something new.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It clearly describes the returned content and notes token efficiency, but it does not explicitly state whether the tool is read-only or whether it has any side effects such as claiming work or updating state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no wasted words. It front-loads the most important usage instruction ('START HERE every session'), lists the output contents compactly, and closes with a concrete advantage (token savings).

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-style informational tool with no output schema and no annotations, the description explains what the result contains, when to use it, and why it is preferable to reading raw project files. It is slightly jargon-heavy ('budgeted pack'), but the following list disambiguates it well.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters thoroughly. The description adds no parameter-specific detail, but it does not need to because the schema covers item, phase, and check_recovery.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Returns a budgeted pack') and enumerates exactly what the pack contains: recoverable work, current item, ready-to-start work, blocked items, and relevant lessons. It clearly establishes the tool's role as the session entry point, which distinguishes it from the many sibling ddflow tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The instruction 'START HERE every session' gives explicit when-to-use guidance, and 'Use this INSTEAD of reading the project's lesson or rule files' names the alternative it replaces. It does not explicitly contrast with sibling tools like ddflow_status or ddflow_session_start, which is a minor gap.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_bug_fixedA

Close a bug. Requires the name of the regression test that would catch it again — write the test, watch it FAIL against the unfixed code, then close.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesBug id.
lessonNoId of an EXISTING lesson this bug belongs to.
lesson_ruleNoThe lesson in full — the transferable rule, not the incident. A future agent on a different task has to be able to apply it.
lesson_titleNoCapture a lesson at the same time.
regression_testYesTest that now guards this.

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry behavioral disclosure. It does reveal the unusual regression-test requirement and the fail-first workflow. However, it does not state side effects, such as whether the bug becomes permanently closed, whether lesson fields trigger lesson creation, or whether the operation is reversible.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

One compact sentence with the core action front-loaded and the critical requirement stated immediately. The em-dash workflow adds necessary context without wasting words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is sufficient for a basic close with id and regression_test, but it ignores the optional lesson-capture parameters entirely and provides no output or side-effect information. For a mutation tool in a large sibling family, an agent could still call it correctly, but would not understand when to attach lesson data.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all five parameters. The description adds some meaning around regression_test ('catch it again') but does not add value for id, lesson, lesson_rule, or lesson_title. This matches the baseline for high schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear action ('Close a bug') on a specific resource, distinguishing it from the sibling ddflow_bug_found at least implicitly. It does not explicitly name the sibling or contrast behaviors, so it stops short of full differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives a clear condition for use: closing requires a regression test, and it explains the expected workflow (write test, watch it fail, then close). It does not explicitly say when not to use it or name alternatives, but the prerequisite is concrete and actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_bug_foundA

Report a bug the moment you find it, BEFORE fixing it. Recording it first is what makes the fix accountable: ddflow_bug_fixed refuses to close one without naming the regression test, so a bug that was never opened is a fix that never had to prove itself. Bug hunts that record nothing look identical to bug hunts that found nothing.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoStable id, e.g. 'B1'. You will cite it when closing.
itemNoThe task it was found in or affects.
summaryYesWhat is wrong, in one line.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the disclosure burden. It does convey workflow behavior: reporting first is necessary for a later fix to be accepted, and unrecorded bugs cannot be proven found. However, it does not describe what actually happens when invoked, such as whether a record is created, whether an ID is auto-generated, or what state changes occur. Moderate transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The core instruction is front-loaded in the first sentence, and the follow-up sentences provide useful context about why reporting first matters. The final analogy is motivational rather than strictly operational, but it reinforces the usage rule without excessive length. Well-structured overall, with minimal waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple three-parameter bug-reporting tool with no output schema, the description covers purpose and workflow, and the schema covers parameters. However, with no annotations and no output schema, it leaves some practical context unstated, such as what happens after reporting and whether the agent must supply or track the id. Adequate but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the parameters are already documented: id, item, and summary. The description reinforces that summary is the one-line statement of what is wrong, but it does not add new field-level meaning beyond the schema. This meets the baseline for fully covered parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific action ('Report a bug') and names the exact resource (a bug record). It explicitly distinguishes itself from ddflow_bug_fixed by saying this tool is for reporting before fixing, while the sibling refuses to close without a regression test. An agent can clearly tell what this tool is for.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description is directive about when to use it: 'the moment you find it, BEFORE fixing it.' It also explains the relationship with ddflow_bug_fixed, making it clear that bug_found precedes bug_fixed and is required for accountability. This is explicit usage guidance with a named alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_cadenceB

Which periodic whole-repo passes are due — integration tests, architecture review, mutation testing, dedupe sweep, lessons compression. Derived from completed work, so there is no state file to drift.

ParametersJSON Schema
NameRequiredDescriptionDefault
ranNoRecord that this cadence just ran.
noteNoWhat the pass did, recorded with it.

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of disclosure. It does provide one useful trait: 'Derived from completed work, so there is no state file to drift.' However, it fails to disclose that passing `ran`/`note` records a pass run, which is a mutating behavior an agent needs to know before invoking the tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler. The first sentence immediately states the core purpose and gives examples; the second adds a valuable behavioral guarantee. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The query behavior is well described and the schema covers the `ran`/`note` parameters, so the core use case is usable. However, the description does not mention the record/completion mode at all, nor clarify what 'due' means in cadence terms, leaving a moderate gap for a tool with no output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds meaning by enumerating concrete pass names and the kinds of work the cadence tracks, which helps an agent infer plausible values for `ran` and the context for `note`. It does not map directly to parameters, but the examples are genuinely useful.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies a specific resource and scope: periodic whole-repo passes that are due, with concrete examples (integration tests, architecture review, mutation testing, dedupe sweep, lessons compression). It is clearly distinct from sibling status/progress tools, though it lacks an explicit verb like 'list' or 'show'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the usage scenario—check which periodic passes are due—but it does not explicitly state when to use this tool versus alternatives or when not to use it. No exclusions or sibling comparisons are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_claimA

Lease an item and create its isolated git worktree. Refuses (exit 3) if another agent holds it or holds an item whose file globs overlap, and names what you could take instead. NEVER steals an expired lease: a crashed agent's worktree often holds finished work.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesItem id to claim.
noteNoWhat you intend to do.
forceNoOverride a refusal. Legitimate for exactly one thing: retrying after `ddflow_recover` has told you a crashed agent's worktree holds nothing. Forcing past a dependency or a live lease is how two agents end up writing the same file, and the override is recorded either way.
globsNoComma-separated path globs this work will write.
no_worktreeNoLease the item without creating a worktree. For work that is not a code change — a research or review task.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses refusal behavior (exit 3 on conflicts), the condition for legitimate force, the recording of overrides, and the rationale behind not stealing expired leases. This goes beyond basic operation and covers important side effects and failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, front-loaded with the core purpose, followed by critical behavioral warnings. Every sentence earns its place, with no fluff. The ALL-CAPS warning is effective for emphasis without being verbose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (5 params, conflict handling, force semantics) and the absence of an output schema, the description covers the essential behavioral context: refusal conditions, force legitimacy, and the no-steal rule. It does not describe the full lease lifecycle (e.g., what happens on success), but that is likely covered by sibling tools. The schema fills parameter details, making this sufficiently complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema covers all 5 parameters with detailed descriptions (100% coverage), so the baseline is 3. The main description adds no parameter-specific meaning, but the schema's descriptions—especially for force—are thorough. The description does not need to repeat what the schema already explains.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('lease') and resource ('an item') plus the concrete action ('create its isolated git worktree'). It is unambiguous and clearly distinguishable from sibling tools like release, recover, or merge, which have different purposes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides clear usage context: it refuses on conflicts, explains when the force override is legitimate (after ddflow_recover), and warns against stealing expired leases. It does not explicitly name alternatives, but the force description references ddflow_recover, giving conditional guidance. This is sufficient for an agent to decide when to call it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_cleanupA

Classify every ddflow worktree and branch: merged (safe to remove), unmerged (carries commits nobody landed), dirty (uncommitted edits — a human looks), orphan, or stale branch. Reports by default; with apply=true it removes merged worktrees and branches and lands commits for items the queue already considers done. A dirty tree is NEVER touched automatically, whatever you pass — it is the only thing here that exists nowhere else.

ParametersJSON Schema
NameRequiredDescriptionDefault
applyNoPerform the safe actions.

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations to rely on, the description carries the full disclosure burden and succeeds: it reveals that apply=true removes merged worktrees/branches, lands commits for queue-done items, and that dirty trees are never modified regardless of apply. This goes well beyond the schema by detailing side effects and the key safety boundary.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, each earning its place: classification categories, default report mode, apply behavior, and the dirty-tree safety warning. The most important information is front-loaded, and there is no filler or repetition of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description gives a full picture for a one-parameter tool with no output schema: categories, default behavior, apply semantics, and a critical safety exception. Minor gaps remain — 'orphan' and 'stale branch' are not formally defined, and 'lands commits' references an internal queue concept — but these do not prevent an agent from invoking the tool correctly or understanding its effects.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already describes apply as 'Perform the safe actions' with 100% coverage, so the baseline is 3. The description adds value by explaining what those safe actions actually are (removing merged items and landing done commits) and by clarifying that the default is a report-only run when apply is not true. This is useful enrichment beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action (classify every ddflow worktree and branch), enumerates the classification categories (merged, unmerged, dirty, orphan, stale), and states the two operational modes: report by default and remove with apply=true. This clearly differentiates it from sibling tools as a cleanup/analysis utility.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly states the default behavior (reports) and when to use apply=true to perform removals, plus an explicit exclusion: dirty trees are never touched automatically. It does not name alternative sibling tools or state when NOT to use this tool in favor of another, but the usage context is still clear for an agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_companionsA

Which companion MCP servers serve this project's gates, which are installed on this machine, and which are wired into an agent's config. ddflow imposes the pipeline; it does not perform the judgement inside most gates — standards wants an automated standards review, research wants documentation to check a claim against, rules wants memory. A project with none of them has agent gates passing on assertion alone. Exit 2 means a default companion is missing or unregistered. Read-only: it detects and advises, it never installs anything.

ParametersJSON Schema
NameRequiredDescriptionDefault
no_probeNoSkip the detection probes (faster, less certain).

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden and meets it well. It explicitly states the tool is read-only, that it detects and advises rather than installs, and it explains what exit 2 means, which is critical for interpreting invocation results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the tool's purpose, and each subsequent sentence earns its place by explaining gate semantics, the no-companion implication, exit-code meaning, and read-only behavior. There is no filler, repetition, or vague padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no required parameters and no output schema, the description provides strong context: what it reports, the domain rules around gates, and a key error condition. It falls just short of perfect because it does not sketch the expected result structure or what a healthy versus unhealthy companion setup looks like beyond exit 2.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already fully documents the single parameter no_probe with 'Skip the detection probes (faster, less certain)', so the description does not need to add much. The description's detection language aligns with the parameter, but adds no new semantic detail beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names the resource (companion MCP servers) and the three dimensions it reports on: which serve the project's gates, which are installed on the machine, and which are wired into an agent's config. It also distinguishes itself from the mutating sibling ddflow_companions_add by framing this as a read-only inventory.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when the tool is relevant: checking companion coverage for gates, understanding the consequences of having no companions, and diagnosing the exit-2 condition. It does not explicitly name alternatives or exclusions, but the read-only inventory purpose is distinct enough that no sibling tool covers the same role.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_companions_addA

Register companion MCP servers that are ALREADY installed into an agent's MCP config, merging rather than overwriting what is there. Refuses (exit 3) to register one that is not installed, because that writes a launch command which fails mid-task, at the moment a gate told the agent to reach for it.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoComma-separated ids; default: every installed one.
agentsNoComma-separated agent keys (default: claude).

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the full disclosure burden. It does so by revealing the merge-not-overwrite behavior, the exit 3 refusal for non-installed servers, and the reason (a launch command that fails mid-task). This is meaningful behavioral information beyond what the schema alone shows.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences front-load the action and merge semantics, and the second justifies the hard-failure behavior without repetition. No filler words or restatements of the tool name.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-optional-parameter tool with no output schema, the description covers the essential invocation context: target state (already installed), behavior on conflict (merge), and failure mode (exit 3). It does not describe the success return or configuration side effects, but those are minor against the clear usage boundary.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, with both id and agents already described inline. The description adds no parameter-level detail beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action ('Register companion MCP servers'), qualifies the target ('ALREADY installed'), and states the merge behavior that distinguishes it from overwrite-style configuration tools. This is enough to separate it from the ddflow_companions family and other config-related siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The opening clause sets the appropriate condition: this tool is for servers already present in the agent's MCP config. It also gives an explicit when-not signal by stating it refuses (exit 3) to register non-installed servers, with the failure rationale. It does not name an alternative tool, but the inclusion/exclusion boundary is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_completeA

Finish an item. Refuses (exit 3) when a required gate has not passed, when a phase still has open tasks, or when no reviewer came from a different model family than the author. Pass your own model as 'model'.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesItem id.
shaNoCommit sha this shipped as.
forceNoComplete over unmet conditions. Every one is recorded in the event log as overridden, so this is visible forever rather than being the quiet way past a gate. Prefer `ddflow_gate_skip` with a reason: it names the single step you are dropping instead of all of them.
modelNoThe AUTHOR's model.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden and does a good job: it discloses the exit code 3 and the three refusal conditions. The force parameter description adds that overridden conditions are recorded in the event log 'visible forever.' It doesn't describe success return values or side effects, but the core behavioral contract is clear.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, no filler. The action is front-loaded, refusal conditions are compactly listed, and the model instruction is a single imperative. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations and no output schema, the description covers the essential preconditions and the one required call-time parameter. It doesn't state what happens on success, but the schema's force description fills in the override behavior. The main gap is success semantics, which is minor for a completion action.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds a usage hint for the model parameter ('Pass your own model'), but this is more of a call-time instruction than new semantic meaning. The schema already documents id, sha, force, and model adequately.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a clear verb and resource: 'Finish an item.' It then defines what finishing means by listing refusal conditions (exit 3 for unmet gates, open tasks, or reviewer model-family mismatch), which makes the tool's role concrete. It doesn't explicitly contrast with siblings like ddflow_release or ddflow_abandon, but the action is distinct enough to avoid confusion.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The main description gives a direct usage instruction: 'Pass your own model as 'model'.' It also implies when not to call by listing refusal conditions. The force parameter's schema description adds strong routing guidance, explicitly preferring ddflow_gate_skip over force and explaining that overrides are permanently visible in the event log.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_configureA

Read or write .ddflow/config.toml. With no arguments it prints every knob, its value, its source and what it does. With toml, it APPENDS that TOML to the config — the usual use is setting your project's test command: [gate.unit_tests] command = "pytest -q" This is how a project is configured without a shell.

ParametersJSON Schema
NameRequiredDescriptionDefault
setNoDotted key to set, e.g. 'gate.unit_tests.command'. Preferred: it edits in place and works whether or not the section exists.
tomlNoA whole TOML block to append. Fails if it would duplicate an existing table — use `set` instead then.
valueNoThe value for `set`.
filterNoOnly show knobs whose name contains this.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It discloses that `toml` APPENDS (mutates the file), that `set` edits in place, and that appending fails on duplicate tables. It also explains the read behavior (prints knobs, values, sources, descriptions). It does not mention permissions, reversibility, or exact output formatting, but the core mutating and failure behaviors are transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is moderately sized, including a practical example that aids understanding. It front-loads the core purpose ('Read or write...') and then details the two main modes. The example adds clarity but could be trimmed slightly; nevertheless, every sentence contributes to usage understanding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a config tool with 4 optional parameters and no output schema, the description covers the primary use cases: reading all knobs, appending TOML, and editing via `set`. It explains failure on duplicate tables and shows a realistic usage pattern. It does not describe the exact format of the printed knobs or exit codes, but the information provided is sufficient for an agent to call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 100%, so the baseline is 3. The description adds value by explaining the `toml` parameter with a full example and the append behavior, and by clarifying that `set` edits in place. The `filter` parameter is not elaborated in the description, but the schema already describes it. Overall, the description enhances the schema's meaning for the key parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states 'Read or write .ddflow/config.toml', specifying the resource and action. It distinguishes itself from siblings by being the configuration tool for the project, with no sibling performing similar read/write config operations. The verb is explicit and the scope is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides clear usage context: with no arguments it prints all knobs, and with `toml` it appends. It gives a concrete example of the typical use case (setting the test command) and notes that this is how a project is configured without a shell. However, it does not explicitly state when NOT to use this tool or compare it to alternatives like ddflow_setup, leaving the exclusion criteria implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_decision_addA

Record an architectural decision so the project stays consistent and the reasoning survives. Use when you or the operator settle a question about HOW the software is built — a data representation, a boundary, a library choice, an invariant.

ALWAYS set globs to the code it governs: that is what lets the decision be surfaced automatically to whoever works those files later, instead of only being findable by someone who already suspects it exists. Record alternatives too — without it the next agent re-proposes what was rejected.

ParametersJSON Schema
NameRequiredDescriptionDefault
byNo'operator' or 'agent' or a name.
idNoStable id, e.g. 'D1'. Choose one: `supersedes`, commit messages and docs all reference it, and a generated id cannot be cited in advance.
itemNoThe task it arose from.
tagsNoComma-separated tags.
globsNoComma-separated paths this governs.
titleYesThe decision as a one-line statement.
statusNoproposed | accepted (default) | superseded. 'proposed' records a decision the operator has not ratified, which is honest about its standing rather than presenting it as settled.
contextNoThe forces: why a decision was needed at all.
decisionYesWhat was DECIDED (not what was discussed).
supersedesNoComma-separated ids this replaces.
alternativesNoWhat was rejected, and why.
consequencesNoWhat it costs, including what it makes harder.

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description carries the burden of behavioral disclosure. It does reveal useful behavior: decisions are persisted, survive, and can be surfaced automatically to future agents via globs. However, it does not disclose side effects, authorization needs, or what happens on invocation beyond recording.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is tight and front-loaded. Every sentence earns its place: the first defines the tool's purpose, the second paragraph gives prioritized operational guidance for the two most decision-critical parameters.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 12 parameters, no annotations, and no output schema, the description wisely focuses on the non-obvious usage context and the two easy-to-miss parameters. The schema already describes all parameters in detail, so the tool is callable correctly after reading both. It could still mention interaction with the decision_supersede/list siblings, but that is not essential for a record operation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds real value beyond the schema by singling out `globs` and `alternatives`, explaining why they matter and instructing the agent to always set them. This helps the agent choose parameters strategically, not just fill required fields.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description starts with a specific verb and resource: 'Record an architectural decision'. It further scopes the tool with concrete examples of what counts (data representation, boundary, library choice, invariant), making it clearly distinct from the decision_list, decision_show, and decision_supersede siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives an explicit trigger: 'Use when you or the operator settle a question about HOW the software is built'. It does not name alternative tools or state when not to use this one, so it falls just short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_decision_applicableA

The architectural decisions that govern a specific item's declared files. CALL THIS BEFORE IMPLEMENTING: it is how a decision reaches the person writing the code, without them having to know it exists. Returns project-wide decisions too.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesItem id.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the burden of behavioral disclosure. It reveals that the tool returns project-wide decisions and acts as the delivery mechanism for decisions to implementers. However, it does not state whether the operation is read-only, describe the response shape, or mention any ordering or filtering behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences deliver purpose, usage, and an important return-scope detail without filler. The usage directive is front-loaded and the wording is efficient, though the opening sentence could be more verb-driven.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a one-parameter lookup with no output schema, the description covers what the tool returns and when to invoke it. It is slightly thin on output format and explicit side-effect disclosure, but adequate for a simple decision-retrieval tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the single id parameter is already documented. The description only connects id to 'a specific item' without adding format, source, or interpretation details beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it returns architectural decisions governing a specific item's declared files, and adds that it also returns project-wide decisions. It is distinct enough from sibling decision tools like ddflow_decision_list or ddflow_decision_show, though it does not explicitly name an alternative.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'CALL THIS BEFORE IMPLEMENTING' is an explicit, strong usage trigger, and 'it is how a decision reaches the person writing the code' explains its role. It provides clear context but does not mention when not to use it or point to alternative tools.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_decision_listA

Every architectural decision in force. Superseded ones are hidden unless you ask for them — they are kept, never deleted, because how the architecture got here is what a rebuild needs.

ParametersJSON Schema
NameRequiredDescriptionDefault
allNoInclude superseded decisions.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses a key behavioral trait: superseded decisions are kept (never deleted) and hidden unless requested, which is valuable for the agent to understand persistence and retrieval semantics. It does not cover permissions or return format, but for a simple read/list operation, this is substantial transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no redundancy. The core purpose is front-loaded, and the rationale for keeping superseded decisions is a meaningful addition that explains the 'all' parameter's existence. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with a single optional boolean, the description is complete enough. It states what is returned (decisions in force), the default filtering, and the rationale for the option. It does not describe the output structure, but no output schema is present, and for a list tool, the agent likely knows the decision entities from related tools. The description is adequate without oversharing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema covers 100% of parameters; the single 'all' parameter already has a clear description ('Include superseded decisions'). The description adds a bit of rationale ('how the architecture got here is what a rebuild needs') but no new semantic meaning about the parameter's format or effect beyond what the schema states. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: listing architectural decisions that are in force, with the option to include superseded ones. It uses a specific verb ('list' implied) and resource ('architectural decisions'), and distinguishes itself from siblings like ddflow_decision_show by focusing on the full set rather than a single decision. The behavior of hiding superseded decisions by default is explicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implicitly communicates when to use the tool: to see currently enforced decisions, and optionally include superseded ones via the 'all' parameter. However, it does not explicitly compare itself to alternatives (e.g., ddflow_decision_show) or state conditions for choosing this tool over others. It provides context about the default behavior but lacks explicit exclusions or alternative guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_decision_showA

Read ONE architectural decision in full — its context, what was decided, the consequences, and what was rejected. ddflow_decision_list gives you the titles; this is what you read before working against one, and especially before proposing something it already considered.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesDecision id.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. The verb 'Read' clearly signals a non-mutating operation, and the description discloses what the tool returns in terms of content. It does not explicitly state side-effect absence, but for a read tool the semantics are sufficiently transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler. The core purpose and content are front-loaded, and the sibling comparison is placed exactly where it adds value. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter read tool with no output schema, the description fully covers what the agent needs: what the tool does, what content it returns, and when to use it relative to its sibling. Nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single parameter 'id' is documented as 'Decision id.' The description does not add additional parameter-level meaning, but it does not need to because the schema already fully covers the parameter. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb ('Read') and a specific resource ('ONE architectural decision in full'), and enumerates the content: context, decision, consequences, and rejected alternatives. It also distinguishes itself from ddflow_decision_list, making the tool's role unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly contrasts with ddflow_decision_list ('gives you the titles') and tells the agent when to use this tool: before working against a decision and especially before proposing something it already considered. This is clear, actionable routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_decision_supersedeA

Mark a decision replaced by a newer one. Decisions are never edited or deleted; a reversal is a new decision that names the old one.

ParametersJSON Schema
NameRequiredDescriptionDefault
byYesThe decision that replaces it.
idYesThe decision being replaced.
reasonNoWhy it changed.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It clearly states that decisions are never edited or deleted, so this operation does not modify or destroy the original decision—it only marks a replacement relationship. This adds meaningful context beyond the schema, though it does not cover reversibility or error cases.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no wasted words. The primary action is front-loaded in the first sentence, and the second sentence provides essential domain context. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with fully described parameters and no output schema, the description sufficiently explains what the tool does and the underlying immutability principle. It could mention what happens after success, but given the low complexity and clear schema, it is adequately complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, with each parameter already defined ('id' as the decision being replaced, 'by' as the replacement, 'reason' as why it changed). The description does not add extra parameter-level detail beyond the overall purpose, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Mark'), a clear resource ('a decision'), and specifies the action ('replaced by a newer one'). It also explains the domain model (decisions are never edited or deleted), which differentiates this tool from editing/deleting and from sibling tools like decision_add or decision_list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use the tool (marking a decision as replaced) and hints at the prerequisite that a newer decision must exist ('a reversal is a new decision that names the old one'), but it does not explicitly name alternatives or conditions for not using it. This is adequate but relies on inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_doctorB

Integrity and health check: log corruption, dependency cycles, unknown dependencies, orphaned worktrees, stale index.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must stand alone, but it never states whether this tool only inspects and reports, or may also repair/correct issues such as stale indexes or orphaned worktrees. It also does not disclose output format, side effects, or failure behavior, which is significant for a 'doctor' tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is one compact, front-loaded sentence that uses a colon-delimited list to pack five concrete check categories. Every item earns its place and no filler is present.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although the tool has no parameters, there is no output schema and no annotation safety profile, so the description should at least indicate whether the tool mutates state and what it returns/reports. It does neither, leaving an agent unable to predict the outcome of invoking ddflow_doctor.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has zero parameters, so there is nothing for the description to explain; the baseline is 4. The description still adds value by listing the focus areas of the check.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states the tool performs an integrity/health check and enumerates specific conditions it examines: log corruption, dependency cycles, unknown dependencies, orphaned worktrees, stale index. This makes the scope clear, though it is phrased as a noun phrase rather than an explicit verb command and does not contrast with ddflow_status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to run this tool versus related tools such as ddflow_status, ddflow_recover, or ddflow_cleanup. The intended use is only implied by the 'health check' label; there are no conditions, prerequisites, or exclusions stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_gate_recordA

Record the outcome of a gate you performed (research, a review, a bug hunt). outcome is one of passed/failed/unavailable/partial/skipped. IMPORTANT: if a reviewer or tool could not run, record 'unavailable' with a reason — recording it as 'passed' is how an entire review silently vanishes. Pass the reviewer's model so family independence can be checked.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesItem id.
gateYesGate id.
modelNoModel that performed it, e.g. 'gemini-2.5-pro'.
reasonNoRequired for failed/unavailable/partial/skipped.
commandNoThe command you actually ran. This and `exit_code` are what make an outcome evidence rather than an assertion; a gate listed in `gates.evidence_required` is rejected without them.
outcomeYespassed | failed | unavailable | partial | skipped
evidenceNoWhat you ran and what it said. Required by some gates.
exit_codeNoThat command's exit code.
output_fileNoPath to its full output. A digest is recorded, so the claim can be checked against the file later rather than taken on trust.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden for this mutation tool, and it does substantial work. It surfaces a non-obvious failure mode (recording 'passed' when a reviewer couldn't run), explains that command+exit_code turn an assertion into verifiable evidence, mentions the digest being recorded against output_file, and notes the family-independence model requirement. This is meaningful behavior an agent could not infer from the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, all load-bearing, with the core action and the critical warning front-loaded before supporting detail. The 'IMPORTANT' caveat is prominent where it matters most. Slightly dense, but every clause earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-param mutation tool with no annotations and no output schema, the description covers the key behavioral traps (unavailable-vs-passed, evidence requirements, digest verification) that an agent must know to avoid silently losing data. It does not describe the return value or confirm whether an id/gate must pre-exist, but it is otherwise sufficient for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so a baseline of 3 applies. The description goes beyond the schema by explaining the semantic relationship between command and exit_code ('what make an outcome evidence rather than an assertion'), the rejection of gates without them (evidence_required), the requirement that reason accompany failed/unavailable/partial/skipped, and the trust-model purpose of output_file ('a digest is recorded... rather than taken on trust'). It adds genuine interpretability over raw field docs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Record the outcome of a gate you performed') and enumerates the exact allowed values (passed/failed/unavailable/partial/skipped). The instruction distinguishes it from sibling gate tools like ddflow_gate_run and ddflow_gate_skip by emphasizing this is the post-hoc recording action, so an agent can tell them apart.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives strong contextual guidance on when to choose 'unavailable' vs 'passed' and explicitly flags the destructive consequence of misrecording (an entire review silently vanishing). It also instructs passing the reviewer's model for family independence. However, it never names alternative tools (e.g. ddflow_gate_skip, ddflow_gate_status) or states when NOT to use this tool, leaving some routing implicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_gate_runA

Execute a command gate (tests, linters) and record the result with its evidence. Agent gates cannot be run this way; they are recorded with ddflow_gate_record.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesItem id.
gateYesGate id, e.g. unit_tests.

TDQS

A4.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure. It does reveal that the tool executes commands and records results, which implies side effects. However, it does not mention failure behavior, permissions, or what 'evidence' means, so some behavioral context is missing for a command-execution tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no wasted words, and the core action is front-loaded. The second sentence directly addresses the most likely confusion with a sibling tool, making the description both compact and informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter tool, the description covers the essential invocation details and differentiates from the relevant sibling. The absence of an output schema means return behavior is not explicitly described, but the tool's purpose and parameters are sufficiently clear for correct use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaningful parameter context by explaining that the gate is a command gate (tests, linters) and that agent gates are not valid here, which goes beyond the schema's simple 'Gate id' description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Execute a command gate') and resource ('tests, linters'), and clarifies the result is recorded with evidence. It also distinguishes itself from ddflow_gate_record by excluding agent gates, making the tool's scope clear relative to a key sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says when not to use this tool: agent gates cannot be run this way and should instead use ddflow_gate_record. This provides clear routing guidance and names the alternative, leaving no ambiguity about the intended use case.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_gate_skipA

Skip a gate ON THE RECORD, with a mandatory reason. This is the auditable escape hatch, and it is the one to reach for: gates.require_outcome means a gate left silent BLOCKS completion, so the alternative to skipping is forcing past everything at once. A skip names the single step you are dropping and why, and that reason is in the event log permanently. Skipping a gate listed in gates.required still blocks — those are not optional.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesItem id.
gateYesGate id.
reasonYesWhy this step does not apply HERE. 'n/a' is not a reason: the next person reads this to decide whether you were right.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It discloses that skipping requires a reason, that the reason is permanently recorded in the event log, and that required gates still block. It could add whether the operation is reversible or what response shape to expect, but the key side effects are transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, with no filler. The core action and mandatory reason are front-loaded in the first sentence, and the rest supplies decision-relevant constraints. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given three simple required parameters and no output schema, the description covers what the tool does, when to use it, what side effects to expect, and which gates cannot be skipped. An agent has enough to select and invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents id, gate, and reason. The description reinforces that reason is mandatory and auditable, but does not add substantial new semantic detail beyond the schema. The baseline of 3 applies because the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Skip a gate ON THE RECORD, with a mandatory reason.' It then distinguishes this tool from the broader act of forcing past everything, making clear it drops a single named step. The phrase 'auditable escape hatch' further differentiates it from gate-related siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says this is 'the one to reach for' when a gate is blocking completion, and contrasts it with forcing past everything at once. It also sets a clear exclusion: gates listed in `gates.required` still block and are not optional, so an agent knows when NOT to use this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_gate_statusA

Where an item stands in its quality pipeline, which gate is next, and the instruction for that gate. Gates marked '?' did not run — that is a coverage gap, never a pass.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesItem id.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of behavioral disclosure. It does add an important interpretation rule: gates marked '?' did not run and are a coverage gap, never a pass. However, it does not explicitly state that this is a read-only operation, mention prerequisites, or describe error/edge-case behavior, leaving some ambiguity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences with no redundancy. The core purpose is front-loaded, and the '?' clarification earns its place by preventing a dangerous misinterpretation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity tool with one parameter and no output schema, the description adequately explains the returned information and includes a key edge-case caveat. It could mention what happens when no gates remain or an item is not found, but those are minor gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already fully documents the single 'id' parameter. The description does not add additional meaning to the 'id' parameter beyond what is already present, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the resource (an item's quality pipeline) and the output: current position, next gate, and gate instruction. It uses 'Where an item stands' rather than an explicit verb, and does not explicitly distinguish itself from siblings like ddflow_status or ddflow_next, so it falls just short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied: call this tool to see an item's current gate status and what gate comes next. However, there is no explicit statement about when to prefer this over sibling tools such as ddflow_status, ddflow_progress, or ddflow_next, and no alternatives are named.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_gate_verifyA

Break what a gate guards and require it to NOTICE. Applies each mutation registered on the gate, runs it, requires a non-zero exit, and restores the file.

This is the anti-vacuous-pass check turned on the checks themselves. A gate that cannot fail is worse than no gate: it reports success on every change and everyone downstream reads that as evidence. Exit 1 means the gate did NOT catch its mutation — or that nobody has registered one, which is the same problem earlier.

A mutation whose old text is absent or ambiguous is a FAILURE, not a skip: the edit never happened, so the gate ran on pristine source and passing proves the opposite of what it claims.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesItem whose worktree to mutate in.
gateYesGate id. Must be a command gate.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full responsibility for behavioral disclosure. It clearly states that it mutates the file, runs the gate, expects a non-zero exit (indicating the gate caught the mutation), and restores the file. It also explains the failure semantics: a mutation with absent or ambiguous 'old' text is a failure, and Exit 1 from the tool indicates the gate did not catch the mutation. This is thorough and unambiguous.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is divided into three paragraphs. The first is concise and action-focused. The second explains the rationale behind the tool, which is useful context but could be shortened. The third clarifies a failure condition. Overall, every sentence adds value, but the middle paragraph is slightly verbose. The structure front-loads the core action and then elaborates, which is effective.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with two parameters, no output schema, and no annotations, the description covers the key behaviors: mutation application, exit-code expectations, file restoration, and failure conditions. It does not explicitly state what the tool returns on success (presumably 0) or how multiple mutations are handled sequentially, but these are inferable from the description. The absence of an output schema reduces the need to explain return values, so the description is sufficiently complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides descriptions for both parameters (id and gate), achieving 100% coverage. The description does not add any additional parameter-specific guidance, such as examples or further constraints beyond 'Must be a command gate.' Since the schema handles parameter documentation, the description's lack of extra detail keeps this at the baseline score.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a concrete, action-oriented statement: 'Break what a gate guards and require it to NOTICE.' It then specifies the exact procedure (apply mutations, run, require non-zero exit, restore file). This clearly distinguishes it from sibling tools like ddflow_gate_run (which presumably runs gates normally) and ddflow_gate_record (which registers mutations).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description establishes clear context: it is the 'anti-vacuous-pass check turned on the checks themselves.' This implies it should be used when verifying that a gate actually catches its registered mutations. While it doesn't explicitly name alternative tools or state when not to use it, the purpose and scenario are evident from the wording and sibling names. It stops short of explicit exclusions or direct comparisons.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_heartbeatA

Renew the lease on an item. Call periodically during long work, or the lease expires and another agent may take the item.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesItem id.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It explains the core effect (renewing a lease) and the consequence of non-action (expiry and takeover), which is valuable. However, it does not disclose failure modes, idempotency, prerequisites, or what the response looks like, leaving some behavioral ambiguity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no wasted words. The action is front-loaded, and the usage guidance follows naturally. Every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with no output schema and no annotations, the description provides the essential what, when, and why. It omits potential edge cases like 'what if the lease already expired?', but given the tool's simplicity, the description is sufficient for the core calling scenario.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already covers the only parameter fully with 'Item id', so the baseline is 3. The description adds minimal semantic value by linking 'the item' to the id, but does not provide additional details such as format, scope, or required state of the item.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific operation ('Renew the lease on an item'), which is unambiguous and distinct from generic verbs. It does not explicitly compare itself to sibling lease-related tools such as ddflow_claim or ddflow_release, but the action and resource are clear enough for an agent to understand what it does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit usage guidance: 'Call periodically during long work'. It also explains the negative consequence of not calling (lease expiry and potential takeover by another agent), which helps the agent decide when to invoke it. It does not mention alternatives or exclusions, but the context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_helpA

What ddflow IS, what it can do, and what the workflow is. Call this first if you have not used it before — the other tool descriptions explain one tool each to someone who already knows which to pick, and the connection instructions describe THIS repository right now. Neither answers 'how am I meant to work here'.

With no argument: the loop from picking work to landing it, what the exit codes mean, and every capability grouped by what it is for. With a topic: workflow, import, gates, parallel, memory, recovery, config.

Read-only. The pages are templates a project can override, so what this returns may be this project's own instructions rather than the defaults.

ParametersJSON Schema
NameRequiredDescriptionDefault
topicNoworkflow | import | gates | parallel | memory | recovery | config. Omit for the overview, which lists them.

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Even though no annotations are provided, the description discloses that the tool is 'Read-only' and warns that returned pages may be project-specific overrides rather than defaults. This is useful behavioral context that prevents the caller from assuming the output is always the canonical template.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with its purpose, then moves clearly through when to call it, what each argument form returns, and the read-only project-override caveat. Every sentence earns its place; the contrast with sibling descriptions is relevant, not padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a help/metadata tool with one optional parameter and no output schema, the description provides everything needed to invoke it correctly: purpose, first-use trigger, argument behavior, allowed topics, side-effect safety, and variability of the response. There are no major gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already documents the optional topic parameter and its allowed values, and the description repeats that list without adding deeper per-topic semantics. Schema description coverage is 100%, so the description contributes little beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states that ddflow_help is the overview tool: 'What ddflow IS, what it can do, and what the workflow is.' It also distinguishes itself from the one-tool-at-a-time sibling descriptions by explicitly saying those assume you already know which tool to pick. The intended role, scope, and argument modes are all immediately clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit trigger: 'Call this first if you have not used it before.' It also explains why the sibling descriptions are not substitutes and tells the caller exactly when to omit the topic versus provide one. This is strong when-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_historyA

ONE timeline of everything that happened, in the order it happened: claims, releases, gates, bugs, decisions, lessons, completions. The other views answer 'what is true now'; this one answers 'how did it get like this', which is the question you have when something looks wrong.

Filter with item for one task's whole life, kind for one family ('gate', 'lease.acquired', 'decision,bug'), since for a time window. Exit 2 means nothing matched — which is an answer, not a failure.

ParametersJSON Schema
NameRequiredDescriptionDefault
itemNoRestrict to one item's timeline.
kindNoComma-separated event kinds or families: 'gate', 'lease.acquired', 'decision,bug'.
limitNoMost recent N entries (default 40).
sinceNoISO timestamp lower bound.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It discloses that results are ordered chronologically and adds a valuable operational detail: exit 2 means nothing matched and should be treated as an answer, not a failure. It does not explicitly state permissions or read-only behavior, but that risk is low for a history view.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense paragraphs with no wasted repetition. The first front-loads the core purpose and selection context; the second delivers filter semantics and the non-obvious exit-code meaning. Every sentence contributes to correct selection or invocation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a history-listing tool with no output schema, the description covers purpose, ordering, filtering, and a non-obvious exit-code behavior. It does not specify the exact return shape, but 'ONE timeline' plus the event-family list is enough for an agent to invoke and interpret it. The main gap is not naming specific sibling alternatives.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already covers all four parameters, so the baseline is 3. The description adds conceptual meaning: `item` means one task's whole life, `kind` groups event families, and `since` defines a time window. It does not add much about `limit`, but the schema already documents that.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly defines a single timeline of events in chronological order and enumerates the event families it covers. It contrasts itself with 'the other views' that answer current state, which helps an agent distinguish it from status-like siblings, though it never names a specific alternative tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit context for when to use the tool: when something looks wrong and the question is 'how did it get like this.' It also explains filter usage and exit code 2 as a valid no-match result. However, it does not name specific sibling tools or state when not to use it, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_hooksA

Inspect or install the enforcement git hook — the one layer of this workflow that does not depend on the agent agreeing. It refuses a commit touching paths no live lease of yours covers. status reports whether it is installed AND whether the policy actually blocks, since a block policy with no hook installed enforces nothing.

ParametersJSON Schema
NameRequiredDescriptionDefault
actionNostatus (default), install, uninstall.

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description must carry the full burden of behavioral disclosure. It explains the hook's behavior (refuses commits on uncovered paths) and clarifies the nuance between installation status and actual enforcement. However, it doesn't describe side effects of install/uninstall (e.g., whether they modify files, require permissions, or are reversible), which would be valuable given no annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, with two sentences that pack essential information. It front-loads the core purpose (inspect/install the hook), explains its uniqueness, and clarifies the critical status nuance. No wasted words; every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with one parameter and no output schema, the description is quite complete. It covers the tool's purpose, the key behavioral nuance (status vs enforcement), and the action options. The only minor gap is the lack of detail on what install/uninstall do in terms of side effects or prerequisites, but given the tool's simplicity, this is a minor omission.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description is 100% covered (the action parameter is fully described with its options). The description adds value by explaining what 'status' reports (installation AND policy effectiveness) and implying install/uninstall actions, going beyond the bare schema. Since schema coverage is high, a baseline of 3 applies, but the description adds meaningful context about the action's semantics, justifying a 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: inspect or install an enforcement git hook that blocks commits touching paths not covered by a live lease. It uses specific verbs ('inspect', 'install') and a specific resource ('enforcement git hook'), and it distinguishes itself from siblings by emphasizing it's the one layer not depending on agent agreement—a unique function among the sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use it: 'status' reports installation and whether the policy blocks, which is useful context. It implicitly suggests using install when a block policy exists but no hook is installed. However, it doesn't explicitly name alternatives or when-not to use it, though the description's focus on the enforcement layer makes the usage context clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_importA

For a project that ALREADY HAS HISTORY and is adopting ddflow now: read its todo checklists, lessons corpus, ADR files and unmerged branches, and propose them as queue items. Reports by default and writes NOTHING until apply is true.

Call this right after ddflow_setup on any repository that is not brand new. A queue that starts empty tells you nothing is in flight about a project that may have three branches in flight.

The proposal is a GUESS about structure — headings became phases, checkboxes became tasks, and almost nothing has globs. Use the import-existing-project prompt, which walks through fixing that with the operator. Exit 2 means nothing was found.

ParametersJSON Schema
NameRequiredDescriptionDefault
applyNoWrite the proposal. Default false: look first.
max_tasksNoRefuse to propose more tasks than this (default 200).
include_doneNoAlso import already-ticked items as completed. Off by default — a finished history is not a queue, and one real project yielded 3,638 of them.

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses the non-destructive default ('writes NOTHING until `apply` is true'), the guess-based nature of the proposal, and the exit code meaning (Exit 2 = nothing found). This is thorough behavioral disclosure for a tool that reads and proposes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is longer than average but every sentence adds value: purpose, usage timing, rationale, guess caveat, and exit code. It is front-loaded with the core purpose and default behavior. Slightly verbose but not wasteful.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema and no annotations, the description covers the essential context: what it does, when to use, what it doesn't do, and how to interpret failure. It does not describe the return format or the structure of the proposal, but that is minor given the other details. Overall, adequately complete for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds context about `apply` (default false, reports only) and mentions the guess nature, but does not significantly enhance the parameter meanings beyond what the schema already provides. It does not explain `max_tasks` or `include_done` beyond schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action: read existing history artifacts (todo checklists, lessons, ADRs, unmerged branches) and propose them as queue items. It also clarifies the default behavior (reports only, no writes until `apply` is true), which distinguishes it from write-oriented tools. The purpose is unambiguous and distinct from siblings like ddflow_import_verify.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use: 'Call this right after `ddflow_setup` on any repository that is not brand new.' It explains why (empty queue hides in-flight work) and points to an alternative (`import-existing-project` prompt) for fixing the guessed structure. This gives clear decision guidance without ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_import_verifyA

Was this project's history imported, is that import still true, and did anyone FINISH it? Read-only; writes nothing.

Three answers in one call. STATUS: how many phases, tasks, branches, lessons, decisions, research notes, journal entries and memories carry import provenance, and when. STILL TRUE: whether the source files have moved on since (and what a re-run would add), and whether any imported item names a source file that no longer exists. FINISHED: the half the import-existing-project prompt asks a human for and nothing else checks — imported tasks with no globs, which the conflict detector cannot protect, and phases whose heading claims the work shipped while a task under them is still open.

Call it after any import, and whenever you are about to hand out imported work. Exit 1 means findings you should put to the operator; exit 2 means nothing was ever imported, which is an answer, not a failure. It does not repeat what ddflow_doctor covers — unresolved dependencies, duplicate globs, cycles — so run that too.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description carries the full behavioral burden. It fully discloses the read-only nature ('Read-only; writes nothing'), the meaning of exit codes, what each of the three answer groups covers, and its relationship to other checks. This is exemplary for a no-annotation tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence carries substantive information: the core question, the three answer categories, exit code semantics, and the boundary with ddflow_doctor. It is front-loaded with the essential purpose and read-only guarantee, then expands in a tightly structured way.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite having no output schema and no annotations, the description tells an agent everything needed to call the tool and interpret results: what it verifies, what exit codes mean, when to run it, and what it deliberately excludes. The absence of an output schema is compensated by the detailed behavioral explanation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and the schema is already complete with an empty object definition, so there is no parameter meaning left to add. This matches the baseline for a no-parameter tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific, multi-part question—'Was this project's history imported, is that import still true, and did anyone FINISH it?'—and then names the exact resources and checks involved. It is clearly distinct from siblings like ddflow_import and ddflow_doctor.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to call it ('after any import, and whenever you are about to hand out imported work'), explains exit code meanings, and directly differentiates it from ddflow_doctor by saying what it does not cover and that ddflow_doctor should still be run. This is unambiguous usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_lesson_addA

Record a lesson so it is never re-learned. Use after any bug, any operator correction, any surprise. Make the rule transferable — a future agent on a different task must be able to apply it.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoStable id you choose. Referenced by `supersedes`, by commit messages and by the reconstruction; a generated id cannot be cited in advance.
howNoHow to apply or detect it.
whyNoWhy it is true / what went wrong.
ruleNoThe rule in full.
tagsNoComma-separated tags.
titleYesThe rule as a one-line statement.
seen_inNoComma-separated item ids where this was hit. What makes a lesson checkable later instead of merely memorable.
supersedesNoComma-separated lesson ids this replaces. The old one is retired, not deleted — retiring is how the corpus stops growing without losing the record of what was once believed.

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It conveys persistence and deduplication intent ('never re-learned') and a transferability quality bar, but it does not describe write semantics such as append-only behavior, idempotency, what gets replaced, or what response to expect.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action and immediately followed by concrete triggers and a quality bar. Every sentence earns its place and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The when-to-use guidance is strong and the schema fully documents all 8 parameters. However, with no annotations and no output schema, the description omits any sense of what the call returns, how the recorded lesson is later retrieved, or how it interacts with sibling tools like ddflow_lesson_search and the reconstruction workflow.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline of 3 applies even though the description adds no parameter-level syntax or relationships. The 'transferable rule' guidance hints at how to fill fields like rule/why/how, but it does not map to specific parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States the action ('Record a lesson'), the resource (a lesson in the corpus), and the intended outcome ('never re-learned'). The verb distinguishes it from sibling read/search tools like ddflow_lesson_search, and the 'lesson' resource separates it from task_add, phase_add, decision_add, and research_add.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit trigger conditions: 'Use after any bug, any operator correction, any surprise.' It also gives a quality requirement for the content ('Make the rule transferable'). It does not explicitly list exclusions or compare against alternatives, so it falls just short of full when-to-use/when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_loopsA

Detect circular references and runtime loops: dependency cycles, an item claimed and given up over and over, a gate whose verdict keeps flipping, work completed and reopened repeatedly, duplicate items writing the same files, and a queue where events keep arriving but nothing advances. CALL THIS WHEN WORK FEELS REPETITIVE — it is the check that tells you to stop and re-plan rather than trying the same thing again. Returns [] when there is nothing wrong.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It describes the behavior well: detects various loop types and returns [] when nothing is wrong. It doesn't explicitly state whether it's read-only or has side effects, but for a detection tool this is implicitly safe. The description adds useful context about the specific loop patterns it identifies.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured: it starts with the core purpose, lists specific loop types, gives a clear usage directive, and ends with the return behavior. Every sentence adds value without redundancy. It's detailed but not bloated, and the front-loaded purpose makes it easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description explains what the tool does, when to use it, and the return value for the no-issue case. However, it doesn't describe the structure of the returned data when issues are found (e.g., does it return a list of issues, objects, etc.). Since there's no output schema, this is a minor gap. Overall it's sufficient for an agent to decide when to call it, but the return format could be clearer.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has 0 parameters, so schema coverage is 100% and there's nothing to explain. The description focuses on the tool's purpose and behavior, which is appropriate. Baseline for 0 params is 4, and the description adds value by explaining what the tool does rather than parameter details.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool detects circular references and runtime loops, listing specific loop types (dependency cycles, repeated claims, flipping gates, etc.). This is a specific verb+resource that distinguishes it from sibling tools like ddflow_status or ddflow_recover, which focus on other aspects of workflow state.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance: 'CALL THIS WHEN WORK FEELS REPETITIVE' and explains it's a check to stop and re-plan. It doesn't name alternative tools or exclusions, but the context makes the usage clear. A slight gap is not mentioning when NOT to use it beyond the implied 'when work doesn't feel repetitive'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_mergeA

Merge an item's branch into the base branch from the primary checkout, without ever switching its branch.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesItem id.
keepNoKeep the worktree after merging, for inspection.
messageNoMerge commit message.
allow_dirtyNoMerge although the worktree has uncommitted changes. They are NOT included — that is the point of the refusal. Only pass this once you have looked at what is dirty and decided it is build output.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations are absent, so the description must carry behavioral disclosure. It does disclose a critical trait: it performs the merge 'without ever switching its branch,' which is a non-obvious behavior. It also specifies 'from the primary checkout,' indicating the working context. However, it does not explicitly state that this is a mutation operation or mention side effects like refusal on dirty worktree (though that is covered by the allow_dirty parameter). The disclosure of the non-switching guarantee is significant and adds value beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, well-structured sentence that places the action and resource first, then states the operational constraint. There is no filler, repetition, or unnecessary detail. It is both concise and front-loaded, making it easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the core action and a key behavioral guarantee, but it omits broader context such as error conditions, prerequisites (e.g., existence of the item), or what happens on merge conflicts. With no annotations and no output schema, an agent might need more detail about potential failures or the merge's impact. However, given the simplicity of the operation and that parameters are well-described, the description is minimally adequate but not rich.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with all four parameters (id, keep, message, allow_dirty) already documented in the input schema. The description adds no additional parameter-level meaning or elaboration beyond what the schema provides. Per the baseline for high coverage, this is a 3. It does not need to compensate for schema gaps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action (merge), a resource (an item's branch), a target (base branch), and a key constraint (without switching branches). This clearly differentiates it from any potential sibling merge tools by specifying the operational context (primary checkout) and the non-switching behavior. It is a precise verb+resource+constraint statement.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus alternatives, nor does it mention any exclusions or prerequisites. It does not reference sibling tools or provide context on selection criteria. The only hint is 'from the primary checkout,' which is a location detail, not usage guidance. An agent would have to infer from the name alone when merging is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_nextA

What may be started RIGHT NOW, and for everything that may not, the reason. Independent items in the ready set can be run in parallel worktrees by separate agents. Returns ready=[] when nothing is actionable — that is a result, not an error, and it never means 'pick something anyway'.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNo'task' (default) or 'phase'.
phaseNoRestrict to one phase (the 'implement phase X' entry point).

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral disclosure burden. It explicitly states that an empty ready set is a valid result rather than an error, and that agents should not pick work arbitrarily when nothing is actionable. It also discloses that ready items are independent and parallelizable, which is meaningful behavioral context beyond a simple query description.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the core purpose. Every sentence contributes: the first defines what the tool returns, the second explains parallel execution, and the third clarifies empty-result semantics. There is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple read-like tool with full schema parameter coverage, the description provides enough context to select and invoke it correctly. It explains the ready set, reasons for non-actionable items, and empty-result behavior. It does not describe the output shape beyond ready=[], but that is acceptable given the tool's simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% for both parameters ('kind' and 'phase'), so the schema already documents them. The description adds no additional meaning about these parameters, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool identifies what can be started right now and the reasons for what cannot. It conveys the core resource (ready tasks/phases) via the ready set and parallel worktree context, though it does not explicitly differentiate from siblings like ddflow_status or ddflow_progress.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear usage context: call this to know what is actionable immediately, and use the ready set to run independent items in parallel. It also explains the critical interpretation of an empty ready set. It does not name alternatives or exclusion conditions, but the guidance is strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_phase_addA

Add a phase to the queue. A phase is a unit of REVIEW: it gets its own research, its own whole-phase test pass and live smoke run, and it merges as one coherent feature. Group tasks into a phase when they only make sense shipped together.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesShort stable id, e.g. 'P2' or 'auth'.
bodyNoDetail, acceptance criteria, context.
tagsNoComma-separated tags.
globsNoComma-separated path globs this phase writes. Set them: they are what lets two agents work different phases in parallel safely, and the phase's own dependencies are INHERITED by every task inside it.
needsNoComma-separated ids this phase depends on.
titleNoOne-line description.
priorityNoLower is offered first (default 100).

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the full behavioral disclosure burden and largely meets it: it reveals that adding a phase triggers its own research, whole-phase test pass, and live smoke run. The globs parameter description adds further behavioral context about parallel-safety and dependency inheritance. It does not cover preconditions, reversibility, or failure behavior, which keeps it below a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: the action, the behavioral concept, and the usage condition. The verb+object is front-loaded and there is no filler or repetition of schema content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 7-parameter tool with a fully self-documenting schema and no output schema, the description explains the domain concept, the trigger condition, and the downstream consequences of the action. The main gap is not explicitly routing the agent to ddflow_task_add for single-task cases and not mentioning session or queue preconditions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds conceptual framing that enriches understanding of parameters like id ('P2') and body (acceptance criteria), but it does not add parameter-specific guidance beyond what the schema already provides.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource ('Add a phase to the queue') and then defines a phase as 'a unit of REVIEW' that gets its own research, test pass, and smoke run, merging as one coherent feature. This conceptual definition clearly differentiates it from the sibling ddflow_task_add and other workflow tools without ambiguity.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The final sentence gives an explicit when-to-use condition: 'Group tasks into a phase when they only make sense shipped together.' This provides clear selection context, though it stops short of naming the alternative (ddflow_task_add) or stating explicit when-not-to-use scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_progressA

What work has ACTUALLY been done, aggregated from the event log: attempts per item, wall-clock held, gate runs, commits produced, and who did them. Use it to answer 'how much effort has gone into this' and to see an item's full gate history including the outcomes that were not passes.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoOne item, with its per-attempt detail.

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral disclosure burden. It reveals the source is the event log, stresses 'ACTUALLY' done work, and explicitly says non-pass outcomes are included, which is beyond what the name suggests. It doesn't state read-only safety or data freshness, but the reporting nature is unambiguous.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences with no filler. The core action comes first and the use case follows immediately, so the most important content is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

It lists the concrete measures returned (attempts, wall-clock, gate runs, commits, who did them) and the failure-inclusive gate history, which is enough for a single-id query without an output schema. The notable gap is the behavior when no id is supplied, since id is optional in the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The single id parameter is 100% covered by the schema ('One item, with its per-attempt detail'). The description reinforces 'per item' and gate history but adds no new parameter-level meaning beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Opens with 'What work has ACTUALLY been done, aggregated from the event log', giving a concrete resource and scope. It enumerates returned measures (attempts, wall-clock, gate runs, commits, actors), so an agent knows what the tool reports, though it doesn't explicitly name a sibling to distinguish from.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit intended questions: 'how much effort has gone into this' and an item's full gate history including non-pass outcomes. It does not state when not to use it or point to alternatives, so exclusions are left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_promptsA

Inspect the prompt templates this project uses, and where each comes from (shipped default, project override, or an explicit config path). Use eject to copy the shipped ones into .ddflow/prompts/ so the project can edit them as plain text — reviewer instructions and workflow commands are operator-tunable behaviour, not code.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoTemplate name, for show/eject.
actionNolist (default), show, or eject.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It clearly discloses that eject copies shipped defaults into .ddflow/prompts/ and that templates may come from defaults, project overrides, or a config path. It does not state whether existing files would be overwritten, but the main side effects and purpose are transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the core purpose, and every clause contributes useful information. The mention of operator-tunable behavior is a brief but valuable rationale for the tool's existence.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter tool with zero required parameters and no nested objects, the description is complete enough: it covers the main inspect action, the source taxonomy, and the eject action with its destination. It doesn't detail list/show return shapes, but the schema enum and 'for show/eject' note keep an agent adequately oriented.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description adds meaningful semantics beyond the schema: it explains what eject does and clarifies the source categories for templates. The schema already defines the action enum, and the description complements it with real-world behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Inspect the prompt templates this project uses') and identifies the exact resource and scope, including where templates come from. It also names the eject behavior, which clearly differentiates this tool from the many ddflow siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides a clear use case: inspect prompt templates and where they originate, and use eject when the project wants to edit shipped templates as plain text. It lacks explicit exclusions or comparisons to sibling tools, but the context is strong enough to guide selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_rebuildA

Re-derive the search index from the event log. The index is a disposable cache; this is never a data-loss operation.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full disclosure burden. It proactively states that the index is a disposable cache and that the operation is never a data-loss operation, addressing the main safety concern. It could add specifics like runtime or locking effects, but the most important behavior is disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences deliver the action, source, and safety guarantee with no filler. The key verb and object are front-loaded, and every sentence adds information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter tool with no output schema, the description covers what it does, where it derives data from, and whether it is safe. Nothing essential is missing for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and an empty input schema, so there are no parameter semantics to document. The baseline of 4 applies, and the description does not need to compensate for any schema gaps.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action ('Re-derive'), a distinct resource ('search index'), and a source ('event log'), making the tool's function unambiguous. It also clarifies the index's disposable nature, which sets it apart from recovery tools like ddflow_recover.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The context is implicit: use this when the search index needs to be rebuilt from the event log. However, no explicit when-to-use or alternative routing is given, and the agent must infer it from the tool's name and description.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_recallA

'HAVE WE BEEN HERE BEFORE?' — one search across everything this project remembers: architectural decisions, lessons learned, research verdicts, past bugs, similar tasks, and the operator's own earlier prompts.

CALL THIS BEFORE STARTING ANY NON-TRIVIAL WORK. It exists so the operator does not have to say the same thing twice and you do not have to learn the same thing twice. Results are labelled by kind, because a binding decision, a transferable lesson and a prompt from three weeks ago should change what you do in different ways. A decision marked superseded names its replacement — follow the replacement.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoHits per source (default 3).
queryYesWhat you are about to do, in plain words.
sourcesNoComma-separated subset: decisions,lessons,research,bugs,items,prompts. Default: all.
max_charsNoTotal budget for the answer. The point of a budget is that recall is called at the START of work, where a long answer costs the context the work itself needs.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the burden of behavioral disclosure. It explains that results are labelled by kind and that a superseded decision names its replacement, which is specific output behavior. It also alludes to the cost of long answers in the max_chars parameter description, indirectly addressing context consumption. It does not mention read-only status or error handling, but for a search tool these are less critical. The disclosure is adequate and adds value beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is longer than typical but every sentence serves a purpose. It opens with a memorable hook, states the primary function, gives an explicit usage directive, explains the rationale, and describes result labelling. It is front-loaded with the core purpose and does not include fluff. A slight deduction for length, but it remains efficient for the tool's complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema, the description sufficiently explains what the agent gets back (results labelled by kind, superseded decisions pointing to replacements). It also explains when to call it and the context cost rationale. It does not describe pagination, error cases, or exact format, but these are minor for a search tool whose main value is the broad recall and guidance on timing. The description is complete enough for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the schema already documents all four parameters with descriptions. The tool description itself does not add parameter-specific semantics beyond what is in the schema; the only added context about max_chars appears in the schema description, not the tool description. Thus, the description does not compensate for any gaps, but none exist, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description explicitly states the tool's function: a single search across all project memory (decisions, lessons, research, bugs, prompts). It distinguishes from siblings by emphasizing 'one search across everything,' whereas siblings like ddflow_lesson_search or ddflow_prompts are narrower. The verb 'search' and the scope 'everything this project remembers' are concrete and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a direct, explicit directive: 'CALL THIS BEFORE STARTING ANY NON-TRIVIAL WORK.' It explains the rationale (avoid repetition) and clarifies how results should be interpreted (by kind, with superseded decisions pointing to replacements). It does not explicitly list when not to use it, but the 'non-trivial' qualifier implies trivial tasks may not need it, and the all-encompassing scope makes it the default recall tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_recoverA

Find work left behind by a crashed agent: expired leases, orphaned worktrees, items stuck running. Reports what each worktree contains and never deletes anything. Run this at the start of any session that follows an interruption.

ParametersJSON Schema
NameRequiredDescriptionDefault
itemNoRestrict to one item.
applyNoAct on the advice: release the leases and remove the worktrees this reports as holding nothing. It only ever touches a tree MEASURED as having no uncommitted and no unmerged work — one that could not be measured is never removed, because 'could not tell' is not 'empty'.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It explicitly states the tool 'never deletes anything' in its default operation, and the apply parameter's behavior is thoroughly described in the schema. However, the description does not mention that apply can remove worktrees, which could be misleading without reading the schema; this slight ambiguity prevents a perfect score.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences with no fluff. It front-loads the primary purpose, then adds safety and usage context. Every sentence adds value, and the structure is clear and efficient.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with two optional parameters and no output schema, the description covers the essential context: what it finds, its safety guarantee, and when to use it. It does not describe the exact report format or the nature of 'advice,' but these are adequately implied and the tool's complexity is low.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and both parameters are already well-documented in the schema. The description adds no additional parameter-specific meaning beyond what the schema provides, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: finding work left behind by a crashed agent, listing specific artifact types (expired leases, orphaned worktrees, stuck items). It explicitly distinguishes itself from deletion tools by stating 'never deletes anything,' and the 'reports' verb clarifies it's a diagnostic tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides a clear usage trigger: 'Run this at the start of any session that follows an interruption.' It does not explicitly name alternative tools or say when not to use it, but the context is specific enough for an agent to recognize the appropriate scenario.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_releaseA

Give up a lease without completing the item — when you are handing off, stopping, or recovering someone else's abandoned work after inspecting it. The note is recorded in the log and is often the only lasting explanation of why a claim was broken.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesItem id.
noteNoWhy you are releasing it.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose a key side effect: 'The note is recorded in the log and is often the only lasting explanation of why a claim was broken.' This warns that the note is persistent and the action breaks a claim. It does not mention reversibility or permissions, but for a simple lease-release action the disclosure is meaningful.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no filler. The core action and primary use cases are front-loaded, and the note's log behavior is added in the second sentence. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple two-parameter tool with no output schema, the description covers purpose, use cases, and a key behavioral consequence. It does not describe what happens to the item after release or whether the action is reversible, but the phrase 'give up a lease' sufficiently implies the item becomes available again. This is adequate for the tool's simplicity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds value beyond the schema by explaining the note's significance: it is recorded in the log and may be the only lasting explanation for breaking the claim. This gives the optional note parameter more practical weight than the schema's 'Why you are releasing it.'

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Give up a lease without completing the item.' This clearly distinguishes release from completion and states the core action. It also names concrete scenarios (handing off, stopping, recovering abandoned work) that make the tool's role unambiguous among many siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use context: 'when you are handing off, stopping, or recovering someone else's abandoned work after inspecting it.' It also implies a when-not via 'without completing the item,' but it does not name alternative sibling tools or state exclusions explicitly, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_removeA

Take an item out of the queue. The log is append-only, so this RECORDS a removal rather than erasing anything — the item stays in the history and in replay, which keeps the record honest about work that was planned and then dropped. Refuses if another item depends on it.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesItem id.
forceNoRemove although it still has open children, or although other items depend on it. Both leave the queue inconsistent in a way the scheduler then reports, so read the refusal before overriding it.
reasonNoWhy.

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of disclosing behavior. It clearly states that the log is append-only, that the item remains in history and replay, and that removal is refused if another item depends on it. This is valuable non-obvious side-effect information, though it does not cover every consequence of using the force override.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences and front-loads the core action before explaining the append-only behavior and dependency guard. The rationale about keeping the record honest is slightly verbose but earns its place by clarifying why the tool is non-destructive.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema and no annotations, the description covers the essential behavioral context: recording instead of erasing, retention in history/replay, and dependency refusal. The schema covers force and id details. The main gap is lack of sibling differentiation, but an agent has enough information to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds no parameter-specific meaning beyond what the schema already provides for id, force, and reason, so it neither gains nor loses points here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Take an item out of the queue.' It further clarifies that this is a recorded removal, not an erasure, which gives the agent an accurate model of the operation. It does not explicitly distinguish itself from siblings like ddflow_abandon or ddflow_workflow_drop, so it loses the fifth point.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the use case of recording planned work that was dropped, and it states a refusal condition when dependencies exist. However, it never names alternative tools or explains when to prefer another tool like ddflow_abandon or ddflow_workflow_drop, so usage guidance is only implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_renderB

Regenerate the human-readable markdown views (queue, lessons, research) under docs/ddflow/.

ParametersJSON Schema
NameRequiredDescriptionDefault
outNoDirectory for the generated views (default: docs/ddflow).
showNoPrint ONE view instead of writing files: lessons, research, or board. This is what the ddflow:// resources are served from.

TDQS

B3.3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of disclosing behavioral traits. It states it 'regenerates' views, which implies writing/overwriting files under docs/ddflow/, but does not explicitly mention that existing files will be overwritten, whether it is safe, or any side effects. For a mutation tool, this is insufficient disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no filler. It front-loads the core purpose and location, making it immediately clear. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple tool with two optional parameters and no output schema, the description is adequate but not complete. It misses side-effect disclosure and usage context. However, since the schema covers parameters and the tool is straightforward, a score of 3 is fair.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides 100% coverage, describing both parameters clearly (out directory and show mode with allowed values). The description adds no additional meaning beyond the schema, so the baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear action (regenerate) on a specific resource (human-readable markdown views) and location (docs/ddflow/). It is unambiguous and distinguishes itself from sibling tools like ddflow_show which prints views rather than writing files.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus alternatives. It does not mention when to regenerate views versus using ddflow_show or ddflow_board, nor does it state any preconditions or use cases. The only hint is the parameter 'show' which offers an alternative mode, but that is not framed as usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_replayA

Reconstruct the project's whole decision history from the log: every operator prompt in order, every architectural decision, every research verdict, every lesson, and the shape of the queue. This is what rebuilds the project if the code is lost — it reproduces the DECISIONS, not the bytes.

ParametersJSON Schema
NameRequiredDescriptionDefault
outNoWrite a recovery kit to this directory.
verifyNoRe-resolve every recorded commit sha against this repository and report the ones that are gone. A reconstruction citing shas nobody can resolve is a narrative, not a record.

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the burden. It discloses that it writes a recovery kit to a directory (via 'out') and that 'verify' checks shas, but it does not reveal whether it modifies the log, requires special permissions, or has side effects beyond writing files. The safety profile is vague.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with the main purpose front-loaded. The first sentence is a long list but remains efficient. No filler words; every phrase adds value.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has two optional parameters and no output schema. The description explains the core function and parameter meanings, but does not discuss prerequisites (e.g., existence of a log) or output format. Given the moderate complexity and schema coverage, it is fairly complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, but the description adds meaningful context beyond the schema: it explains the purpose of the recovery kit and gives a rationale for the verify flag ('A reconstruction citing shas nobody can resolve is a narrative, not a record'). This goes beyond the schema's basic descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (reconstruct) and resource (project's decision history) and enumerates the exact contents (operator prompts, decisions, research verdicts, lessons, queue shape). This clearly distinguishes it from narrower sibling tools like ddflow_history or ddflow_prompts, which focus on subsets of this history.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides a clear context: 'This is what rebuilds the project if the code is lost.' This tells the agent when to use it, but it does not explicitly name alternative tools or state when not to use it, so it lacks exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_research_addA

Record a research finding. verdict MUST be CONFIRMED, REFUTED or THEORETICAL, and CONFIRMED/REFUTED require a probe — a verdict with no probe behind it is an opinion. A REFUTED entry is as valuable as an adopted one: it stops the next session re-researching it.

ParametersJSON Schema
NameRequiredDescriptionDefault
idNoStable id you choose. Referenced by `supersedes`, by commit messages and by the reconstruction; a generated id cannot be cited in advance.
itemNoThe task this research is for.
claimNoThe falsifiable claim.
probeNoThe command you ran.
budgetNoWhat you allowed yourself, e.g. '30 min, no GPU'.
sourcesNoComma-separated URLs/DOIs you actually opened.
verdictYesCONFIRMED | REFUTED | THEORETICAL
questionYesWhat was asked.
falsifierNoThe single observation that would kill it.
mechanismNoWHY it would work in this repo. The middle field of the triple — claim, mechanism, falsifier — and the one most often skipped.
probe_outputNoIts output, verbatim.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses the key behavioral constraint that CONFIRMED/REFUTED require a probe and that a verdict without a probe is an opinion, plus the value of REFUTED entries. It does not mention side effects, errors, or the nature of the saved record, but the critical rule is surfaced.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no redundancy. The purpose is front-loaded, and the most important rule (probe requirement) is stated immediately after. Every sentence contributes actionable information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 11-parameter tool with no output schema or annotations, the description covers the critical business rule but leaves the rest to schema descriptions. It adequately orients the agent on the verdict/probe requirement and the purpose, but does not guide on how to populate optional fields or what the result of saving is. This is a moderate level of completeness given the complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so each parameter is documented. The description adds meaning beyond the schema by tying verdict to probe and emphasizing the importance of REFUTED, which is not present in the field descriptions. This adds genuine value but does not extensively elaborate on other parameters.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear verb+resource: 'Record a research finding.' It specifies the exact domain and implies a distinct function from siblings like ddflow_lesson_add or ddflow_decision_add, though it does not explicitly name alternatives. The added detail about verdict and probe reinforces the tool's unique role.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage: when you have a research finding to record. It provides conditional guidance on verdict values and the mandatory probe for CONFIRMED/REFUTED, which are effectively usage rules. However, it does not explicitly state when to use this tool versus alternatives or when not to use it, leaving some inference required.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_reviewA

Run the configured cross-family reviewer over an item's diff and record the result. This is the critic gate performed by ddflow rather than claimed by you — it calls a real endpoint, parses the verdict, and records the evidence. If no reviewer is configured, or the endpoint is unreachable, or the model returns no verdict, it records UNAVAILABLE and never a pass.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesItem whose diff to review.
baseNoRef to diff against (default: the item's base branch).
gateNoGate to record under: critic (default) or rubber_duck.
intentNoWhat the change is MEANT to do. The reviewer flags where the diff and the intent disagree, so without it there is nothing to disagree with. Defaults to the item's title and body.
contextNoExtra context to hand the reviewer.

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does so thoroughly. It discloses that the tool calls a real endpoint, parses the verdict, records evidence, and handles three distinct failure cases by recording UNAVAILABLE and never passing. This goes well beyond a generic 'review this item' statement.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short, front-loaded with the action, and keeps critical failure semantics in a compact second sentence. It is slightly repetitive ('record the result' and 'records the evidence'), but there is no filler or irrelevant detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is nearly complete for invocation: it explains the tool's role, mechanism, and failure behavior, with all parameters documented in the schema. The main gap is that, since there is no output schema, it never states what a successful call returns or how the success verdict is surfaced, and it does not mention whether item state changes beyond recording.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the input schema already documents all five parameters. The description adds no per-parameter meaning beyond reinforcing that the reviewer needs an intent to catch disagreements, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Run') and resource ('configured cross-family reviewer over an item's diff') and states the outcome: record the result. It hints at differentiation from agent-claimed critic gates but does not explicitly name or contrast sibling tools like ddflow_gate_run or ddflow_gate_record, so it stops short of a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear situational context: use this when the real configured reviewer should be invoked by ddflow and the evidence recorded, rather than claiming the gate yourself. It also warns that a missing reviewer or unreachable endpoint yields UNAVAILABLE, which helps an agent decide whether to call it, but it never explicitly names alternatives or exclusion conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_reviewers_detectA

Probe well-known local ports for an OpenAI-compatible model server (ollama, vLLM, LM Studio, llama.cpp, sglang) and report what is serving, with each model's pretraining family. Use this to find a reviewer from a DIFFERENT family than yourself — which the critic gate requires. Pass write=true to add what it finds to .ddflow/config.toml.

ParametersJSON Schema
NameRequiredDescriptionDefault
writeNoAppend the discovered reviewers to the config.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the disclosure burden. It clearly states the probing/reporting behavior and the conditional side effect of write=true (adding findings to .ddflow/config.toml). It does not cover error or timeout behavior, but for a simple one-flag tool this is adequate.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no wasted words: the operation is front-loaded, then the use case, then the parameter behavior. Every clause contributes to correct invocation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema, the description states what will be reported (serving model servers and pretraining families), why it matters (critic gate), and how to persist findings. It does not specify return format or no-server-found behavior, but the core facts needed to invoke correctly are present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 100% description coverage for the single write parameter, so the baseline is 3. The description adds slight extra context by naming the exact config file path, but the parameter meaning is already well documented in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific action ('Probe well-known local ports'), a concrete target (OpenAI-compatible model servers), and a distinctive deliverable (each model's pretraining family). It clearly differentiates this from siblings like ddflow_reviewers_list by emphasizing local detection rather than listing known reviewers.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly frames when to invoke the tool: 'Use this to find a reviewer from a DIFFERENT family than yourself — which the critic gate requires.' This is clear usage context, though it does not name alternatives or state when not to use it, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_reviewers_listA

Show the configured reviewers, their families and which gates they serve.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It tells the agent that the tool shows configuration data and is therefore read-only in nature, but it does not explicitly state that no changes occur, nor does it provide details like response format or reliance on configuration. This is minimally adequate for a zero-parameter list tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single concise sentence with no filler, and it front-loads the primary action ('Show') followed by the three pieces of content (reviewers, families, gates). Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has no parameters, no annotations, and no output schema, the description sufficiently explains the return value's content (reviewers, families, gates). It is slightly vague about what 'families' means and whether additional configuration context is included, but it is adequate for a simple list command.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has no parameters, and the schema confirms this, so there are no parameter semantics to clarify. The description adds value by stating what information is displayed, meeting the baseline for a parameterless tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Show') and a specific resource ('configured reviewers'), and goes beyond a simple name restatement by enumerating what is displayed: their families and the gates they serve. This distinguishes it from siblings like ddflow_reviewers_detect, which implies detection rather than listing configured data.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is a read-only listing operation, so an agent can infer it is for inspecting configured reviewers. However, it provides no explicit when-to-use guidance, no conditions, and no mention of alternatives such as ddflow_reviewers_detect or other reviewer-related commands.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_session_endA

Close a session with a summary of what it achieved. The summary is what a later reader sees before deciding whether to open the whole transcript, so write it for someone who was not there.

ParametersJSON Schema
NameRequiredDescriptionDefault
sessionYesSession id.
summaryNoWhat this session achieved.

TDQS

A3.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full behavioral disclosure burden. It mentions the action (close session) and the purpose of the summary, but does not disclose side effects (e.g., termination of session, irreversible closure), permissions required, or the return value. This is insufficient for a potentially destructive operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, no fluff. The primary action is front-loaded, and the rationale for the summary immediately follows. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (two parameters, no output schema), the description is fairly complete: it explains the purpose and gives guidance on summary writing. However, behavior like session closure effect and any implicit requirements (e.g., session must exist) are missing, but the description is sufficient for most use cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema documents both parameters. The description adds contextual guidance for the 'summary' parameter (write for a later reader), which is valuable. However, it does not clarify the 'session' parameter's format or any constraints, so it doesn't go beyond schema in that aspect. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action: 'Close a session' with a summary of achievements. It distinguishes the tool by focusing on session closure and the purpose of the summary, which helps differentiate it from ddflow_session_start and ddflow_session_note.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains that the summary is for a later reader deciding whether to open the transcript, implying when to use this tool (at the end of a session) and how to write the summary. However, it does not explicitly contrast with sibling tools like ddflow_session_note or ddflow_abandon, which might also be considered for ending or noting sessions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_session_noteA

Record something that happened during a session which is neither an operator prompt nor a decision — a surprise, a dead end, why you changed approach. It goes into the reconstruction alongside the prompts, and a dead end recorded is a dead end nobody walks down twice.

ParametersJSON Schema
NameRequiredDescriptionDefault
itemNoItem it concerns.
textYesThe note.
sessionYesSession id from ddflow_session_start.

TDQS

A4.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral disclosure burden. It does disclose that the content 'goes into the reconstruction alongside the prompts,' giving a useful side-effect hintokus, but it does not mention whether the record is append-only, editable, or how it affects later reconstruction. The second sentence adds motivational framing but no concrete behavioral detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, with the core action and boundary conditions in the first sentence and a reinforcing rationale in the second. Every clause earns its place and no space is wasted on repetition of the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple create-note tool with three parameters and no output schema, the description covers the purpose, content guidance, and the relationship to the reconstruction. It could have explicitly noted whether the note is immediately persisted or appended, but the essential invocation context is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3 ate. The description adds meaning beyond the schema by clarifying that 'text' should contain surprises, dead ends, or rationale for changes – content guidance not present in 'The note.' It provides some help for 'session' implicitly via the session start context but adds little for 'item.'

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Record something that happened during a session' – a specific verb and resource – then sharpens the scope by explicitly excluding operator prompts and decisions. It gives concrete examples (surprise, dead end, changed approach) and the mention of 'reconstruction alongside the prompts' distinguishes it from sibling note-taking flows.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly indicates when to use this tool: for unexpected events, dead ends, and strategy changes. It also says what it is not for ('neither an operator prompt nor a decision'), which indirectly routes an agent away from ddflow_session_prompt and ddflow_decision_add, though it does not name those siblings directly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_session_promptA

Record the operator's prompt verbatim. This is what makes the project reconstructible from the log alone if everything else is lost. Secrets are redacted before anything touches disk. Call it once per operator turn.

ParametersJSON Schema
NameRequiredDescriptionDefault
itemNoItem it concerns.
textYesThe prompt, verbatim.
sessionYesSession id from ddflow_session_start.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It mentions that secrets are redacted before touching disk, which is important security context. However, it does not disclose what happens on repeated calls, whether the operation is idempotent, or any side effects beyond recording. This is adequate but not rich.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three concise sentences, each adding distinct value: the primary purpose, the importance for reconstructability, and the security redaction plus usage frequency. It is front-loaded with the core action and avoids any redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple logging tool with 3 fully documented parameters and no output schema, the description is fairly complete. It covers the purpose, the importance, the security redaction, and the expected call frequency. It does not mention error handling or repeated call consequences, but these are minor gaps for such a tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all three parameters (item, text, session). The description adds no additional parameter-level detail beyond implying that 'text' contains the verbatim prompt. This meets the baseline but does not exceed it.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb 'Record' and the resource 'the operator's prompt verbatim', which is specific and unambiguous. It does not explicitly differentiate from sibling tools like ddflow_prompts or ddflow_history, but its focus on verbatim recording of operator turns is distinctive enough.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly states 'Call it once per operator turn', which is a clear when-to-use guideline. It does not mention alternatives or when not to use it, but the frequency instruction provides useful context for correct invocation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_session_startB

Open a session for provenance logging. Returns the session id.

ParametersJSON Schema
NameRequiredDescriptionDefault
toolNoYour harness, e.g. 'claude-code'.
modelNoYour model id.

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of disclosing behavioral traits. It mentions that a session is opened and that an id is returned, but it does not state whether this mutates persistent state, whether the call is idempotent, what side effects occur, or whether a corresponding session_end is required. For a state-changing tool, this is a notable gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences deliver the core purpose and return value with no wasted words. The description is efficiently front-loaded and easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple 2-parameter tool with no output schema, the description covers the basic call and return, but omits important context such as the session lifecycle, whether setup is required, and how the returned id is used by other ddflow tools. It is minimally viable but has clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, with both 'tool' and 'model' already described in the input schema. The description adds no additional meaning about these parameters, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Open a session for provenance logging' and identifies the return value ('Returns the session id'). It distinguishes from session_end, session_note, and session_prompt by signaling the start of a session, though it does not explicitly name or contrast with siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided about when to use this tool, what prerequisites exist, or how it relates to sibling tools like ddflow_setup, ddflow_session_end, or ddflow_recover. The description implies a session-opening role but gives no explicit when-to-use or when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_setupA

Install ddflow into this repository: creates .ddflow/, writes the driver and the AGENTS.md section, and registers nothing else. Run this ONCE per project, then set your test command with ddflow_configure. Safe to re-run — it updates a managed block and leaves your own prose alone.

ParametersJSON Schema
NameRequiredDescriptionDefault
agentsNoComma-separated agents to write driver deltas for: claude,gemini,codex,copilot,kilo,cursor. Default: all.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full disclosure burden and does so well: it states what gets created (.ddflow/, driver, AGENTS.md section), what does not ('registers nothing else'), and the re-run behavior ('updates a managed block and leaves your own prose alone'), which signals idempotency. This goes well beyond a generic 'Installs ddflow' phrasing.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences with zero waste: effects, usage cadence, and re-run safety each get exactly one sentence. The most decision-relevant information (what the tool does) is front-loaded, and nothing is repeated from the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a low-complexity tool (one optional param, no output schema, no annotations), the description covers the essentials: actions performed, when to run, idempotency guarantee, and the next step in the workflow. Minor omissions are the undo path (ddflow_remove exists among siblings but isn't referenced) and what 'driver' refers to, neither of which blocks a correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% — the `agents` parameter is fully documented in the schema with its comma-separated format, accepted values, and default. The description adds no parameter-level meaning, which matches the baseline-3 expectation when the schema already carries the detail.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Install ddflow into this repository') and enumerates the concrete effects: creates .ddflow/, writes the driver and the AGENTS.md section, 'registers nothing else.' The scope is precise enough to distinguish from siblings like ddflow_configure, ddflow_update, and ddflow_remove without needing their schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit cadence ('Run this ONCE per project') and names the follow-up tool ('then set your test command with ddflow_configure'). It also resolves the re-run question with 'Safe to re-run.' It stops short of a full 5 because it doesn't explicitly state when not to use it versus other siblings such as ddflow_update or ddflow_remove.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_showA

Everything known about one phase or task: state, dependencies, declared globs, the lease and who holds it, the worktree path you can cd to, and every gate's outcome with its evidence. Use it to check your own work before calling ddflow_complete.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesItem id.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral burden. It goes well beyond a vague 'show' by detailing exactly what the tool reveals, including the lease holder and worktree path, and implies a read-only inspection operation. It does not explicitly state 'does not modify state,' but the informational content and 'show' semantics make that strongly evident.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, information-dense sentence that front-loads the tool's purpose, enumerates what it returns, and closes with an actionable use case. There is no redundant or filler content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter inspection tool with no output schema, the description provides enough detail about the expected result categories and the calling context. It does not cover error cases or prerequisites, but the low complexity and rich content enumeration make the definition sufficiently complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already documents the single parameter 'id' with 100% coverage. The description adds slight context by framing the item as 'one phase or task,' but does not materially expand on the meaning, format, or source of the id. Baseline 3 is appropriate given full schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a concrete action and resource: it shows everything known about one phase or task, and enumerates the specific categories of information returned (state, dependencies, globs, lease holder, worktree path, gate outcomes). This clearly distinguishes it from broader overview tools like ddflow_status or ddflow_workflow by scoping it to a single item.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit guidance on when to use the tool: 'Use it to check your own work before calling ddflow_complete.' It does not name alternatives or state when not to use it, but the use case is clearly contextualized.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_splitA

Split an item into sub-tasks IN PLACE when the work turns out to be two things. Use this the moment you discover it — mid-task discovery is the normal case, not an exception.

The original keeps its id and history and becomes an umbrella that completes when its children do; closing it and opening two new ones instead would lose the thread between what was planned and what happened. Children inherit the parent's globs, so give each its own afterwards if they write different files — until then they cannot run in parallel.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesThe item to split.
intoYesComma-separated 'sub-id=title' pairs. At least two.
globsNoGlobs for the children (default: inherit).
needsNoDependencies for the FIRST child. The others chain from it if you set theirs with ddflow_update.

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full responsibility for behavioral disclosure. It explains the key behavioral traits: the original keeps its id and history and becomes an umbrella, children inherit globs, and the parallelism constraint (cannot run in parallel until they get their own globs). This is rich, actionable behavior beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and well-structured. It front-loads the primary action and purpose, then adds behavioral nuance and parameter context. Every sentence adds value; there is no redundancy or filler. The two-paragraph structure is easy to scan.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with 4 parameters, no output schema, and no annotations, the description covers the essential context: when to use it, what happens to the original and children, and how parameters affect behavior. It does not explicitly state return values or error handling, but those are less critical for a mutation tool and are not required given the lack of an output schema. It is sufficiently complete for an agent to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaning by explaining how 'globs' are inherited and how to give each child its own, and it notes that 'needs' sets dependencies for the first child while others chain from it via ddflow_update. This provides context that the bare schema lacks, raising the score to 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb (split) and resource (item into sub-tasks), and explicitly says 'IN PLACE' to indicate the core behavior. It also distinguishes itself from the alternative of closing and reopening items, so an agent knows exactly what this tool does and what it doesn't.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit guidance: 'Use this the moment you discover it — mid-task discovery is the normal case, not an exception.' It also contrasts with the alternative (closing and opening two new ones) and explains why that would lose the thread, making the when-to-use and when-not-to-use crystal clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_statusA

The state of the whole project in one answer: how many tasks are done and which, what is in flight and who holds it, what is ready to start, what is blocked, how many agent-hours and commits went in, and whether anything is looping or waiting to be recovered. This is the tool for 'what is the status of this project?' and 'what has been completed?'.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral disclosure burden. It thoroughly discloses what information the status answer includes and implies this is a read-only query through the word 'answer' and 'state'. It could go further by explicitly stating that it makes no changes, but the content coverage is strong.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with no filler. The first sentence front-loads the core value ('state of the whole project in one answer') and packs in the meaningful categories; the second sentence anchors it to natural user queries.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple with no parameters, and the description covers the major dimensions of a project status report, including completion, in-flight ownership, blockers, effort metrics, and loops. It could be more complete by noting relationship to more granular sibling tools, but nothing critical is missing for an agent to invoke this correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes no parameters, and the schema coverage is 100%, so the baseline is 4. The description still adds useful context about the content of the returned status, though parameter semantics are trivial here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool provides an overall project status, enumerating tasks done, in-flight work, blocked items, hours, commits, and loops/recovery. It does not explicitly differentiate itself from siblings like ddflow_progress or ddflow_board, but the scope of 'whole project in one answer' is specific.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly maps the tool to the questions 'what is the status of this project?' and 'what has been completed?', giving clear usage context. It does not name alternative tools or provide exclusion criteria, but the usage context is strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_task_addA

Add a task to a phase. ALWAYS set globs to the paths this task will write: they are what lets two agents work in parallel safely, and an unset glob means the conflict detector cannot protect you.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesShort stable id, e.g. 'P2.T1'.
bodyNoDetail and acceptance criteria.
tagsNoComma-separated tags.
globsNoComma-separated path globs this task writes.
needsNoComma-separated ids this task depends on.
phaseNoOwning phase id. Give this OR `parent`.
titleNoOne-line description.
parentNoOwning phase id OR another TASK's id — a task parent makes this a SUB-TASK, which carries its own globs and dependencies and runs in parallel with its siblings like any other task. Same field as `phase`; both names exist because the CLI has both, and an argument that exists in one surface and not the other is a trap.
priorityNoLower is offered first (default 100).

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of behavioral disclosure. It reveals that globs are critical for conflict detection and parallel safety, which is valuable. But it does not describe other behaviors such as mutation effects, validation, or return values, so transparency is partial.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, with the purpose in the first and a critical usage instruction in the second. It is front-loaded and every sentence earns its place, with no fluff.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (9 parameters, no output schema, no annotations), the description only addresses one aspect (globs) and leaves out broader context such as how tasks relate to phases, prerequisites, or typical usage patterns. It is adequate but not comprehensive for an agent to fully understand when and how to use it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds significant meaning to the 'globs' parameter by explaining its purpose and consequence of omission, which goes beyond the schema's simple description. Other parameters are not elaborated, but the added emphasis on globs justifies a 4.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the verb ('Add') and resource ('a task to a phase'), which is specific and unambiguous. However, it does not explicitly differentiate from sibling tools like ddflow_phase_add or other 'add' tools, so it lacks explicit sibling distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides strong guidance on setting globs for parallel safety, which is useful context for using the tool correctly. However, it does not mention when to use this tool versus alternatives or any exclusions, leaving tool selection to implication.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_updateA

Change an item's fields. MOST IMPORTANT USE: widening globs when your work turns out to touch files outside what you claimed. Do that BEFORE writing them — the conflict detector and the commit hook both work from the declared globs, so an undeclared file is a file no one is protecting and the commit will be refused.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesItem id.
bodyNoNew detail / acceptance criteria.
tagsNoComma-separated tags.
globsNoComma-separated path globs this item writes.
needsNoComma-separated ids it depends on. Pass an EMPTY string to clear them — that is how you break a dependency cycle the loop detector found.
titleNoNew title.
priorityNoLower is offered first (default 100).

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the transparency burden. It does disclose a meaningful behavioral consequence: undeclared files are unprotected and commit will be refused, which is valuable. However, it does not describe return values, error behavior, persistence semantics, or whether the update is partial or full.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded: first the general purpose, then the most important operational guidance. Every sentence earns its place, and the critical warning is stated directly without filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the essential update purpose and the key `globs` caveat, making it usable in the highlighted scenario. However, with no output schema and no annotations, it is incomplete about what happens after a successful update, how errors surface, and whether unspecified fields are left untouched or reset.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds real semantic value for `globs`, explaining that it should be widened before writing and why declared globs matter for protection and commit acceptance. It does not need to add detail for the other parameters since the schema already documents them.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool as changing an item's fields, which is a specific verb+resource. It also highlights the most important use case, widening `globs`, which helps differentiate intent, though it does not explicitly contrast with any sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a concrete when-to-use rule: widen `globs` before writing files that fall outside the originally claimed set. It explains why this matters via the conflict detector and commit hook, but it does not name alternatives or state when not to use the tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_workflowA

The rules THIS project runs by, in one answer: the gates every task and phase passes through in order, which are commands and which you perform yourself, which are required, which need evidence, which need a different-family reviewer, which have been PROVEN able to fail — plus the completion rules, the parallelism caps, the reviewers, and where each value came from (a default, this project's config, or the environment).

Call it before your first ddflow_claim in a session, and after any workflow change: the connection instructions are computed once when the server starts, so a pipeline edited mid-session is not reflected there.

Exit 1 means the workflow does not hang together — most importantly a pipeline naming a gate that has no definition, which blocks every item that reaches it forever. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does so thoroughly: it declares the tool read-only, explains exit code 1 semantics (workflow does not hang together, with the concrete blocking consequence), and discloses the staleness behavior of the computed connection instructions. This goes well beyond a typical minimal disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: content enumeration, invocation timing with staleness caveat, and error semantics with read-only note. The core purpose is front-loaded. The first sentence is dense, but that enumeration is the tool's actual deliverable, so the length is justified.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter, read-only informational tool with no output schema and no annotations, the description is complete: an agent knows exactly when to call it, what it returns, that it fails fast with exit 1, and when its content is stale. Nothing needed for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters and an empty schema, so the baseline is 4. The description adds value by explaining what the response contains (gate rules, commands, reviewers, provenance of values) and when results are current, which is useful context beyond the empty schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a clear resource (the project's complete workflow rules: gates, commands, reviewers, completion rules, parallelism caps) with an implied retrieval verb, and positions it as the comprehensive 'one answer' overview. It doesn't explicitly name a sibling it differs from, but its holistic scope clearly sets it apart from per-gate (ddflow_workflow_gate) and pipeline-editing (ddflow_workflow_pipeline) tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-call guidance: 'Call it before your first ddflow_claim in a session, and after any workflow change.' It also states when not to rely on it — after a mid-session pipeline edit, since connection instructions are computed once at server start — which is a clear exclusion and correctness caveat.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_workflow_dropA

Take a gate out of both pipelines, and out of required so it does not become a requirement that quietly requires nothing. WRITES to config.

The gate's DEFINITION is left in place, so putting it back is one call. Exit 2 means it was in neither pipeline.

Ask the operator first: a gate in a pipeline is a check somebody added on purpose, and removing it weakens every future item.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesThe gate id to remove from the pipelines.
dry_runNoReport the change and write nothing.

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses that the tool WRITES to config, leaves the gate's definition in place for easy re-adding, and returns exit code 2 when the gate is in neither pipeline. This is strong behavioral disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the core action, followed by side effects, exit-code behavior, and operator caution. Every sentence adds useful information with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a two-parameter mutation tool with no output schema and no annotations, the description covers the write side effect, reversibility, exit-code interpretation, and a necessary human-approval caution. The schema covers the parameters, so nothing essential is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both 'id' and 'dry_run'. The description adds no parameter-specific meaning beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Take a gate out of both pipelines, and out of required'. It is clear what the tool does and it is distinguishable from siblings by its action, though it does not explicitly name an alternative tool.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear usage context: removing a gate weakens future items, so the operator should be asked first. It also explains the 'Exit 2' meaning. It does not explicitly contrast with sibling tools or state when not to use it, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_workflow_gateA

Define or change one gate, and optionally put it in a pipeline. WRITES to this project's config.

command makes it a COMMAND gate: ddflow runs it and the exit code is the evidence. prompt makes it an AGENT gate: you perform it and record what you did. A gate needs one of the two — one with neither tells an agent nothing and gives a reviewer no contract, so it is refused.

into adds it to a pipeline (after places it; default is last). required means an item cannot complete without it.

Ask the operator first, and prefer dry_run to show them the change.

ParametersJSON Schema
NameRequiredDescriptionDefault
idYesThe gate id, e.g. 'lint' or 'security_scan'.
cwdNo'worktree' (default) or 'repo'.
intoNoAdd to the 'task', 'phase' or 'both' pipeline(s).
afterNoPlace it after this gate. Default: last.
titleNoHuman-readable name.
promptNoWhat an agent must do. Makes it an agent gate.
commandNoShell command to run. Makes it a command gate.
dry_runNoReport the change and write nothing.
timeoutNoSeconds before the command counts as unavailable.
requiredNoAn item cannot complete without it.
reviewerNo'different_family' to require a reviewer from another model family, or 'same_family_ok'.
applies_toNo'task', 'phase' or 'both'.

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It explicitly states 'WRITES to this project's config', making the mutation clear. It also warns that a gate without command or prompt is refused, and suggests dry_run for previewing changes. This is substantial transparency for a write tool, though it does not detail permission requirements or reversibility.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured, front-loaded with the purpose and write warning, then organized by key behaviors. It uses formatting (backticks) to reference parameters and keeps sentences purposeful. It is not overly verbose, though the operator note at the end could be seen as slightly tangential.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 12 parameters and no output schema, the description covers the core behaviors: write operation, gate type selection, pipeline integration, and required semantics. It also includes a usage note. The schema fills in the rest, so the description is reasonably complete for an agent to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so parameters are documented. The description adds semantic value by explaining the distinction between command and prompt gates, the pipeline placement via into/after, and the required flag. It also clarifies that exactly one of command/prompt is needed. This goes beyond the schema's per-field descriptions, earning a score above the baseline of 3.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb and resource: 'Define or change one gate, and optionally put it in a pipeline.' It distinguishes between command and prompt gates, making the tool's role unmistakable. It also implies a contrast with sibling tools like ddflow_gate_run or ddflow_gate_record, which operate on existing gates rather than defining them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explains when to use the tool: when defining or modifying a gate, and it clarifies the required choice between command and prompt. It also advises asking the operator first and preferring dry_run. However, it does not explicitly contrast with sibling gate tools (e.g., gate_run, gate_record) or state conditions to avoid using it, leaving some inference to the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

ddflow_workflow_pipelineA

Set the ordered list of gates a task or a phase must pass. WRITES to this project's config.

Validated before anything is written: a gate id with no definition is REFUSED and the error names the near miss, because an undefined gate in a pipeline blocks every item that reaches it and cannot be recorded or skipped. Define the gate first with ddflow_workflow_gate.

Ask the operator before changing a pipeline. It governs every future item, not the one you are working on, and removing a gate removes a check somebody added deliberately. Use dry_run to show them what it would do.

ParametersJSON Schema
NameRequiredDescriptionDefault
gatesYesComma-separated gate ids, in the order they run.
whichYes'task' or 'phase'.
dry_runNoReport the change and write nothing.

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and meets it: it discloses that the tool WRITES to project config, validates gate definitions before writing, refuses with a near-miss error when a gate is undefined, and explains the downstream consequences (undefined gates block every item and cannot be recorded or skipped). The need to ask the operator is also explicitly stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core action and write effect, then gives precise validation and operator-consent guidance in compact paragraphs. Every sentence contributes either a rule, a rationale, or an action, with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 3-parameter config-writing tool with no output schema, the description covers the prerequisites, the validation failure mode, the human-consent requirement, and the dry-run escape hatch. No critical information an agent needs to invoke this tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already covers all three parameters at 100%, so the baseline is 3; the description adds value by explaining that `gates` must reference already-defined gates and that `dry_run` is the preview path for getting operator approval. It does not restate the schema descriptions, which keeps this dimension above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Set the ordered list of gates a task or a phase must pass,' immediately clarifying this is a config-writing operation. It also distinguishes itself from related gate tools by naming `ddflow_workflow_gate` as the prerequisite for defining gates before a pipeline is set.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit operating instructions: define gates first with `ddflow_workflow_gate`, ask the operator before changing a pipeline, and use dry_run to preview changes. It also explains why these precautions matter, so an agent knows when to pause and when to proceed.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 63 tool updatesv0.1.0
    • First observedddflow_abandon
    • First observedddflow_block
    • First observedddflow_board
    • First observedddflow_brief
    • First observedddflow_bug_fixed
    • First observedddflow_bug_found
    • First observedddflow_cadence
    • First observedddflow_claim
    • First observedddflow_cleanup
    • First observedddflow_companions
    • First observedddflow_companions_add
    • First observedddflow_complete
    • First observedddflow_configure
    • First observedddflow_decision_add
    • First observedddflow_decision_applicable
    • First observedddflow_decision_list
    • First observedddflow_decision_show
    • First observedddflow_decision_supersede
    • First observedddflow_doctor
    • First observedddflow_gate_record
    • First observedddflow_gate_run
    • First observedddflow_gate_skip
    • First observedddflow_gate_status
    • First observedddflow_gate_verify
    • First observedddflow_heartbeat
    • First observedddflow_help
    • First observedddflow_history
    • First observedddflow_hooks
    • First observedddflow_import
    • First observedddflow_import_verify
    • First observedddflow_lesson_add
    • First observedddflow_lesson_search
    • First observedddflow_loops
    • First observedddflow_merge
    • First observedddflow_next
    • First observedddflow_phase_add
    • First observedddflow_progress
    • First observedddflow_prompts
    • First observedddflow_rebuild
    • First observedddflow_recall
    • First observedddflow_recover
    • First observedddflow_release
    • First observedddflow_remove
    • First observedddflow_render
    • First observedddflow_replay
    • First observedddflow_research_add
    • First observedddflow_review
    • First observedddflow_reviewers_detect
    • First observedddflow_reviewers_list
    • First observedddflow_session_end
    • First observedddflow_session_note
    • First observedddflow_session_prompt
    • First observedddflow_session_start
    • First observedddflow_setup
    • First observedddflow_show
    • First observedddflow_split
    • First observedddflow_status
    • First observedddflow_task_add
    • First observedddflow_update
    • First observedddflow_workflow
    • First observedddflow_workflow_drop
    • First observedddflow_workflow_gate
    • First observedddflow_workflow_pipeline

TDQS

A3.6/5.0

Scored across 63 tools

Disambiguation3/5

The tools are grouped by domain and the descriptions are unusually explicit about boundaries, but at 63 tools several close pairs exist (release/abandon/remove/block, status/progress/board/next, recall/lesson_search/history/replay) whose distinctions are clear only after reading long descriptions. Some overlap remains and the descriptions do the disambiguating work.

Naming Consistency3/5

All names share the ddflow_ prefix and many use a noun_verb compound (task_add, gate_skip, session_start), but there is a large set of bare verbs (release, abandon, update, claim, complete, import) and bare nouns (status, board, workflow, history) that do not follow that pattern. The convention is readable but mixed.

Tool Count2/5

Sixty-three tools is far beyond even the 16–25 heavy range and would be difficult for an agent to navigate in one namespace. The domain is genuinely broad, so the tools are not redundant, but the surface is still too large for practical selection.

Completeness5/5

The surface covers the full lifecycle: task/phase creation through claim, gate execution, completion, and removal; plus sessions, bugs, decisions, lessons, research, history, recovery, import, and configuration. I can identify no dead-end operations or significant missing capabilities for this workflow domain.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables pipeline-driven task management for AI coding agents, with stage-gated workflows, dependency tracking, artifact versioning, and multi-agent collaboration.
    35 npm
    20
    MIT
  • F
    license
    Not graded
    quality
    A
    maintenance
    Enables orchestrating multiple AI coding tools (Claude Code, Codex CLI, Gemini CLI) on shared repositories with file locking, knowledge capture, task planning, and drift detection via an MCP server.
    161
    -
  • A
    license
    C
    quality
    A
    maintenance
    Enables multiple CLI-based AI agents to collaborate as a coordinated team through shared task queues, shared memory, and a message bus, with DAG orchestration, rate-limit avoidance, parallel dispatching, and long-task management.
    63
    2
    MIT