Fagan
Fagan's MCP server runs a gated, autonomous software-delivery pipeline: decompose goals into plans, dispatch story agents, review code, adjudicate risk, and merge verified changes.
Manage plans: save, list, ingest into Plane, patch role config, pause/resume.
Decompose goals into epics/stories via the product-analyst persona.
Track stories: list ready stories, check status, mark in progress/done, set status, patch story fields.
Dispatch/control agents: spawn headless coding agents in isolated worktrees, resume interrupted stories, checkpoint progress, interrupt.
Advance pipeline: one tick or all plans; dispatch ready stories, run tests, review, open PRs, adjudicate merges.
Review/merge: run code-reviewer, open PR or request changes, approve merge with rebase, CI polling, acceptance re-run.
Escalate decisions: request overlord rulings, read decision audit log.
Inspect config and usage: effective configuration/provenance and subscription usage.
Record repository security audits at a given commit.
Fagan
Spend tokens on judgment, not typing.

A real story (STE-1, PR #986) crossing the board: implemented by an open-weight model, sent back once by review, merged. 14 minutes, time-lapsed.
Try it (macOS; Linux via Ollama or LM Studio), then see the Quickstart:
curl -fsSL https://raw.githubusercontent.com/motock/fagan/master/scripts/remote-install.sh | bashFrontier models cost money per token and are excellent at judgment. Local models run free and are adequate at typing. This pipeline splits software engineering along exactly that line: a frontier model decomposes the work, plans it, reviews the diff, and adjudicates anything risky — while a local model writes the implementation at no marginal cost.
What makes the cheap half trustworthy is inspection. In Michael Fagan's 1976 IBM study, formal inspection found 82% of the defects in the released product — 38 per KLOC, against 8 per KLOC for unit testing. Quality lives in the gate, not in the author. So this project spends its budget on gates: TDD enforced before implementation, an independent review pass, acceptance-oracle grading, a risk-tiered overlord that stops for a human on anything irreversible, and a merge gate that re-runs the suite against the rebased branch before anything lands.
The goal is narrow and specific: enterprise-grade engineering discipline — decomposition, TDD, code review, dependency-ordered delivery — on a $20/month budget.
For detailed reference material, see REFERENCE.md.
Before you start: read Reliability & limitations below. This is an autonomous coding pipeline with real, documented failure modes — it is not a hands-off "describe a feature, get a PR" tool yet.
Platform support
Developed and run day-to-day on macOS. The core (MCP server, dashboard, Claude-backend dispatch/review, the full test suite) is plain Python and CI tests it on Ubuntu across Python 3.12–3.14 on every push. Two pieces are macOS-only:
launchd/*.plist— the scheduler/MLX-supervisor/usage-poller are packaged as launchd jobs on macOS. On Linux, render the systemd equivalent withscripts/generate_systemd_units.sh(see Scheduler below) instead of hand-rolling init files, or run the entry points directly in a foreground terminal/tmuxsession.MLX (
PIPELINE_LOCAL_PROVIDER=mlx) — Apple Silicon only. Local dispatch works fine on Linux via Ollama or LM Studio instead (PIPELINE_LOCAL_PROVIDER=ollama/lmstudio).
Windows is untested.
Related MCP server: Vibecoders MCP
Quickstart
One-line install
curl -fsSL https://raw.githubusercontent.com/motock/fagan/master/scripts/remote-install.sh | bashThis clones the repo to ~/.fagan (override the location with
FAGAN_INSTALL_DIR, and the source URL with FAGAN_REPO_URL) and runs
scripts/install.sh inside it -- equivalent to the manual clone-and-run
steps below, minus the typing. Re-running it later updates the existing
checkout (git pull --ff-only) instead of re-cloning.
Piping a remote script into bash means trusting whatever that URL serves
at fetch time. If you'd rather read it first:
curl -fsSL https://raw.githubusercontent.com/motock/fagan/master/scripts/remote-install.sh -o remote-install.sh
less remote-install.sh # or open it in an editor
bash remote-install.shEither way, cd into the install directory it reports (~/.fagan by
default); it has already done steps 1–3 below, so restart Claude Code (step 5). Prefer a manual clone? Use the
steps below instead.
This gets the MCP server registered and a first plan running end-to-end.
A first run needs no local model at all: with nothing configured, dispatch
and review fall back to the claude backend, which shells out to the Claude
Code CLI. That fallback is the starting configuration, not the intended one
— the cost split described above only happens once you deliberately route
the implementation role to a local model, which is why the shipped registry
ships no roles block of its own: see Provider selection & authorization
below for how to make that choice when you're ready.
# 1. Clone and install the Python environment
git clone https://github.com/motock/fagan.git
cd fagan
scripts/install.sh # creates .venv, installs requirements.txt
# 2. Register the MCP server with Claude Code (adjust the path to where you cloned it)
claude mcp add -s user pipeline "$(pwd)/.venv/bin/python3" "$(pwd)/app/pipeline_mcp_server.py"
# 3. Copy the persona subagents and decision policy into place
# (cp -n skips any file you already have — e.g. a customized code-reviewer.md —
# instead of silently overwriting it; diff before removing -n if you do want the update)
mkdir -p ~/.claude/agents
cp -n agents/*.md ~/.claude/agents/
cp -n overlord-policy.md ~/.claude/overlord-policy.md
# 4. (Optional) Install the global rules bundle for your agent CLIs
# scripts/install_global_rules.py --tools=claude,codex,opencode
# Opt-in: nothing is written unless --tools is passed. It writes the bundle into
# ~/.claude/CLAUDE.md, ~/.codex/AGENTS.md and ~/.config/opencode/AGENTS.md, copies
# the rule files into the sibling fagan-rules/ directory, and backs up an existing
# file as <name>.fagan-bak-<UTC timestamp>. Re-running refreshes only the fenced
# block between the fagan:begin and fagan:end markers.
# 5. Restart Claude Code (or start a new session) so it picks up the MCP serverscripts/install.sh creates the .venv, installs requirements.txt and
requirements-dashboard.txt (the dashboard's fastapi/uvicorn deps, installed
on every run; a --dev install uses requirements-dev.txt, which already
includes the dashboard deps), and reports on the tools the pipeline shells out
to — required: git, gh, and the claude CLI; optional: ollama and
docker — with graceful-degradation messaging, and is safe to re-run. It does not register the MCP server, set environment
variables, or install the persona subagents — steps 2–3 above cover those. With
nothing but the claude backend configured, ollama/docker being absent is
expected, not an error.
From a Claude Code session in the project you want the pipeline to work on:
Ask the
product-analystsubagent to turn a goal into epics/stories, or hand-write a plan per the schema.mcp__pipeline__save_plan(oringest_plan) with that plan and arepo_rootpointing at the target project — not this pipeline repo.mcp__pipeline__list_ready_storiesto see what's unblocked, thenmcp__pipeline__dispatch_storyto claim and start one.Watch progress with the dashboard:
scripts/dashboard.sh start, then openhttp://localhost:8000.For unattended operation, run the scheduler so ready stories advance without you calling
advance_pipelineby hand:.venv/bin/python3 -m pipeline.scheduler_daemon(foreground, or under launchd/systemd/tmux — see Scheduler below).
Start with PIPELINE_AUTONOMY=dry-run (plans and logs only, nothing is
dispatched or merged) until you've watched one plan run and trust the gates —
see Autonomy levels.
Only using the claude backend? The PIPELINE_LOCAL_* and
PIPELINE_BACKEND_*=ollama/lmstudio/mlx variables, and Ollama/MLX/LM Studio
setup, only matter if you opt a role into local-model dispatch — but provider
selection itself is still a required setup step (the shipped registry routes
nothing; see Provider selection & authorization below), and even the
claude path needs two credentials before the first dispatch: gh auth login
(the pipeline opens and merges PRs through the GitHub CLI) and the Claude Code
CLI's own login. See
Minimal configuration for the handful of
variables actually worth setting on day one, versus the ~100 that exist purely
for tuning.
Provider selection & authorization
Provider selection is a required setup step. The shipped
model_registry.json deliberately declares which models exist per provider
but ships no roles routing: this project decouples from any single
provider, so the operator chooses. There are two supported ways to select a
provider per role, checked in this order by resolve_role:
Plan role config — a plan's per-role
provider/modelbeats everything below.A
rolesblock in a registry file — the single source of truth for role routing; see below.PIPELINE_BACKEND_<ROLE>environment variables — consulted only when the registry has no entry for the role (the empty-state path, so a fresh clone still boots); e.g.PIPELINE_BACKEND_DISPATCH=ollamaopts the dispatch role into Ollama.The caller's own fallback — for dispatch/review this is the
claudebackend.
For an interactive alternative to editing registry JSON by hand, run the
picker: .venv/bin/python scripts/choose_providers.py. It walks through all
nine roles one at a time, showing each role's current provider/model and where
that setting came from, and lets you switch it by typing an option number —
each of the nine roles is configured independently, and every change is
validated against the registry before it is written. It is safe to re-run any
time: re-running just re-reads the current routing, and pressing Enter keeps a
role's existing setting.
The same two registry files work for both selection styles:
PIPELINE_MODEL_REGISTRY_PATHpoints the pipeline at any registry JSON you like.model_registry.local.json(repo root) is the convention for a personal registry: it is gitignored, so your per-role routing stays out of the repo. PointPIPELINE_MODEL_REGISTRY_PATHat it, or copy it overmodel_registry.jsonlocally if you prefer not to set the variable.
A roles block names a provider and a friendly model name per role; the
friendly name must exist under that provider's models in the same file, and
the concrete tag is resolved from there. A typo raises an error rather than
silently falling back.
Authorization matrix. Selecting a provider also selects which credentials
you must establish first — scripts/install_checks.py probes these and
reports unauthorized (remedy: a login, not an install) where it can:
Provider / tool | Credential needed | How to establish it |
| GitHub auth (the pipeline opens and merges PRs through |
|
| Claude Code CLI's own login |
|
any | An ollama.com account, signed into the local daemon |
|
| Per-vendor API keys | |
on-device ollama / lmstudio / mlx tag | Nothing extra | — |
On the :cloud rows: those calls are proxied through https://ollama.com by
the local ollama daemon, which sends its own credential — the pipeline sends
no credential of its own. :cloud tags are the only ollama tags that need
a sign-in; purely on-device tags need nothing beyond the daemon running.
Getting-started walkthrough
The walkthrough works with whatever dispatch provider you have configured —
PIPELINE_BACKEND_DISPATCH (set it explicitly, or add a roles block to a
local registry — the shipped registry routes nothing; see Provider selection
& authorization above). With claude configured, dispatch and review shell
out to the Claude Code CLI; with a local provider such as ollama configured,
they run on that local model instead.
Install — one command:
scripts/install.sh(see the quickstart above for what it does and does not do).Register the MCP server and personas — quickstart steps 2–3 above (
claude mcp add ...plus copyingagents/*.mdand the overlord policy), then restart Claude Code.Start the dashboard —
scripts/dashboard.sh start, then openhttp://localhost:8000and pick your target project in the workspace picker.Decompose a tiny goal — ask the
product-analystsubagent (or the dashboard's decompose action) to turn a one-liner goal into epics/stories, thenmcp__pipeline__save_planthe result with itsrepo_rootfield pointing at your target project — not this pipeline repo.Dispatch the first ready story —
mcp__pipeline__list_ready_stories, thenmcp__pipeline__dispatch_storyon the first one, and watch the story advance across the kanban board in the dashboard.Watch it merge — with
PIPELINE_AUTONOMY=gated(the default), a risk-lowstory that passes review merges unattended. Start withPIPELINE_AUTONOMY=dry-runfirst, per the quickstart advice above.Prefer the scripted path? —
.venv/bin/python scripts/smoke_getting_started.pyruns the same flow end-to-end without the dashboard, in a scratchPLAN_DIRthat never touches your real plans. The smoke is provider-neutral: it runs on your configured dispatch provider (PIPELINE_BACKEND_DISPATCH, defaultclaude) and announces the resolved provider, model and source up front, so you always know which backend it validated. Exit codes:0PASS (the story reachedtests_passed),1the resolved provider isclaudeand theclaudeCLI is missing, exit 2 means the configured provider is empty or unrecognised — a configuration error, not a refusal of a local provider —3the bounded poll timed out,4the story failed. Honest caveat: PASS depends on the configured model actually completing the story, so a failure on a weak local model reflects that model, not a broken pipeline.
For what can still go wrong, see Reliability & limitations.
Companion MCP server (overlord + acceptance-oracle only)
Not ready to adopt the whole orchestrator? pipeline/companion_server.py is a
second, smaller MCP server (pipeline-companion) exposing two ideas that
stand on their own without adopting the rest of the pipeline:
escalate_decision (the overlord decision path) and the acceptance-oracle
helpers classify_oracle_outcome / acceptance_digests. It imports the real pipeline.overlord and
pipeline.oracle_gate modules rather than duplicating them, so it stays in
sync with the main server. Add it alongside the main server as a second
mcpServers entry:
{
"mcpServers": {
"pipeline": {
"command": ".venv/bin/python3",
"args": ["app/pipeline_mcp_server.py"]
},
"pipeline-companion": {
"command": ".venv/bin/python3",
"args": ["-m", "pipeline.companion_server"]
}
}
}The adoptable specs this server exports live in docs/specs/:
OVERLORD_POLICY_SPEC.md (the overlord decision path),
ACCEPTANCE_ORACLE_PATTERN.md (the acceptance-oracle grading pattern), and
DOCKER_SANDBOX.md (the opt-in Docker sandboxing behavior).
Running standalone (dashboard + scheduler, no MCP server)
The dashboard exposes the same operations as the MCP tools — save/ingest a plan,
decompose a goal, dispatch a story, advance, review, approve merge — so the
pipeline can run without registering an MCP server at all. That parity lives at
the HTTP API, not in the UI: the dashboard UI directly surfaces chat (including
drafting a plan), browsing plans, stories, journals and logs, the workspace
picker, the worktree-patch review/apply flow, role configuration, and ingesting a
saved plan. Dispatch, advance, review and approve-merge have UI-less API routes
(/api/plans/{plan_name}/stories/{story_key}/dispatch and friends) available for
scripting, and for the standalone flow the scheduler is the intended driver:
draft and ingest a plan from the dashboard, then let the scheduler dispatch,
advance, review and merge ready stories on its own. The
supported path is one command:
scripts/standalone-setup.sh upup provisions a scratch data dir (default ~/pipeline-standalone), writes
the shared operator env file with absolute paths, starts the dashboard and the
scheduler through their existing helper scripts, and then refuses to report
success until GET /api/health answers with an empty config_mismatch and
the intended plan_dir. Main options: --data-dir DIR (default
~/pipeline-standalone), --target-repo DIR (default: a scratch repo under
the data dir), --port PORT (default 8001), --autonomy MODE (default
dry-run), plus --repo-root and --force. down stops both processes and
leaves the scratch data in place; status prints the resolved paths and both
processes' state.
Both long-running processes read the same operator env file:
scripts/dashboard.sh and scripts/scheduler.sh both source
.pipeline.env (gitignored; see .pipeline.env.example) first, then
.dashboard.env (gitignored; see .dashboard.env.example) second, so
existing dashboard-only installs keep their current last-write precedence —
.dashboard.env still works and simply overrides .pipeline.env where they
overlap.
Because the dashboard and the scheduler are separate processes, PLAN_DIR
must match between the two: the scheduler writes a config fingerprint to
<plan_dir>/.scheduler_health.json, and /api/health reports
config_mismatch listing the fields where the dashboard's resolved config
differs from that fingerprint. A non-empty config_mismatch means the UI and
the scheduler are working different plan stores — check that both were
started with the same PLAN_DIR (the standalone script writes one env file
for exactly this reason, and fails hard on a non-empty config_mismatch).
The normal prerequisites still apply in standalone mode: gh auth login for
the PR/merge path (the pipeline opens and merges PRs through the GitHub CLI),
and provider authorization for whichever backend is configured — see
Provider selection & authorization above.
Components at a glance
Piece | Location | Role |
Persona subagents |
| The SDLC roles agents play |
Decision policy |
| How the overlord decides |
Pipeline MCP server |
| All pipeline tools + orchestration; |
Backend seam |
| Per-role driver routing ( |
Local agent loop |
| Native-tool-calling write loop for local dispatch (subprocess) |
Monitoring dashboard |
| FastAPI status/lifecycle viewer; in standalone mode (see "Running standalone" below) it also drives save/ingest/dispatch/review/merge directly |
Install / deps |
| venv + dependency setup |
Tests |
|
|
Plans / manifests / logs |
| Plan, manifest, decisions, notifications |
Worktrees |
| Isolated per-story branches |
Issue tracker | Plane (external, optional) | Mirror of story state; skipped entirely when unconfigured (manifest is the source of truth) |

The dashboard's Comms view — ask what's blocked, draft a plan, or approve a merge, all routed through the same gated API the kanban board's own buttons call. More screenshots (the live kanban board and the workspace picker) are in docs/DEMO.md.
Architecture
┌───────────────────────────────────────────────────────────┐
│ Orchestrator loop (cron / /loop skill) │
│ advance_pipeline(plan) — one idempotent tick │
└───────────────────────────┬───────────────────────────────┘
│ ready stories (deps satisfied)
▼
┌───────────────┐ resolve backend + ┌───────────────────────────────┐
│ Plan/Manifest │ persona/model │ Dispatch │
│ (JSON, Plane) │──────────────────────►│ claude -p OR local loop │
└───────────────┘ │ (tech-lead plans for local → │
│ .agent_plan.md) │
└───────────────┬───────────────┘
▼
┌───────────────────────────────┐
│ Headless story agent, TDD- │
│ first, in an isolated git │
│ worktree │
└───────────────┬───────────────┘
local fail → escalate │ tests +
to claude (`auto`) │ acceptance oracle
▼
┌───────────────────────────────┐
│ code-reviewer: VERDICT, │
│ opens a PR │
└───────────────┬───────────────┘
▼
low → decide silently ┌───────────────────────────────┐
medium → decide, notify the user │ Overlord adjudicates risk │──► decisions log
high → park, wait for a human │ (blocked decisions, merge, │ (audit trail)
│ scope disputes) │
└───────────────┬───────────────┘
▼ approved
┌───────────────────────────────┐
│ Merge gate: rebase on master, │
│ force-push, poll CI, re-run │
│ the suite on the rebased │
│ branch │
└───────────────┬───────────────┘
▼
masterPersonas (~/.claude/agents/)
Each persona is a Claude Code subagent: a markdown file with YAML frontmatter
(name, description, model, and optionally memory: user) and a
system-prompt body. The pipeline reads the body and dispatches a headless agent
with it as the role.
memory: user injects the user-memory directory into the system prompt on
every Claude call — high-leverage context but expensive in tokens. The
reviewer personas (code-reviewer, security-engineer) deliberately omit
it: their job is a mechanical check (run tests, read diff, emit VERDICT),
the CLAUDE.md rules they need are in the persona body, and skipping the
~132 KB memory injection shaves ~30-40% off every review call's input tokens.
The dispatch and overlord personas keep it because they benefit from project
context and are lower-volume.
Persona | Default model | Responsibility |
| opus | Decompose a goal into epics/stories with acceptance criteria, dependencies, and per-story |
| opus | General system design, tech selection, API design (delegates mobile to |
| sonnet | Default TDD implementer for non-mobile work |
| opus | Threat modeling and security review (OWASP, Secure by Design) |
| sonnet | Build/CI, branch & worktree hygiene, releases |
| sonnet | Reviews a branch, emits a |
| haiku | Docs for externally visible changes |
| opus | The decision authority (see below) |
Existing mobile specialists (mobile-architect, mobile-engineer,
ux-mobile-principal, qa-test-engineer) are unchanged and used for mobile work.
To change a persona's behavior or default model, edit its .md file. The
frontmatter model: line is the fallback model when a story does not specify one.
The overlord and the decision policy
The overlord (~/.claude/agents/overlord.md) rules on the user's behalf when
a story agent is blocked, two personas disagree, or a gate needs adjudication. It
follows ~/.claude/overlord-policy.md (plus an optional per-repo
<repo>/.overlord-policy.md override).
Decision tiers:
Routine / reversible → decide silently (naming, internal structure, a library within the approved stack, refactors).
Notify-async (
risk: medium) → decide, proceed, flag the user (new dependency, schema change, additive API change).Park-and-ping (
risk: high) → do not act unattended; hold for human review and notify. Anything irreversible, security/auth, money, production config, or breaking changes. Held for a human in dry-run and gated; infullthe overlord adjudicates it and records the ruling for post-hoc audit.
The overlord returns a structured ruling (RULING / TIER / RISK /
RATIONALE / NOTIFY_USER) that is parsed and written to the plan's decisions
log as an audit record.
Reference
See REFERENCE.md for the full MCP tools reference, the plan/story JSON schema, per-role provider/model configuration, guided decomposition and TDD-split details, every PIPELINE_*/LOCAL_AGENT_* environment variable, the end-to-end workflow, safety controls, the usage gate, and development/testing instructions.
For a worked end-to-end example of the pipeline developing this repository itself — the install command, the real pull requests it produced, and an honest account of what it can't do yet — see docs/DEMO.md.
For how a release is cut, see docs/RELEASING.md.
Prerequisites
Python 3.10+ and the project venv. CI tests 3.12–3.14 on Ubuntu and macOS on every push; 3.10/3.11 aren't part of the CI matrix, so treat them as likely-fine but unverified.
git on PATH.
GitHub CLI (
gh).Claude Code CLI (
claude).
Scheduler
The advance-scheduler runs as a long-lived daemon rather than a periodic
launchd tick. launchd's role is limited to crash-restarting it via KeepAlive.
Environment Variables
PIPELINE_SCHEDULER_INTERVAL_S – default reconcile sweep interval (default 60 seconds).
PIPELINE_SCHEDULER_HEALTH_PATH – optional path where the daemon writes its health JSON each iteration.
Rendering the launchd files for your machine
The committed launchd/*.plist files and launchd/pipeline-logs.newsyslog.conf
are a reference copy: they carry the maintainer's own absolute paths (a
/Users/<name>/... home directory, a specific model cache path) and will not
work unedited on another machine. On a fresh install, regenerate them yourself
with scripts/generate_launchd_plists.sh (install.sh does not run this for
you) — it fills the templates in launchd/
(launchd/com.fagan.pipeline.*.plist.template) from three flags:
--repo-root— the pipeline checkout the rendered files should point at (default: the repo that contains the script).--out-dir— where the rendered files are written (default:<repo-root>/launchd).--mlx-model-path— the local MLX model directory baked into the mlx-supervisor plist. As an alternative to the flag you can set theMLX_MODEL_PATHenvironment variable; the flag wins when both are given. The script fails closed — it exits with an error — when neither is supplied.
The same script also renders launchd/pipeline-logs.newsyslog.conf from
launchd/pipeline-logs.newsyslog.conf.template, substituting only the repo root.
scripts/generate_launchd_plists.sh \
--repo-root "$HOME/.claude/mcp-servers/pipeline" \
--out-dir "$HOME/.claude/mcp-servers/pipeline/launchd" \
--mlx-model-path "$HOME/.cache/qwen2.5_coder_14b_manual"These launchd files are macOS-only - see Platform support.
Rendering the systemd units for Linux
scripts/generate_systemd_units.sh renders the equivalent systemd user-unit
and logrotate files from systemd/*.template, the same way
scripts/generate_launchd_plists.sh does for launchd – minus MLX, which is
Apple Silicon-only:
scripts/generate_systemd_units.sh \
--repo-root "$HOME/fagan" \
--out-dir "$HOME/fagan/systemd"Install as per-user systemd units (no root required):
mkdir -p ~/.config/systemd/user
cp systemd/com.fagan.pipeline.advance-scheduler.service ~/.config/systemd/user/
cp systemd/com.fagan.pipeline.usage-poller.service ~/.config/systemd/user/
cp systemd/com.fagan.pipeline.usage-poller.timer ~/.config/systemd/user/
systemctl --user daemon-reload
systemctl --user enable --now com.fagan.pipeline.advance-scheduler.service
systemctl --user enable --now com.fagan.pipeline.usage-poller.timer
# Optional: let these run even when you are not logged in
loginctl enable-linger "$USER"Log rotation (needs root, one-time):
sudo cp systemd/pipeline-logs.logrotate.conf /etc/logrotate.d/com.fagan.pipelineReliability & limitations
This pipeline runs real autonomous coding loops, and they fail in specific, documented ways — read this before pointing it at anything you care about.
Local (non-Claude) model dispatch is the weak point. It works well for small, mechanically-scoped stories (one concern, ≤2 production files) and degrades sharply on anything bigger: large-file edits, multi-function stories, and anchored inserts into long existing functions reliably cause step-cap timeouts, stalls, or file corruption from stale line-number edits.
docs/plans/*.mdandretros/*.mdin this repo are the actual incident record this finding comes from, not a marketing claim — read a few before trusting local dispatch on anything non-trivial.PIPELINE_BACKEND_DISPATCH=autoexists specifically to escalate a struggling local attempt to Claude rather than let it loop.The "$20/month" framing is the design goal the gates are built around, not a benchmarked result yet. The one full model-comparison run on record (
tests/benchmark/FINDINGS.md) was contaminated mid-run by rate limits and credit exhaustion, so there is no clean apples-to-apples success-rate/cost comparison across backends published yet. The cleanest number there is narrow —gpt-oss:20bon-device, 2 T1 tasks, 2/2 success with the independent oracle passing on the merged code, one trial each — and is directional, not a quality comparison. Read that file for exactly what is and isn't known before citing a number from it.A green test suite is not proof of a correct or complete change. An executor (local or Claude) converges to the minimum diff that turns its own tests green, and can write a self-consistently wrong test that encodes the same bug as its implementation. See
.claude/rules/code-review.md's "Merge-gate and AI-review lessons" section — every lesson there came from a real merged regression, not a hypothetical.A story marked
doneis not proof its title's full scope shipped. A "migrate everything" or "remove all X" story can pass review and merge having only done part of the job, because review grades the story's own tests, not the title's claim. See.claude/rules/agent-dispatch-story-sizing.md.The overlord's
park-and-pingtier is a real safety floor, not a suggestion — high-risk decisions (irreversible actions, auth/security, money, production config, breaking changes) stop for a human indry-runandgated; infullthe overlord adjudicates the high-risk merge hold and records the ruling for post-hoc audit. Start any new deployment atPIPELINE_AUTONOMY=dry-runand read the decisions log before trustinggatedorfull.This is a single-maintainer research project, not a maintained product with an SLA. The test suite and CI are real gates, but expect rough edges, and expect the failure-mode catalog to keep growing as new ones are found.
If you hit a new failure mode, it's worth documenting (see retros/ for the
existing format) rather than working around it silently — the whole value of
this project's design is that failure modes get named and fed back into how
stories are sized and reviewed.
License
Licensed under the Apache License, Version 2.0 — see LICENSE and NOTICE.
Available Tools
26 toolsadvance_all_plansA
Run advance_pipeline on every plan that has been ingested (has a manifest), keyed by plan name. Plans saved but not yet ingested (no manifest) are skipped. Intended for a recurring scheduler (cron/launchd or /loop) so newly ingested plans are picked up automatically with no hardcoded plan name to maintain.
NOTE on zombie reaping: the per-plan advance_pipeline polling phase already handles dead-pid in_progress stories via check_story_status (which falls through to test-running on dead pids). Running an external reap pass BEFORE the polling would clobber that and silently leave stories re-dispatching forever without ever running the test (manifest observation 2026-06-28: 3 e2e stories hit dispatch_attempts= MISSING because the reap ate the polling opportunity). The reap helper _reap_zombie_in_progress_stories is kept for callers that need a one-shot cleanup (e.g. tests, ops CLI) but is NOT wired in here.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does so thoroughly: it explains skipping non-ingested plans, how dead-pid stories are handled via check_story_status, and why an external reap pass is deliberately not wired in. This is detailed, non-obvious behavior that an agent would not otherwise know.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded and the skip condition is immediately clear. The zombie-reaping note is valuable but contains more incident detail than an agent needs for selection, so the definition is slightly verbose rather than perfectly concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter batch tool with an output schema, the description covers what it does, when to use it, which plans are included, and important behavioral caveats. Nothing essential to calling it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, so the schema already exhaustively covers the input contract. The description adds relevant context by explaining that no plan name needs to be passed because the tool iterates all ingested plans.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Run advance_pipeline') and a precise resource scope ('every plan that has been ingested'), and distinguishes itself from the per-plan sibling advance_pipeline by emphasizing 'every plan' and the manifest condition. It also clarifies what is excluded (plans without a manifest).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear intended use case: a recurring scheduler that picks up newly ingested plans automatically, and notes the no-hardcoded-plan-name benefit. It does not explicitly name the alternative for single-plan advancement, though 'every plan' strongly implies the boundary with advance_pipeline.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
advance_pipelineA
Run one orchestration tick: dispatch every ready story (deps satisfied), advance finished stories through test -> review -> PR, and adjudicate merges against the risk threshold. Idempotent; designed to be called repeatedly by a scheduler (/loop or cron). In PIPELINE_AUTONOMY=dry-run it plans and logs only, taking no actions.
Honors a per-backend resource gate: dispatch and review are gated independently by their own backend's resource_status() (see _role_resource_ok). If the dispatch backend is gated, in-progress stories are interrupted (checkpointed, resumable) and no new dispatch starts; if the review backend is gated, review is deferred. Each is independent, so a Claude usage pause no longer freezes local-backed dispatch. Merge adjudication always runs (no model usage). "interrupted" stories are dispatch-eligible like "todo" ones, so they resume automatically once the dispatch backend frees up.
Also honors MAX_CONCURRENT_AGENTS: dispatch is capped to the number of free slots remaining (limit minus agents already in_progress across all plans), so a tick never starts more agents than the configured ceiling. Stories left undispatched this tick stay "todo"/"interrupted" and are picked up on a later tick as slots free up.
Skips entirely (returns {"ok": True, "skipped": "locked"}) if another tick for this same plan is already running - see _plan_lock.
| Name | Required | Description | Default |
|---|---|---|---|
| plan_name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries full responsibility and does so thoroughly. It discloses idempotency, dry-run behavior, per-backend gating with independent dispatch/review gates, interruption/resume semantics, MAX_CONCURRENT_AGENTS cap, and the locked-skip response. This is exceptional transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence adds essential behavioral context. It is front-loaded with the core purpose, then progressively details gating, concurrency, and locking. There is no filler or redundancy; each clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex orchestration tool, the description covers all critical aspects: what actions are taken, idempotency, gating logic, concurrency limits, and lock behavior. The existence of an output schema covers return-value details, so nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter, plan_name, has zero schema description. The description doesn't explicitly define it, but the tool name and the phrase 'Run one orchestration tick' make it obvious that plan_name identifies the plan to advance. Given the single-parameter simplicity, this is acceptable; a baseline 4 applies due to low schema coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a precise verb and resource: 'Run one orchestration tick' and then enumerates the specific actions (dispatch ready stories, advance through test->review->PR, adjudicate merges). It clearly differentiates from siblings like advance_all_plans (one plan vs all) and dispatch_story (single story vs orchestrated batch).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says it is idempotent and designed for repeated scheduler calls, and describes the skip-on-lock behavior. It doesn't explicitly name alternatives (e.g., 'use advance_all_plans for all plans'), but the one-plan scope is clear from the purpose and the lock behavior.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
approve_mergeA
Merge a reviewed story's PR into the default branch, right now, on the caller's explicit approval — this is the human/overlord merge decision itself, not a status check.
plan_name: the plan's name, as returned by list_plans or passed to save_plan/ingest_plan. story_key: the story's key within that plan's manifest. Must currently be "parked" or "pr_open" with review_verdict == "APPROVE" — any other state (not yet reviewed, still in progress, already merged) returns {"ok": False, "error": ...} without changing anything.
On success this: rebases the story's branch onto the current default branch, force-pushes it (--force-with-lease) to origin, polls real CI (gh pr checks) — auto-retrying once on a cancelled run, and failing closed on a fail/cancelled/still-pending result — re-runs the story's acceptance fixtures and a build check against the rebased code, then merges the PR, deletes the branch/worktree, marks the story "done" in the manifest and its ticket, and notifies the user to restart the MCP server if the story touched the pipeline's own source. Any failure at any of those steps aborts the merge and returns the specific reason instead of partially completing it.
This does more than a plain gh pr merge (which skips the rebase,
force-push, and re-verification) — prefer this tool over a manual
merge for exactly that reason. It force-pushes and merges regardless
of any local test run you've done yourself, so only call it once you
actually want this specific story merged now; there is no separate
confirmation step after this call.
| Name | Required | Description | Default |
|---|---|---|---|
| plan_name | Yes | ||
| story_key | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure, and it does so thoroughly: it enumerates the entire merge pipeline (rebase, force-push with --force-with-lease, CI polling with retry, acceptance re-run, merge, branch/worktree cleanup, manifest/ticket update, server restart notice), emphasizes atomic abort-on-failure, and warns that force-push and merge happen regardless of local test runs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but every sentence earns its place: the core purpose is front-loaded, parameter semantics follow, then the detailed behavioral pipeline, then usage cautions. There is no filler or repetition; the length is justified by the destructive and irreversible nature of the operation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity, the absence of annotations, and the presence of an output schema, the description is complete: it covers preconditions, side effects, failure semantics, alternatives, and success/error behavior. Nothing an agent needs to safely decide to call this tool is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, so the description must compensate, and it does: plan_name is defined as returned by list_plans or passed to save_plan/ingest_plan, and story_key is defined as the story's key within that plan's manifest with the exact state constraints ('parked' or 'pr_open' with 'APPROVE'). Both parameters receive clear, operational meaning beyond their bare schema types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Merge a reviewed story's PR into the default branch, right now, on the caller's explicit approval.' It also clarifies what the tool is not ('not a status check') and distinguishes itself from a plain `gh pr merge`, so an agent can tell exactly what action it performs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use the tool ('only call it once you actually want this specific story merged now'), states the required precondition state ('parked' or 'pr_open' with review_verdict == 'APPROVE'), and names the alternative ('plain `gh pr merge`') while explaining why this tool is preferred. It also warns there is no separate confirmation step.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
checkpointA
Deprecated alias for checkpoint_story — identical behavior and parameters (see checkpoint_story's docstring for the full description of plan_name/story_key/step/summary/next_hint). Kept only for backward compatibility with agents/prompts still calling the old name; prefer checkpoint_story in new code.
| Name | Required | Description | Default |
|---|---|---|---|
| step | Yes | ||
| summary | Yes | ||
| next_hint | No | ||
| plan_name | Yes | ||
| story_key | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, but the description declares the behavior as identical to checkpoint_story and directs the agent to that tool's docstring for full details. This is sufficient for an alias, though it does not independently describe the behavior here.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Extremely concise, front-loaded with the deprecation status, and every sentence earns its place. It conveys the essential information (alias, identical behavior, preference for the canonical tool) with zero waste.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a deprecated alias, the description is complete: it tells the agent that the tool is a backward-compatible shim, points to the full documentation, and provides the parameter names. Nothing more is needed for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
While schema coverage is 0%, the description explicitly lists the parameter names (plan_name/story_key/step/summary/next_hint) and points to the canonical docstring for full semantics. This compensates for the missing inline documentation by directing the agent to the authoritative source.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as a deprecated alias for checkpoint_story, stating it has identical behavior and parameters. This is specific and distinguishes it from the canonical sibling, making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use this tool (backward compatibility only) and when not to (prefer checkpoint_story in new code). It also names the alternative directly, leaving no ambiguity about routing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
checkpoint_storyA
Record a durable checkpoint for a dispatched agent's own progress on a story. Commits any uncommitted work in the story's worktree as a WIP commit and appends an entry to the story's journal (..journal.json).
plan_name: the plan's name this story belongs to. story_key: the story's key within that plan's manifest — the same key this agent was dispatched with. step: a short label identifying this step (e.g. "wrote-failing-test", "implemented-fix") — becomes part of the WIP commit message, so keep it terse and distinct from other steps in this story. summary: a sentence describing what was actually done in this step, for whoever (human or resumed agent) reads the journal later. next_hint: optional — what to do next if this run is interrupted right after this checkpoint. Leave empty if there's nothing beyond "continue the story normally."
Call this after completing each idempotent step of a story (not mid- step) so a killed or interrupted agent resumes from the last checkpoint via dispatch_story instead of starting the story over from scratch.
| Name | Required | Description | Default |
|---|---|---|---|
| step | Yes | ||
| summary | Yes | ||
| next_hint | No | ||
| plan_name | Yes | ||
| story_key | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden and mostly meets it: it discloses that the tool commits uncommitted work as a WIP commit and appends to the journal. It does not cover failure behavior or permission requirements, but the principal side effects are explicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose, then moves through mechanics, parameters, and usage timing. Every sentence adds decision-relevant detail, including the journal path and step-label constraints, with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a five-parameter mutation tool with no annotations, the description covers what it does, how it mutates state, what each parameter means, and when to call it. The output schema exists for return-value details, so no required context appears to be missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description defines all five parameters in detail, including the journal filename pattern, example step labels, WIP-commit naming impact, and the optional nature of next_hint. This fully compensates for the empty schema descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific action (record a durable checkpoint), a scoped resource (a dispatched agent's own progress on a story), and the concrete mechanism (WIP commit + journal append). The 'dispatched agent's own progress' framing distinguishes it from generic checkpoint/story-state siblings like checkpoint or set_story_status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit timing: call after each idempotent step, not mid-step. Also explains the recovery flow via dispatch_story, making it clear when the checkpoint matters and how it ties to the alternative path.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_story_statusA
Check whether a dispatched agent has finished. If complete, runs tests in the worktree and reports pass/fail without auto-merging.
| Name | Required | Description | Default |
|---|---|---|---|
| plan_name | Yes | ||
| story_key | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It states that it runs tests in the worktree, reports pass/fail, and explicitly notes it does not auto-merge. This is significant behavioral information. However, it does not disclose what happens if the agent hasn't finished, or any other side effects, leaving some gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, concise and front-loaded with the primary purpose. It avoids filler and directly states the key behavior. Every sentence adds value, and the structure is optimal for quick scanning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is relatively simple with two parameters and no output schema, but the description leaves gaps: it does not explain the parameters, does not mention what happens when the agent is not complete, and does not differentiate from sibling status tools. It covers the core behavior but lacks contextual details that would help an agent use it correctly in a workflow.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for the lack of parameter documentation. It does not explain what 'plan_name' or 'story_key' refer to, nor how they are used. The agent is left to infer from names alone, which is insufficient given the zero coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the verb 'check' and the resource 'whether a dispatched agent has finished', and adds specific behavior (runs tests, reports pass/fail, no auto-merge). This distinguishes it from sibling tools like set_story_status or mark_story_done, which focus on status changes rather than checking and testing.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use the tool: to check if a dispatched agent has finished. It implies the intended workflow (post-dispatch status check) but does not explicitly mention when not to use it or name alternatives. Since it gives clear usage context without exclusions, it earns a 4.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_usageA
Probe current subscription usage (current session + current week) via a headless /cost call and persist it to USAGE_STATE_PATH.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full disclosure burden, and it covers both the mechanism ('headless /cost call') and the side effect ('persist it to USAGE_STATE_PATH'). It could add more about failure behavior or prerequisites, but the core behavioral traits are disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, front-loaded with the primary purpose, and every clause earns its place by adding scope, mechanism, or persistence detail. No redundant text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter probe with an output schema, the description covers purpose, scope, mechanism, and persistence. The main gap is the absence of any note about when not to call it or what the persisted state is used for, but this is minor given the low complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so the description isn't required to add parameter-level meaning; the baseline for zero-param tools is 4. The description's mention of scope is not parameter documentation but it reinforces what the tool measures.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific action ('Probe'), a concrete resource ('current subscription usage'), and a precise scope ('current session + current week'). This distinguishes it from the unrelated sibling tools and leaves no doubt about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the invocation context clear: an agent calls this tool when it needs current subscription usage, scoped to session and week. It doesn't name explicit exclusions or alternatives, but no sibling tool overlaps with this usage-checking function.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
decompose_planA
Turn a raw goal/feature request into epics/stories JSON via the product-analyst persona, run on whichever provider the "decompose" role is configured for (PIPELINE_BACKEND_DECOMPOSE env var, or a "decompose" entry in model_registry.json - defaults to Claude when neither is set). This is a separate, additional path from the interactive product-analyst subagent (invoked via the Agent tool, which is always Claude) - that path remains available and is still the default choice for Claude-quality decomposition; this tool exists so decomposition can also run on a local provider when desired.
Does NOT call save_plan itself - review the returned plan the same way you would review the interactive subagent's output, then save_plan it yourself.
Returns {"ok": True, "plan": {...}} on success. On failure, returns {"ok": False, "error": ...}, with "raw": included whenever the backend actually returned text that failed to parse (never raises).
| Name | Required | Description | Default |
|---|---|---|---|
| request | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden and does so thoroughly. It discloses the success and failure return contracts, includes raw output on parse failures, states that it never raises, and clarifies that it does not persist the plan via save_plan.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than average, but every paragraph serves a distinct decision point: core action, provider configuration, relationship to the alternative path, save_plan responsibility, and return contract. The main purpose is front-loaded in the first sentence, and there is no obvious filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given one parameter, no annotations, and the complexity of provider selection and output handling, the description is complete. It tells the agent exactly what to pass, what provider will be used, what the return shape is, what failure looks like, and what the agent must do after calling it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, so the description must define the sole request parameter. It does this by calling it a 'raw goal/feature request', which is sufficient for a single free-form string parameter, though it does not give examples or length/format constraints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb-resource pair: 'Turn a raw goal/feature request into epics/stories JSON'. It further distinguishes itself from the interactive product-analyst subagent and clarifies that it is a separate decomposition path, so an agent can tell it apart from siblings like save_plan and the Agent tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly explains when to use this tool versus the interactive product-analyst subagent: the interactive path is still the default for Claude-quality decomposition, while this tool is for running decomposition on a local provider when desired. It also gives a clear behavioral instruction: do not expect it to call save_plan; review the output and save it yourself.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
dispatch_storyA
Spawn a headless Claude Code agent to work on a single story.
For a fresh story, creates a git worktree on a new branch. For a story left "interrupted" (or whose worktree already exists from a prior run), reuses the existing worktree/branch instead and seeds the agent's prompt with the checkpoint journal so it continues rather than starting over. Transitions the Plane issue to In Progress. Returns the subprocess PID; completion is async.
Acquires _plan_lock so direct MCP tool calls serialize across MCP
server processes - without this guard, two Claude sessions (each with
their own MCP server PID) can both call dispatch_story on the same story
in the same window, and the second one treats the first one's
half-built worktree as resumable and spawns a second agent into the
same directory. That race is what produced the repeated zero-output
agent deaths logged in 2026-06-27's e2e-decentralized-messaging run.
| Name | Required | Description | Default |
|---|---|---|---|
| plan_name | Yes | ||
| story_key | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It covers the core action (spawn agent), the worktree lifecycle (create vs. reuse), the transition to In Progress, the async completion and PID return, and a detailed explanation of the _plan_lock race condition including a concrete failure example. This is exceptionally transparent and goes beyond typical descriptions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose, then explains the fresh/interrupted branching, and finally a lengthy race-condition explanation. The race-condition paragraph is detailed but arguably excessive for an agent deciding whether to call the tool; it could be condensed to a caution about locking. Overall it is well-structured but slightly verbose, so a middle score is appropriate.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations and no output schema shown, the description covers the essential behavioral aspects: what it does, how it handles different states, the async nature, the returned PID, and the concurrency risk. The only major omission is the meaning of the two parameters, which is a notable gap. Still, for a tool of moderate complexity, it is largely complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate by explaining the parameters. It never explicitly defines plan_name or story_key. While the names are somewhat self-explanatory (plan identifier and story key), the description does not add any detail about their format, allowed values, or relationship. For a tool with two required parameters and zero schema coverage, this is a significant gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a clear verb-resource statement: 'Spawn a headless Claude Code agent to work on a single story.' It then distinguishes fresh versus interrupted stories and explains worktree reuse, which differentiates it from sibling tools like mark_story_in_progress or interrupt_story. The purpose is unambiguous and specific.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly explains when to use the tool (for a fresh or interrupted story) and what happens in each case. It does not explicitly name alternatives or say 'use this instead of X,' but the behavioral context (spawning an agent, worktree management) makes the use case evident. It also cautions about the lock race condition, which is relevant to safe usage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_effective_configA
Read-only diagnostic snapshot of the pipeline's effective configuration: every role in config_provenance.PIPELINE_ROLES with its resolved (provider, model) and provenance, every cataloged env var's resolved value and provenance, any unrecognized/ignored env vars present, and which config-source files were actually consulted (and whether each exists). Pure read - makes no changes and writes nothing.
"restart_required" on a role/env entry means that entry's winning value came from an env var or the launchd plist/mcp_server_env layer, so a change there only takes effect after the scheduler/MCP server is restarted. By contrast, a plan's role_config and model_registry.json are both read fresh on every call, so edits to either are live immediately with no restart needed.
A role entry carrying a non-None "error" key is misconfigured (e.g. no model configured for it anywhere, or its provider/model pairing isn't declared in model_registry.json) - never raises for this; the bad role just reports its error inline while the rest of the roles resolve normally.
Pass plan_name to additionally layer in that plan's role_config overrides (same effect as get_effective_config's plan_name); a plan_name whose manifest doesn't exist degrades to "no plan overrides" rather than raising.
| Name | Required | Description | Default |
|---|---|---|---|
| plan_name | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and handles it excellently. It explicitly states 'Pure read - makes no changes and writes nothing', explains the meaning and restart implications of 'restart_required', describes how misconfigured roles surface via an inline 'error' key without raising, and documents the degraded behavior for missing plan manifests. This is far beyond typical behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than average, but nearly every sentence carries substantive behavioral or parameter semantics. The opening sentence front-loads the core purpose and read-only guarantee. Minor redundancy exists ('Read-only diagnostic snapshot' vs 'Pure read - makes no changes and writes nothing') and the self-referential 'same effect as get_effective_config's plan_name' could be tightened, but overall it is well-organized into scannable paragraphs.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity and the presence of an output schema, the description covers everything needed to call it correctly: what it returns, how the single optional parameter behaves, error-handling behavior, and restart semantics. The output schema handles return-value structure, so no further description is needed there.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must explain plan_name, and it does so thoroughly: optional, layers role_config overrides, missing manifest degrades to no overrides. It also clarifies the semantic distinction between values that require a restart and those that are live immediately, which adds meaning far beyond the raw schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource ('pipeline's effective configuration') and a specific diagnostic purpose ('Read-only diagnostic snapshot'), then enumerates exactly what it reports: resolved role configs, env vars, provenance, and consulted files. This clearly differentiates it from the sibling tools, which are all mutation or status operations on stories/plans.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit context for the optional plan_name parameter: pass it to layer in role_config overrides, and a missing manifest degrades gracefully instead of raising. It does not explicitly name alternatives or when-not-to-use, but the tool is unique among siblings and its read-only diagnostic role is clear. The phrase 'same effect as get_effective_config's plan_name' is slightly self-referential and confusing, but does not undermine the guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
ingest_planA
Push a saved plan into Plane. Creates epics first, then issues linked to their parent epic. Optionally restrict to specific epic summaries via only_epics. Returns a manifest mapping local IDs to Plane UUIDs.
Re-ingesting an already-ingested plan merges into the existing manifest rather than replacing it: epics/stories not touched this call (including everything only_epics excludes) are preserved verbatim, a story whose key already exists gets its authored fields (summary, agent_instructions, dependencies, persona, model, acceptance, risk (only while todo), backend, tdd_split, files) refreshed while its runtime state (status, pr_url, escalated, ...) is kept, and top-level manifest keys outside epics/stories/repo_root (paused, local_model_fallback, final_rework_escalation, ...) carry over untouched. Pass overwrite=True to restore the old wholesale-replace behavior (drops anything not produced by this call).
| Name | Required | Description | Default |
|---|---|---|---|
| overwrite | No | ||
| plan_name | Yes | ||
| only_epics | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden and does so in detail: it explains the creation order, the optional epic restriction, the local-to-UUID manifest return, and the critical merge-vs-replace semantics on re-ingest. It also discloses the destructive effect of overwrite=True (drops anything not produced by this call), which is exactly the kind of side-effect information an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core action and return, then moves into necessary merge semantics. It is dense and somewhat long, but the complexity of the operation justifies most of the length. Minor parenthetical-heavy passages could be tightened, but no sentence is wholly wasteful.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of annotations, the presence of an output schema, and the tool's mutation complexity, the description is largely complete: it covers side effects, merge behavior, and return format. The main remaining gap is the missing semantic definition of plan_name and the lack of explicit routing guidance against sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for all parameters. It clearly explains only_epics as restricting to specific epic summaries and overwrite=True as restoring wholesale-replace behavior. However, the required plan_name parameter is never explicitly defined (e.g., whether it is a saved plan name, path, or identifier), leaving one of three parameters only implicitly covered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Push a saved plan into Plane.' It also describes the resulting actions (creates epics, then linked issues) and the return manifest, giving the agent a precise sense of what the tool does distinct from siblings like save_plan or list_plans.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Implies usage when a plan has been saved and should be pushed into Plane, and it clarifies when to use only_epics and overwrite. However, it does not explicitly name alternatives (save_plan, decompose_plan, patch_plan) or state when not to use this tool, leaving the agent to infer routing from sibling names.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
interrupt_storyA
Stop a dispatched agent and leave its story resumable.
Sends SIGTERM to the agent's process (a no-op if it has already exited), commits any uncommitted work in its worktree as a checkpoint, and marks the story "interrupted" rather than "failed" so a later dispatch_story call resumes it instead of starting over. The worktree and branch are left in place.
Acquires _plan_lock for the same reason dispatch_story does - two MCP
servers can race here too, with one calling interrupt while the other
calls dispatch on the same story, producing a manifest write race that
leaves the worktree in an inconsistent state.
| Name | Required | Description | Default |
|---|---|---|---|
| plan_name | Yes | ||
| story_key | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully discloses behavior: sends SIGTERM (no-op if exited), commits uncommitted work as a checkpoint, marks story 'interrupted' rather than 'failed', leaves worktree and branch in place, and acquires _plan_lock due to a race condition. This is exceptionally transparent about side effects and concurrency concerns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured: the main action is front-loaded, followed by detailed mechanics and a separate paragraph on locking. Each sentence adds value, though the lock explanation is a bit verbose. It is appropriately sized for a tool with this complexity and avoids redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the core behavior, side effects, and concurrency rationale, and the existence of an output schema likely handles return values. However, it does not explain the parameters or state prerequisites (e.g., the story must be currently dispatched). Given the detailed nature, it is mostly complete but leaves parameter semantics unresolved.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate by explaining plan_name and story_key. It mentions 'plan' and 'story' in the text but never explicitly defines what these parameters represent or how they are used. An agent would have to infer their meaning from the tool name and context, which is insufficient for unambiguous invocation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Stop a dispatched agent and leave its story resumable.' It details the mechanism (SIGTERM, commit, marking interrupted) which clearly distinguishes it from siblings like mark_story_done (which would finalize) or dispatch_story (which starts). The resumability is a unique, well-articulated trait.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly implies when to use the tool (to stop a running agent while preserving progress) and mentions that a later dispatch_story call resumes it. However, it does not explicitly state when not to use it (e.g., for permanent termination) or compare against other status-changing siblings like mark_story_done. The context is clear but lacks explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_decisionsA
Return the overlord's decision log for a plan: every ruling ever made by request_decision on this plan, oldest first, as an audit trail. Read-only — makes no changes to the plan or any story.
plan_name: the plan's name, as returned by list_plans or passed to save_plan/ingest_plan.
Each entry corresponds one-to-one with a prior request_decision call and carries at least the story_key, question, the ruling made, and a timestamp. Returns an empty list if the plan has no decisions logged yet — this is normal for a plan with no blocked stories, not an error.
Call this to check for precedent before escalating a similar decision with request_decision, or when a human wants to review what the overlord has ruled on so far for a plan.
| Name | Required | Description | Default |
|---|---|---|---|
| plan_name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does it well: it states the tool is read-only and makes no changes, explains the one-to-one correspondence with request_decision calls, lists the fields each entry carries, and clarifies that an empty list is normal rather than an error.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose, then adds parameter semantics, return behavior, and usage guidance in a logical order. Every paragraph earns its place without redundant wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the low complexity, one parameter, and presence of an output schema, the description covers everything an agent needs: what the log contains, how to read it, when to use it, and what empty results mean. Nothing essential is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, but the description compensates by defining plan_name as the plan's name as returned by list_plans or passed to save_plan/ingest_plan. This gives the agent concrete sources for valid values despite the sparse schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: return the decision log (audit trail) for a plan, with every ruling made by request_decision, oldest first. It clearly differentiates itself from sibling tools like list_plans and request_decision by focusing on historical rulings rather than plans or new decisions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly says to call this before escalating a similar decision with request_decision, or when a human wants to review the overlord's rulings. It does not list exclusions or when-not-to-use conditions, but the intended contexts are clearly described.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_plansA
List plan names found in the plan directory (~/.claude/plans, or $PLAN_DIR). No parameters.
Returns the '.json' stem of every file in that directory, not just fresh, ingestable plan sources — an already-ingested plan's '.manifest.json' surfaces as '.manifest', and a story journal or notification log contributes its own noisy stem too. Treat a returned name as a candidate to inspect, not a guarantee it is a valid target for save_plan/ingest_plan.
Call this to check whether a plan name is already taken before save_plan, or to confirm a plan file actually landed on disk after saving it.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It candidly discloses that results are noisy, that manifest files surface as '.manifest', that story journal/notification logs contribute stems, and that returned names are only candidates, not guaranteed ingestable plans. This is rich and honest behavioral detail.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than minimal, but every sentence adds distinct value: directory location, no parameters, noisy-stem behavior, candidate semantics, and concrete use cases. It is structured with a clear purpose, caveat, and usage; only mild redundancy about 'not a guarantee' could be trimmed.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given zero params and the presence of an output schema, the description covers everything needed to call and interpret results: directory source, return stem semantics, noise sources, validity caveat, and practical call scenarios. Nothing important is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters and the description explicitly says 'No parameters.' With no params, there is nothing to document; the baseline 4 applies because the description accurately confirms the schema's empty parameter set.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'List plan names found in the plan directory'. It distinguishes itself from fresh/ingestable plan sources by explaining the directory scan includes noisy stems, which differentiates it from sibling tools like list_ready_stories.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly tells when to call the tool: to check if a plan name is taken before save_plan, or to confirm a plan file landed on disk. It does not name an alternative tool to use instead, but the guidance is clear and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_ready_storiesA
Return the stories in this plan that are unblocked and available to dispatch: status == "todo" and every entry in the story's own dependencies list refers to an already-completed story.
plan_name: the plan's name, as returned by list_plans or passed to save_plan/ingest_plan. Returns [] if the plan has no manifest yet (not yet ingested) rather than raising.
Returns a list of {"key": ..., "summary": ...} — just enough to choose a story key for dispatch_story or check_story_status, not the full story record (agent_instructions, acceptance, etc. are omitted; use check_story_status for those). A story already in_progress, parked, or done is never included, and neither is a "todo" story whose dependencies aren't all done yet — call this again after a dependency completes rather than assuming today's list stays valid.
| Name | Required | Description | Default |
|---|---|---|---|
| plan_name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it delivers: it discloses non-raising behavior for missing manifests, the intentionally limited return shape, exclusion of in_progress/parked/done stories, and the dynamic/volatile nature of the ready list. This goes well beyond what the name alone conveys.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is informative and front-loaded with the core behavior. It slightly repeats the dependency condition in the final paragraph ('neither is a todo story whose dependencies aren't all done yet') after already stating it in the first sentence, but overall every sentence adds meaningful operational detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the exact readiness condition, return format, omitted fields, error/missing-manifest behavior, parameter provenance, and temporal validity. With no annotations and only one simple parameter, this is sufficient for an agent to call the tool correctly without needing to infer anything.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and there is only one parameter, so the description must define plan_name. It does so meaningfully by explaining it is 'the plan's name, as returned by list_plans or passed to save_plan/ingest_plan,' which gives the agent a concrete source and provenance for the value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Return the stories in this plan that are unblocked and available to dispatch.' It precisely defines readiness via status == 'todo' and fully-completed dependencies, which clearly differentiates it from siblings like list_plans, check_story_status, and dispatch_story.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says to use check_story_status for full story records and indicates this tool returns only enough to choose a key for dispatch_story or check_story_status. It also instructs the agent to call again after a dependency completes rather than trusting a stale list, providing clear when-to-use and when-to-recall guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mark_story_doneA
Mark a story finished: sets its ticket (Plane, when enabled) to Done, sets manifest["stories"][story_key]["status"] to "done", and clears any stale parked_reason. If this was the plan's last remaining story, fires the plan-completion notification.
plan_name: the plan's name, as returned by list_plans or passed to save_plan/ingest_plan. story_key: the story's key within that plan's manifest, as returned by list_ready_stories, check_story_status, or dispatch_story.
Call this only after the story's PR has actually been reviewed and
merged — approve_merge already calls this internally as its last step,
so you normally only need to call mark_story_done directly for a
merge that happened outside the pipeline (e.g. a manual gh pr merge
you've already confirmed passed CI). It does not merge or verify
anything itself; it only records that the work is done. If another
dispatch/ingest/interrupt holds the plan's lock, this is a no-op that
returns {"ok": True, "skipped": "locked", ...} rather than blocking or
raising — retry the call rather than assuming failure.
| Name | Required | Description | Default |
|---|---|---|---|
| plan_name | Yes | ||
| story_key | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully carries the behavioral burden. It discloses all side effects (ticket status, manifest status, stale parked_reason clearing, plan-completion notification), what it does not do ('does not merge or verify anything'), and lock behavior (no-op returning skipped instead of blocking). This is thorough and honest.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than average but every sentence earns its place: side effects, parameter provenance, when-to-use, non-behavior, and lock semantics are all relevant. It front-loads the core action and then layers necessary detail without fluff.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, the description covers the full calling context: prerequisites, side effects, lock behavior, and parameter sources. An output schema exists, so return-value details are not required. Nothing critical is missing for an agent to select and invoke this tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must add meaning. It does: plan_name is sourced from list_plans/save_plan/ingest_plan, and story_key is sourced from list_ready_stories/check_story_status/dispatch_story. This gives the agent exactly the provenance needed to fill both parameters correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Mark a story finished' with concrete effects on the ticket, manifest status, and parked_reason. It also differentiates itself from approve_merge by noting that approve_merge calls it internally, making the tool's unique role clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says when to call it directly ('only after the story's PR has actually been reviewed and merged'), when not to call it ('approve_merge already calls this internally'), and names the alternative path. This is exemplary usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mark_story_in_progressA
Transition the ticket to In Progress and update the local manifest's story status.
plan_name: the plan's name, as returned by list_plans or passed to save_plan/ingest_plan. story_key: the story's key within that plan's manifest, as returned by list_ready_stories or dispatch_story.
Use this before writing any code for a story. Call it once, right after dispatch_story (or after claiming a story for manual work) — it only records status, it does not create a worktree/branch or start an agent itself (dispatch_story does that). Calling it on a story that doesn't exist in the local manifest returns {"ok": False, "error": ...} rather than raising. If another dispatch/ingest/interrupt holds the plan's lock, this is a no-op that returns {"ok": True, "skipped": "locked", ...} — retry rather than assuming the status change happened.
| Name | Required | Description | Default |
|---|---|---|---|
| plan_name | Yes | ||
| story_key | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it does so thoroughly. It discloses that the call only records status, describes the no-op/locked behavior with a specific return shape, and explains that a nonexistent story returns {'ok': False, 'error': ...} instead of raising. This gives the agent a clear model of side effects and failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose and then adds only high-value operational detail. Every sentence carries meaning: parameter provenance, invocation timing, non-behavior, error semantics, and lock handling. There is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity and the presence of an output schema, the description covers everything an agent needs: when to call, what the side effects are, how to handle lock contention, and what happens for invalid stories. The lock-retry guidance is especially valuable because it prevents a false assumption about success.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must fully compensate. It explains plan_name as the name returned by list_plans or passed to save_plan/ingest_plan, and story_key as the key within the plan's manifest returned by list_ready_stories or dispatch_story. This is exactly the provenance an agent needs to populate the two required parameters correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: transition the ticket to In Progress and update the local manifest's story status. It also distinguishes itself from dispatch_story by explicitly pointing out that it does not create a worktree/branch or start an agent, which separates it from nearby sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit timing guidance: use it before writing code, call it once right after dispatch_story, and it is also appropriate after manually claiming a story. It does not enumerate alternatives like set_story_status, but the use context is clear enough for an agent to select this tool correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
patch_planA
Edit a plan manifest's top-level fields (currently only role_config) without hand-editing the manifest JSON.
Hand-editing the manifest directly races the scheduler's 60s advance_all_plans tick - a read-modify-write on either side can silently clobber the other's write. This tool acquires the same _plan_lock the scheduler and dispatch_story use, so the edit is atomic with respect to it. Only role_config may be set; every other top-level field (repo_root, epics, stories, ...) is rejected fail-closed before the lock is taken, so an unknown field can never reach disk.
| Name | Required | Description | Default |
|---|---|---|---|
| fields | Yes | ||
| plan_name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With zero annotations, the description carries the full burden and excels: it discloses the atomicity guarantee (acquires the same _plan_lock as the scheduler and dispatch_story), the concurrency race it prevents, and the fail-closed validation that rejects unknown fields before the lock is taken so nothing invalid reaches disk. These are exactly the behavioral traits an agent needs to trust a mutation tool, all beyond what any structured field would reveal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The purpose is front-loaded in the first sentence, and each subsequent sentence earns its place: the race explanation justifies why the tool exists, the lock detail establishes the atomicity guarantee, and the fail-closed note covers safety. It runs to four sentences, slightly longer than minimal, but none are filler, so it reads as dense and purposeful rather than verbose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The presence of an output schema covers return values, and the description addresses the essential concerns for a patch operation: what can be edited, why it's safe, and the validation boundary. Minor gaps remain - the precise shape of role_config and exact error behavior on rejection aren't detailed - but for a two-parameter mutation tool the description covers the high-risk aspects comprehensively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for both parameters. It clarifies that fields may only contain role_config and that setting any other top-level field is rejected fail-closed, which maps directly to the 'fields' object parameter. plan_name is implied by context as the plan identifier. It doesn't spell out plan_name's role or role_config's internal shape, but it adds substantial meaning the empty schema lacks, earning a 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (edit) and resource (plan manifest top-level fields), and immediately narrows scope to 'currently only role_config'. This clearly distinguishes it from sibling patch_story (stories vs plan manifests) and save_plan/ingest_plan (whole-plan creation/import vs targeted field edit). An agent can select it correctly without ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly frames the tool as the correct alternative to hand-editing the manifest JSON, warning of the read-modify-write race with the scheduler's 60s advance_all_plans tick. It clearly states only role_config may be set and every other top-level field is rejected, defining what not to use it for. It stops short of naming specific sibling tools as alternatives for other operations, so it's a 4 rather than a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
patch_storyA
Edit a story's plan-authored fields (agent_instructions, model, persona,
risk, dependencies, acceptance, pr_url, summary, tdd_split, backend)
without hand-editing the manifest JSON. files is not patchable: change
the plan's file list via ingest_plan and re-ingest.
Hand-editing the manifest races the scheduler's 60s advance_all_plans tick - a read-modify-write on either side can clobber the other's write. This tool takes the same _plan_lock the scheduler and dispatch_story use, so the edit is atomic with respect to it. Only the fields above may be set; status transitions go through set_story_status, not this tool.
| Name | Required | Description | Default |
|---|---|---|---|
| fields | Yes | ||
| plan_name | Yes | ||
| story_key | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden and delivers rich context: the 60s advance_all_plans tick race, the shared _plan_lock, atomicity with the scheduler and dispatch_story, and the consequence of hand-editing. This is exactly the behavioral context an agent cannot get from schema or annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, front-loaded with the core action and field list, followed by the concurrency rationale. Efficient overall, though the field enumeration is dense and could be tightened; nothing is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists, so return values need not be explained. For a nested-object mutation tool with no annotations and 0% schema coverage, the description supplies the exact editable fields, the excluded field, and the concurrency semantics — everything an agent needs to call it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate, and it does by enumerating the allowed keys inside the free-form `fields` object (agent_instructions, model, persona, risk, dependencies, acceptance, pr_url, summary, tdd_split, backend) and explicitly excluding `files`. Parameters plan_name and story_key remain undocumented but are self-evident from context.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (patch/edit) and resource (a story's plan-authored fields) and enumerates the exact editable fields. It also names the sibling set_story_status for status transitions and ingest_plan for file list changes, so an agent can distinguish it from siblings without reading schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use this tool (patch plan-authored fields without hand-editing JSON) and when not to (status transitions go through set_story_status; files are not patchable, use ingest_plan instead). Alternatives are named with the condition that selects each.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
pause_planA
Stop advance_pipeline/advance_all_plans from touching this one plan - no new dispatch, review, or merge - while leaving every other ingested plan's scheduler ticks unaffected. Any story currently in_progress is interrupted (checkpointed and left resumable) so a paused plan isn't quietly burning usage in the background. Resume with resume_plan.
| Name | Required | Description | Default |
|---|---|---|---|
| plan_name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the transparency burden and handles it well: it discloses that in_progress stories are interrupted, checkpointed, and left resumable, and that no background usage is burned. It stops short of describing already-paused behavior or persistence details, but the key side effects are covered.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences front-load the action and scope, then cover side effects and reversal. No filler, restatement, or redundant schema repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with one required string parameter and an output schema, the description provides everything needed to invoke it correctly: what is stopped, what is preserved, what happens to in-progress work, and how to undo it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage for the single plan_name parameter, and the description never mentions the parameter by name or its format. However, phrases like 'this one plan' and 'paused plan' make the target of the operation clear enough to compensate at a basic level.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: pause a single plan so advance_pipeline/advance_all_plans no longer dispatch, review, or merge it. It clearly differentiates the tool from global pipeline operations and from resume_plan.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly names the sibling tools this operation blocks and clarifies that other plans remain unaffected, which defines the precise use case. It also directs the agent to resume_plan as the paired reversal operation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
record_security_auditA
Record that a repository was security-audited at a given commit.
Parameters: repo_root: Absolute path to an existing directory containing the git repo. sha: Optional 4-64 hex-character commit id. Defaults to HEAD.
Returns:
On success, {"ok": True, "repo_root": ..., "last_audited_sha": ..., "last_audited_at": ...}. On failure, {"ok": False, "error": ...}.
| Name | Required | Description | Default |
|---|---|---|---|
| sha | No | ||
| repo_root | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden: it discloses that this is a write/state-recording operation, the sha format constraint (4-64 hex), the HEAD default, and the success/failure response shape. However, it does not say whether an existing audit record is overwritten (the 'last_audited_sha' field implies it is), nor any permission/auth requirements, leaving key mutation semantics to inference.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded purpose sentence, then a clean Parameters/Returns layout with no filler. The Returns block duplicates the existing output schema, which is slight waste, but sizing is reasonable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter mutation tool with an output schema, the definition covers the purpose, both parameters, and success/failure behavior. The residual gap is the lack of any annotation coverage plus no statement about overwrite behavior or required permissions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it does: repo_root is defined as an absolute path to an existing git directory, and sha as an optional 4-64 hex commit id defaulting to HEAD. That covers both parameters with constraints the bare schema lacks.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Record'), resource ('repository security-audited'), and scope ('at a given commit'). No sibling in the list does anything similar, so an agent can select it unambiguously from the name plus description.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the description (call it to record an audit), and the sha default to HEAD hints at the common case, but there is no explicit 'use this when / instead of X' guidance or stated prerequisite. Adequate but with a clear gap.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
request_decisionA
Escalate a blocking decision to the overlord, which rules on the user's behalf per the decision policy. The ruling is appended to the plan's decisions log (audit trail, readable via list_decisions) and returned.
plan_name: the plan's name this story belongs to.
story_key: the story's key within that plan's manifest — the story
that's actually blocked.
question: the specific question you need answered, stated so a ruling
of "pick one of these options" fully resolves it.
options: the mutually exclusive choices the overlord may rule between,
as plain strings (e.g. ["hand-roll a parser", "add a dependency"]).
Not free text — the ruling should select one of these verbatim.
context: optional — anything the overlord needs to rule correctly that
isn't in question itself (constraints, tradeoffs you've already
found, why the choice matters). Defaults to empty; provide it
whenever the bare question is ambiguous without it.
On success returns a dict with at least "ruling" (the chosen option's text), "rationale", "risk", "tier", "action", and "notify_user" — act on "ruling", not on your own preference. Call this from a story agent when you are blocked on a choice the user would normally make; do not guess.
Fails open: if the overlord backend errors, the story is parked for a human and a single-line escalation message (a plain string, not the dict above) is returned instead of raising — check whether the return value is a str before reading dict keys off it.
| Name | Required | Description | Default |
|---|---|---|---|
| context | No | ||
| options | Yes | ||
| question | Yes | ||
| plan_name | Yes | ||
| story_key | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description fully carries the burden. It discloses the side effect ('appended to the plan's decisions log'), the return contract (dict with specified keys), and the failure mode ('Fails open... returns a single-line escalation message... check whether the return value is a str'). This is unusually complete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Although lengthy, the description is densely packed and well-structured: purpose first, then parameter semantics, then usage, then failure handling. Every sentence contributes actionable information, and the return-type caveat is front-loaded in the failure paragraph.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a complete standalone description despite having no annotations. It covers what the tool does, when to call it, all parameter meanings, the success return format, the failure return format, and how to distinguish the two. An agent has everything needed to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description documents every parameter in detail. It explains that options must be mutually exclusive plain strings selected verbatim, and that context should be supplied whenever the bare question is ambiguous. This adds meaning far beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Escalate a blocking decision to the overlord, which rules on the user's behalf per the decision policy.' It clearly distinguishes this tool from the siblings by focusing on obtaining a decision rather than modifying stories, plans, or statuses.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to call it: 'Call this from a story agent when you are blocked on a choice the user would normally make; do not guess.' This gives a clear trigger condition and tells the agent not to guess, which effectively directs behavior away from alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
resume_planA
Clear a pause set by pause_plan so this plan's stories are eligible for dispatch/review/merge on the next advance_pipeline tick again.
| Name | Required | Description | Default |
|---|---|---|---|
| plan_name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description must carry the behavioral burden. It discloses the timing (next advance_pipeline tick) and the effect (stories eligible for dispatch/review/merge), but does not discuss reversibility, side effects, or whether there are any prerequisites beyond having a paused plan. This is a basic resume operation, and the description adds some context beyond just 'resume'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, efficient sentence that immediately states the primary action (clear a pause) and the consequence (stories become eligible). No filler words; every part carries meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is simple with one parameter and no output schema expectations explained, but an output schema exists (possibly indicating a return value). The description could mention what the response looks like, but given the tool's simplicity and the known pipeline context, it is largely complete for an agent to use correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With only one parameter, plan_name, and 0% schema description coverage, the description provides no additional meaning beyond the schema. Since the parameter name is self-explanatory, the baseline of 3 is appropriate; no further compensation is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: clearing a pause set by pause_plan, and its effect: making stories eligible for dispatch/review/merge on the next advance_pipeline tick. It references the specific sibling (pause_plan) and the pipeline mechanic, distinguishing it from other tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explains the intended use case: resuming a plan that was paused. It implicitly indicates it should be used after pause_plan, but doesn't explicitly state when not to use it or list alternatives. However, given the clear pair with pause_plan, the guidance is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
review_storyA
Run the code-reviewer persona over a dispatched story's branch. On APPROVE, open a PR via gh and set status to pr_open; otherwise set status to changes_requested. Does not merge — merge is the overlord's decision.
Only reviewable when story["status"] == "tests_passed" - any other status (a stale/duplicate call, e.g. a second tick racing an already-merged story) is a no-op skip; see README.md's "Review & merge" section.
| Name | Required | Description | Default |
|---|---|---|---|
| plan_name | Yes | ||
| story_key | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, and it delivers: it discloses the concrete side effects (opens PR via gh, sets status to pr_open or changes_requested), explicitly states it does not merge, and explains no-op behavior for invalid statuses. This is strong behavioral disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tightly packed paragraphs that front-load the main behavior, then add the critical precondition and exclusion. Every sentence adds usable information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with an output schema, the description covers the essential call-time facts: action, status transitions, non-merge guarantee, and invalid-call handling. It also references the README for deeper review/merge detail, so an agent has enough to invoke it safely.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description does not define plan_name or story_key beyond their names. The general story/status context helps indirectly, but the required parameters are left mostly implicit.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: running the code-reviewer persona over a dispatched story's branch. It clearly distinguishes itself from merge-related siblings by explicitly saying 'Does not merge — merge is the overlord's decision.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit precondition: only reviewable when story status is 'tests_passed', and any other status is a no-op skip. It even anticipates stale/duplicate racing calls and points to README for more context, making when-to-use versus when-to-avoid unmistakable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
save_planA
Save a generated project plan to disk. Plan should be JSON matching the schema: { "epics": [ { "summary", "stories": [...] } ] }. Call this after generating a plan so the user can review before ingestion.
workspace is optional. When supplied, it is validated and (WS-11) the model-authored repo_root in plan_json is OVERWRITTEN with the server-validated resolved path (the server overwrites the model-authored value; this tool never resolves or rewrites plan_json itself). When omitted, the plan's own repo_root is trusted, exactly as before. The tool deliberately does NOT fall back to the dashboard's persisted active workspace: the MCP server and the dashboard are separate processes, and silently coupling them through shared durable state is out of scope (that fallback lives only in the dashboard HTTP route).
| Name | Required | Description | Default |
|---|---|---|---|
| plan_json | Yes | ||
| plan_name | Yes | ||
| workspace | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and excels. It discloses key behavioral traits: the overwriting of repo_root when workspace is supplied (WS-11), that the tool never resolves or rewrites plan_json itself, and that it deliberately does not fall back to the dashboard's persisted workspace, explaining the separate-process rationale. These details exceed what annotations would typically cover and prevent misleading assumptions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is longer than typical but well-structured: purpose, format, then a detailed workspace behavior paragraph. The rationale for the no-fallback decision is valuable and earns its place. It is front-loaded with the core purpose and scoping before diving into edge cases. No wasted sentences.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (optional workspace, overwriting, process boundary), the description covers the critical aspects: the JSON schema, the overwriting rule, the fallback decision. An output schema exists to document return values, so that omission is acceptable. It does not address error conditions or interactions with siblings beyond the ingestion hint, but for a save operation the provided details are largely sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate. It provides extensive semantics for workspace (validation, overwriting behavior, fallback rationale) and clarifies plan_json's required schema. However, plan_name is not explicitly described beyond being a required string, though its meaning as a name is likely inferred. The description adds significant value for two of three parameters, nearly fully compensating for the schema gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action: 'Save a generated project plan to disk.' It specifies the expected JSON format and differentiates from ingestion by noting 'so the user can review before ingestion,' distinguishing it from the sibling ingest_plan. The verb, resource, and purpose are all explicit.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: 'Call this after generating a plan so the user can review before ingestion.' This implies the correct sequence (generate → save → review → ingest) and indicates when to use it. It does not name alternative tools explicitly, but the 'before ingestion' phrase sets the stage, and the workspace behavior explains when to omit or supply it. No explicit exclusions are given, but the guidance is sufficient.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
set_story_statusA
Transition a story to an explicit status without hand-editing the manifest JSON (e.g. resetting a "parked" story to "interrupted" so the scheduler retries it).
Acquires _plan_lock for the same reason patch_story does. Only accepts the fixed set of statuses the pipeline itself assigns (todo/in_progress/running/interrupted/failed/tests_passed/pr_open/ changes_requested/parked/done) - this is a sanctioned status change, not a way to invent pipeline state the rest of the code doesn't expect.
| Name | Required | Description | Default |
|---|---|---|---|
| status | Yes | ||
| plan_name | Yes | ||
| story_key | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden and does well by disclosing that it acquires _plan_lock and only accepts pipeline-defined statuses. This tells the agent about locking side effects and invariants. It does not cover possible failures or permission requirements, but the key behavioral constraints are present.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the primary action and includes an example before adding lock and status constraints. The status list is long but necessary because the schema lacks enums. Minor indirectness ('for the same reason patch_story does') keeps it from being perfectly self-contained.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema, so return-value details are not required. The description gives enough context for a correct call: what the tool does, the exact allowed statuses, an example use case, and lock behavior. The only gaps are minor details about the plan/story parameters and failure conditions.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% and no enum metadata exists, but the description supplies the complete allowed status set, which is the critical parameter meaning. plan_name and story_key are left to their self-explanatory titles, and the description's story/plan context partially covers them. This is solid compensation for a low-coverage schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Transition a story to an explicit status') and contrasts it with hand-editing the manifest JSON. The fixed status list and 'sanctioned status change' clarify exactly what this tool is for, making it distinguishable from related story tools by defining its scope rather than by name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a concrete usage example ('resetting a parked story to interrupted so the scheduler retries it') and says the tool is the sanctioned path rather than inventing state. It does not explicitly name alternative sibling tools or state when not to use it, so it lacks full exclusion guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.6.0- Added
record_security_audit
5 tool updates
v0.4.0- Added
check_story_status - Added
checkpoint_story - Removed
get_role_config - Added
patch_plan - Changed
request_decision4 fields changed- removed
Output schema / additionalPropertiesRemoved value: -true - added
Output schema / propertiesAdded value: +{ + "result": { + "anyOf": [ + { + "additionalProperties": true, + "type": "object" + }, + { + "type": "string" + } + ], + "title": "Result" + } +} - added
Output schema / requiredAdded value: +[ + "result" +] - changed
Output schema / titlePrevious value: -"request_decisionDictOutput"New value: +"request_decisionOutput"
23 tool updates
v0.1.0- First observed
advance_all_plans - First observed
advance_pipeline - First observed
approve_merge - First observed
check_usage - First observed
checkpoint - First observed
decompose_plan - First observed
dispatch_story - First observed
get_effective_config - First observed
get_role_config - First observed
ingest_plan - First observed
interrupt_story - First observed
list_decisions - First observed
list_plans - First observed
list_ready_stories - First observed
mark_story_done - First observed
mark_story_in_progress - First observed
patch_story - First observed
pause_plan - First observed
request_decision - First observed
resume_plan - First observed
review_story - First observed
save_plan - First observed
set_story_status
TDQS
Scored across 26 tools
Several tools overlap in purpose: checkpoint is a deprecated alias of checkpoint_story, and mark_story_in_progress/mark_story_done/set_story_status all mutate story status while dispatch_story also transitions to In Progress. Detailed descriptions mitigate confusion, but an agent could still misselect among the status-transition tools.
Nearly all tools follow a consistent snake_case verb_noun pattern (e.g. list_plans, dispatch_story, approve_merge). The main deviation is the deprecated 'checkpoint' alias, which drops the object noun and breaks the pattern slightly.
26 tools is on the heavy side, but the server covers a broad multi-agent pipeline domain (planning, dispatch, review, merge, decisions, config, usage, audit). The count is justified though a few redundant tools (deprecated alias, overlapping status setters) inflate it.
The surface covers the full plan-to-merge lifecycle including planning, ingestion, dispatch, checkpointing, review, merge, decisions, and operational diagnostics. Some gaps remain, such as listing all stories in a plan or managing plan/story deletion, but core workflows are complete.
Maintenance
Related MCP Connectors
Control plane for autonomous software labor. Agents claim objectives over MCP with audit trail.
AI-native git hosting — repos, PRs, issues, CI gates, and AI code review over MCP (60 tools).
A paid remote MCP for OpenAI Codex agent coordination MCP, built to return verdicts, receipts, usage
Real-time chat for AI agents. Claude Code, Cursor, Cline and Codex join channels over MCP.
Related MCP Servers
- AlicenseAqualityDmaintenanceMCP server that spawns autonomous Claude Code agents in GitHub repos, enabling task delegation with persistent state, multi-step workflows, and job monitoring.47109 npm2Apache 2.0
- AlicenseAqualityDmaintenanceOne MCP that turns Claude Code into your whole dev stack by swallowing other MCP servers, delegating to Codex & Gemini on your CLI subscriptions, remembering projects in a searchable knowledge graph, and carrying setup across sessions — secret-free by design.234MIT
- AlicenseAqualityDmaintenanceMulti-agent orchestration MCP server that lets Claude Code delegate backend, frontend, and tooling tasks to specialized AI agents (Codex, Kimi, Grok) with contract-first sequencing and workspace mutation guards.47MIT
- AlicenseNot gradedqualityCmaintenanceAn MCP server that lets ChatGPT or any MCP client securely delegate coding tasks to a local Claude Code instance, with git checkpointing, approval gates, and structured results. Supports code review, test running, and rollback via simple tool calls.15 npmMIT