hamgoose
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@hamgoosestart a mission to migrate auth from session cookies to JWT"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
🦆 Factory-Droid-style Mission orchestration for Goose
Type a goal → get a structured plan → approve → watch isolated workers build it, get validated, get corrected — until it's done and proven.
USER GOAL → ANALYSIS → STRUCTURED PLAN → FEATURES + DEPS + MILESTONES → APPROVAL
→ DEPENDENCY-AWARE EXECUTION (isolated workers) → REAL CODE
→ SCRUTINY + USER-FACING VALIDATION → AUTOMATIC CORRECTION
→ FINAL VALIDATION → MISSION COMPLETED ✅hamgoose is a genuine Goose extension — a standalone stdio MCP server on the
official mcp/FastMCP model, not a recipe, todo wrapper, or delegation
prompt. Code enforces the orchestration mechanics; models do the semantic
reasoning.
📖 Contents
Related MCP server: Codex Orchestrator
✨ Features
🚦 Approval gate | Nothing is implemented until you approve the plan |
ðŸï¸ Isolated leaf workers | Each feature runs in its own |
ðŸ•¸ï¸ Dependency-aware scheduling | A DAG of features with path-overlap conflict detection and a hard concurrency cap (your provider's limits, enforced in code) |
🔠Two validators | Scrutiny distrusts the worker's claims and inspects diff + tests; user-testing exercises the app from the user's perspective |
🔠Automatic correction | Failed validation becomes corrective features; the bounded loop repeats until the milestone passes |
🧯 Crash recovery | Atomic JSON state + append-only event log — kill Goose mid-mission, reopen, it reconciles and continues |
📡 Live progress | Long calls stream MCP progress notifications, and the guided flow runs in short visible bursts — you always see movement, never a silent wait |
🧠Steering & replanning | Change course mid-mission without losing completed work |
🔠Secrets redacted | Every persisted artifact scrubbed of keys, tokens, credentials |
🪶 Per-repo state | Lives in |
🚀 Quickstart
Node — via npm:
npm i -g @cooked-ham/hamgoose
hamgoose registerPython — via pip:
pip install git+https://github.com/cooked-ham/hamgoose.git # Python 3.11+
hamgoose registerEither world, one shot (no install): npx @cooked-ham/hamgoose register
That's the whole install — two commands, no repo wiring, no config surgery.
Pickone channel (npm or pip); both provide the same hamgoose command.
Uninstall: hamgoose unregister + npm uninstall -g @cooked-ham/hamgoose
(or pip uninstall -y hamgoose).
🎯 The walkthrough
Then, in any repository you're working in:
$ goose
You: /prompt start_mission
goose: What's the goal?
You: Migrate the auth module from session cookies to JWT.
goose: Any rules or constraints? (concurrency, provider/model, git, validation)
You: My provider only allows 3 concurrent agents at a time.
goose: Plan: 2 milestones, 6 features, workers capped at 3 concurrent. Approve?
You: Approve.
goose: MS01 1/3 … passed scrutiny … MS02 2/3 …
✅ Mission COMPLETED — changes on branch mission/base with per-feature commits.Rules are recorded verbatim on the mission (visible in every status and
plan view), translated into execution config ("max 3 concurrent" →
max_concurrent_workers: 3), and handed to every worker as context.
Mid-mission you can just say "pause", "don't touch config files", or
"replan around X" — steering and replanning never lose completed work.
No slash command? Just say"start a hamgoose mission" in plain English.
/prompt start_mission runs the guided setup; plain English does the same thing.
ðŸ—ï¸ How it works
flowchart LR
U["👤 You<br/>goal + rules"] --> G["Goose session"]
G <-->|MCP stdio| H["🦆 hamgoose<br/>orchestrator<br/>(deterministic code)"]
H -->|isolated goose run| W1["Worker F001<br/>🌳 worktree"]
H -->|isolated goose run| W2["Worker F002<br/>🌳 worktree"]
H --> V["🔠Validators<br/>scrutiny + user-test"]
W1 -->|merge + commit| R[("repo<br/>mission/base")]
W2 -->|merge + commit| RstateDiagram-v2
[*] --> CREATED
CREATED --> ANALYZING
ANALYZING --> PLANNING
PLANNING --> AWAITING_APPROVAL
AWAITING_APPROVAL --> RUNNING : approve
RUNNING --> PAUSED
RUNNING --> BLOCKED
PAUSED --> RUNNING : resume
BLOCKED --> RUNNING : resolve + resume
RUNNING --> VALIDATING
VALIDATING --> RUNNING : corrective work
VALIDATING --> COMPLETED : all pass ✅
CREATED --> CANCELLED
AWAITING_APPROVAL --> CANCELLED
RUNNING --> FAILED
COMPLETED --> [*]
FAILED --> [*]
CANCELLED --> [*]Why it's different from "just let the agent do it": the orchestrator is deterministic code — state machines, DAG scheduling, retries, Git bookkeeping, persistence are enforced, not hoped for. The LLM only does what LLMs are good at: understanding intent and writing code. A confused model can't corrupt the mission state, skip the approval gate, or double-dispatch a feature. See ARCHITECTURE_REPORT.md for the full design analysis.
📦 Install
Requires goose (≥ 1.40) on your PATH and git (for Git missions).
Option | For | Command |
1. From GitHub â | Everyone |
|
2. Goose's own menu | No extra commands |
|
3. From a clone | Contributors |
|
4. Per-run | One-off experiments |
|
5. npm (Node world) | Node-first machines |
|
Pin a release once tags exist:pip install "git+https://github.com/cooked-ham/hamgoose.git@v0.1.2".
# config.yaml — path printed by `goose info`
extensions:
hamgoose:
enabled: true
type: stdio
name: hamgoose
description: Mission orchestration for Goose
cmd: hamgoose
args: []ðŸ› ï¸ The lifecycle (tools)
Operation | Tool |
Create + analyze repo (guided setup) |
|
Generate the plan (approval gate) |
|
Approve & start |
|
Execute the control loop (resumable) |
|
Pause / resume |
|
Steer (priority / guidance) |
|
Replan (new constraint) |
|
Retry a feature / validate now |
|
Cancel |
|
Read status / plan / events / list |
|
Resources (read): mission://{id}/status|plan|events|features|milestones|validation
Prompts (list with /prompts, run with /prompt <name>): start_mission (guided setup) · plan_mission · resume_mission · validate_milestone
ðŸ—„ï¸ Where state lives
<repo>/.goose/hamgoose/<mission-id>/
├── mission.json # canonical atomic state
├── mission.yaml # human-readable mirror
├── plan.md # plan mirror
├── events.jsonl # append-only event log
├── workers/ # redacted worker transcripts
├── validation/ # validation reports
├── worktrees_base/ # mission/base worktree (merged result)
└── worktrees/<F> # per-feature worktreesYour current branch is never modified — mission/base accumulates the
merged, validated result for you to merge. Add /.goose/hamgoose/ to your
repo's .gitignore.
🧪 Development
git clone https://github.com/cooked-ham/hamgoose.git && cd hamgoose
uv venv .venv
uv pip install -p .venv -e ".[dev]"
.venv/bin/python -m pytest -m "not realgoose" # fast, deterministic (no LLM)
.venv/bin/python -m pytest -m "realgoose" # real Goose + LLM (slower)📚 Docs
ARCHITECTURE_REPORT.md — design decisions & compliance analysis
ARCHITECTURE.md — component architecture
MISSION-LIFECYCLE.md — state machines & control loop
CONFIGURATION.md — config reference, registration, known limitations
TESTING.md — test strategy
WINDOWS.md — Windows gotchas & hardening notes (cmd, cp1252, CRLF, tree kills, pipe hang, pytest temp)
FIX_PLAN.md — the v0.1.8 hardening fix plan & implementation record (HG-01…HG-17)
remainingwork_hamgoose.md — the original failure evidence from 5 missions
PUBLISHING.md — PyPI release runbook
📜 License
MIT — see LICENSE.
Available Tools
19 toolsmission_apply_suggestionsA
H10: apply the config deltas recorded at mission create when the model preflight flagged the worker model (e.g. worker_timeout>=900 for SMALL-OUTPUT-BUDGET models). One call; the applied values are echoed.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | ||
| mission_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Since no annotations are provided, the description carries the burden of behavioral disclosure. It conveys that the tool mutates mission configuration, that it applies pre-recorded deltas, that it is meant to be called once, and that the applied values are echoed. It does not mention side effects, reversibility, or the no-delta case, but the core behavior is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and action-first, with no filler. The cryptic 'H10:' prefix is unexplained and adds little value, which prevents a perfect score, but the rest is tightly written and informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the trigger, the action, the one-call behavior, and the echoed output, which is good for a focused mutation tool. However, it omits any explanation of the repo parameter and does not situate this tool against the many mission_* siblings, leaving some context incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, and the description does not explain the purpose or format of mission_id or repo. The mission_id is inferable from context and the sibling tool naming, but repo is entirely undocumented. With such low schema coverage, the description needed to compensate and did not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('apply the config deltas recorded at mission create') on a concrete resource, with a triggering condition (model preflight flagged the worker model). It is clearly distinct from the sibling mission_* tools, which handle creation, planning, running, and validation rather than applying recorded suggestions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives a clear when-to-use condition, including a concrete example (worker_timeout>=900 for SMALL-OUTPUT-BUDGET models). It does not explicitly name alternatives or state when not to use the tool, so it stops short of full routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mission_approveA
Approve the plan and begin dependency-aware execution. Safe to call only once, from AWAITING_APPROVAL.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | ||
| mission_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses that the tool is non-idempotent ('Safe to call only once') and state-dependent ('from AWAITING_APPROVAL'), and that it triggers dependency-aware execution. It does not detail all side effects, but the key transition behavior is transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, tight sentence with no filler. The action is stated first, and the critical safety constraint is included in the same sentence. Every word contributes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a state-transition tool with an output schema and a simple input schema, the description covers the core semantics and the key precondition. The main gap is the unexplained optional repo parameter, but the tool is otherwise sufficiently specified for an agent in the AWAITING_APPROVAL state.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not mention either parameter. mission_id is reasonably inferable from the tool name, but repo is left entirely unexplained, and the description does not clarify how either parameter affects the approval or execution behavior.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Approve') and resource ('the plan'), and clearly states the resulting behavior ('begin dependency-aware execution'). It also distinguishes this tool from siblings like mission_plan or mission_validate by tying it to the approval transition.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when it is safe to call: 'only once, from AWAITING_APPROVAL.' This gives a clear precondition and cautions against repeated calls, though it does not explicitly name alternative tools or the action to take when not in that state.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mission_cancelC
Cancel a mission and clean up its worktrees.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | ||
| mission_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries full responsibility for behavioral disclosure. It does reveal a key side effect: cancellation plus worktree cleanup, which signals destructiveness. However, it does not state whether the operation is irreversible, what happens to running processes, or what cleanup entails in practice.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single efficient sentence with no filler. The core action is front-loaded, and the cleanup behavior is mentioned immediately.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the destructive nature of cancellation and an undocumented optional repo parameter, this minimal description leaves important context missing. The output schema exists, so return values need not be described, but the tool's effect on mission state and worktrees is only vaguely conveyed.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description adds no meaning to the parameters. It does not explain that mission_id identifies the mission to cancel, nor does it clarify the optional repo parameter's role in locating or cleaning up worktrees. The agent must guess at the intended semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Cancel') and a clear resource ('mission'), and adds a distinctive behavioral detail ('clean up its worktrees'). It is reasonably distinguishable from sibling tools like mission_pause or mission_resume, though it does not explicitly differentiate itself.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance on when to use this tool versus alternatives such as mission_pause or mission_complete_feature. There is no mention of prerequisites, when cancellation is appropriate, or when a different mission lifecycle tool should be chosen.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mission_complete_featureA
Record work on a feature that was implemented OUTSIDE the worker pipeline (by you, the lead agent, or a human). Use this instead of editing mission state files by hand. Verifies the commit exists, runs the feature's validation_commands, runs a real scrutiny validation on the diff, appends proper events, and continues the normal milestone flow.
Args: summary: what was implemented and why it meets the acceptance criteria. commit: the git hash of the implementation (required for git missions). changed_files: JSON list of paths (auto-derived from the commit if omitted). tests: JSON list of verification commands you ran and their outcome.
NOTE (H1): this runs a REAL scrutiny validation and can take minutes. If your client sandbox times the call out, the completion still proceeds server-side - verify via mission_events / mission_status instead of re-calling blindly.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | ||
| tests | No | ||
| commit | No | ||
| summary | Yes | ||
| feature_id | Yes | ||
| mission_id | Yes | ||
| changed_files | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden, and it does so excellently. It discloses that the tool verifies the commit, runs validation_commands, executes a real scrutiny validation, appends events, and continues milestone flow. It also warns about long runtime, sandbox timeouts, and server-side completion, which is high-value behavioral context.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured: purpose first, then parameter guidance, then a critical timeout note. Every sentence earns its place, and the most important warnings are front-loaded. The note about not re-calling blindly is especially useful and not redundant.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations and no schema-level descriptions, this is unusually complete. It covers when to use it, what side effects occur, parameter semantics, and failure-mode behavior. The output schema exists, so not describing return values is acceptable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it adds meaningful detail for summary, commit, changed_files, and tests. It explains formats, defaults, and the special requirement that commit is needed for git missions. It does not explicitly explain repo or the required mission_id/feature_id, but those are largely self-explanatory from their names and schema types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: record work on a feature that was implemented outside the worker pipeline, and it contrasts this with editing mission state files by hand. It is specific about the verb and resource, and the context makes it distinguishable from generic mission tools. However, it does not explicitly differentiate itself from sibling tools such as mission_retry_feature by name.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear context for when to use the tool: when a feature was implemented outside the worker pipeline, and it explicitly advises using this instead of hand-editing mission state files. It does not, however, mention exclusions or directly compare against sibling tools like mission_validate or mission_retry_feature.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mission_createA
Create a hamgoose mission. Use this to START a mission for the user.
Guided setup protocol:
If the user has not stated a clear goal, ask for it (one short question).
Ask ONCE whether they have rules/constraints worth recording (e.g. concurrency limits, provider/model for workers, git on/off, validation toggles). If they say none, proceed with defaults - do not interrogate.
Pass their rules VERBATIM in
rules(persisted on the mission; shown in plan and status; given to every worker as context).Translate rules into
configoverrides with this map:"max N concurrent agents/workers/subagents" -> {"execution": {"max_concurrent_workers": N}}
"one worker at a time" / "sequential" -> {"execution": {"max_concurrent_workers": 1}}
"same for validators" -> {"validator": {"provider": ..., "model": ...}}
"use / for planning" -> {"planner": {"provider": ..., "model": ...}}
"no git / no worktrees" -> {"git": {"enabled": false, "use_worktrees": false}}
"skip user-facing testing" -> {"validation": {"user_testing": false}}
"no scrutiny validation" -> {"validation": {"scrutiny": false}} Note: workers are always isolated
goose runleaf processes (never nested delegation); max_concurrent_workers caps how many run simultaneously.
Returns the mission id, a readiness report and next steps. Next: mission_plan, present the plan, get approval, mission_approve, mission_run.
| Name | Required | Description | Default |
|---|---|---|---|
| goal | Yes | ||
| repo | No | ||
| rules | No | ||
| config | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description carries the burden. It discloses that rules are persisted, shown in plan/status, and given to every worker, and that workers are isolated 'goose run' leaf processes capped by max_concurrent_workers. It also states that the response includes mission id, readiness report, and next steps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is long but organized into a numbered setup protocol with a config override map and a workflow handoff. It front-loads the core purpose, and nearly every sentence carries operational value, though it is denser than necessary.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is nearly complete for a mission-creation tool: it covers goal discovery, rule handling, config translation, worker behavior, and next steps, and mentions return contents even though an output schema exists. The main omissions are `repo` semantics and explicit defaults, leaving minor ambiguity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema coverage, the description compensates substantially for `rules` and `config` by requiring verbatim rules and providing a clear natural-language-to-config map. However, `repo` is never explained, and `goal` is only implied, leaving two of four parameters with little semantic guidance.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States 'Create a hamgoose mission' and explicitly says 'Use this to START a mission for the user,' making the resource and verb concrete. The 'Next: mission_plan...' handoff distinguishes it from sibling tools like mission_plan and mission_run.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The guided setup protocol gives clear when-to-do steps: ask for the goal if absent, ask once about rules, proceed with defaults if none. It also routes to the follow-up workflow. It does not explicitly state when not to use this tool in favor of an alternative, but the workflow context is clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mission_eventsC
Recent mission events (append-only event log).
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | ||
| tail | No | ||
| mission_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the burden of behavioral disclosure. 'Append-only event log' usefully suggests that events are immutable and reading them is safe, but it does not disclose ordering, tail semantics, repo filtering, or whether the tool is strictly read-only.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very short, front-loaded, and contains no filler words. It is efficiently phrased, though its brevity means it omits operational details that other dimensions must compensate for.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations, 0% schema description coverage, and a large sibling set, the description is too minimal. It does not explain how tail works, what repo does, what event ordering to expect, or how this differs from mission_status or mission_plan_view.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description mentions none of the three parameters. mission_id, tail, and repo remain semantically undocumented beyond their names and types, so the description adds no value for parameter understanding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The phrase 'Recent mission events' clearly identifies the resource being accessed, and 'append-only event log' distinguishes this from mission planning, execution, and retry siblings. It lacks an explicit verb like 'retrieve' or 'list,' but the meaning is strongly implied.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies this tool is for reading recent event history for a mission, especially given the 'append-only event log' phrasing. However, it provides no explicit guidance about when to use this tool versus alternatives or when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mission_gcA
H11 housekeeping: list terminal and long-stale missions that clutter mission_list. With archive=true, non-terminal stale missions older than max_age_days are cancelled (their data and event history are kept). Terminal missions are never re-touched.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | ||
| archive | No | ||
| max_age_days | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden. It transparently discloses that archive=true cancels non-terminal stale missions older than max_age_days while preserving data and event history, and that terminal missions are never modified. Core destructive behavior is clear, though ancillary details like permissions or error behavior are absent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three compact sentences with the main action front-loaded in the first clause. It contains no filler, and each sentence adds distinct information: the listing purpose, the conditional archive behavior, and the terminal-mission guard.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given that an output schema exists and no annotations are present, the description covers the core behavior, side effects, and the terminal-mission exception well enough for typical calls. The main gap is the undocumented repo parameter, which would require an agent to infer its meaning from the name alone. Overall, it is a fairly complete definition for a tool of this complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for parameter meaning. It does explain archive and max_age_days via 'archive=true' and 'older than max_age_days', but it never explains the repo parameter. This is partial compensation: two of three parameters receive meaningful semantic grounding.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description immediately identifies the tool as 'H11 housekeeping' and states a specific action: list terminal and long-stale missions cluttering mission_list, with archive mode cancelling eligible non-terminal missions. It clearly pairs a verb with a resource and distinguishes itself from normal mission listing by the cleanup/housekeeping framing. The terminal-mission exclusion further sharpens the scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The 'H11 housekeeping' prefix gives useful context, and the archive=true conditional plus 'Terminal missions are never re-touched' clarify when mutation happens and what is excluded. However, the description does not explicitly name sibling alternatives such as mission_list or mission_cancel, nor does it say when to prefer them. Usage guidance is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mission_listB
List all missions in the repository.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. 'List' strongly implies a read-only operation and 'all missions in the repository' conveys scope, but it does not mention optional repo behavior, pagination, permissions, or any other side-effect details. The output schema covers return structure, so that omission is less critical.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, clear, front-loaded sentence with no filler. It is appropriately concise for a simple list operation, though it relies on the schema and output schema for additional detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description plus the available output schema provide a minimally usable picture for a simple list tool. However, the optional repo parameter is unexplained, and the lack of sibling differentiation means an agent cannot fully determine when mission_list is the correct choice among the many mission_* tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage for the optional repo parameter, and the description does not explain how to use it, what values are valid, or what the default of null means. The word 'repository' provides a loose hint that the repo parameter identifies the repository, but it does not compensate for the missing parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the operation ('List') and the resource ('missions'), and adds the scoping phrase 'all ... in the repository'. It is distinguishable from the sibling tools by the list verb, though it does not explicitly contrast with related tools like mission_status or mission_plan_view.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no guidance about when to choose this tool over the many sibling mission_* tools, and it mentions no alternatives or exclusions. The imperative 'List' implies a use case, but the agent is left to infer when this is the right tool.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mission_pauseB
Pause an active mission. No new workers are launched while paused.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | ||
| reason | No | ||
| mission_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral burden. It discloses one concrete effect: 'No new workers are launched while paused.' However, it does not clarify whether currently running workers are stopped, whether the pause is reversible, or what state changes occur beyond the pause flag.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences provide the core action and a key behavioral consequence without any filler. The description is front-loaded and every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a relatively simple pause operation, the description covers the main purpose and the most important behavioral nuance. An output schema exists, so return-value details are not required. It is slightly incomplete in that it leaves the repo parameter unexplained and does not mention how pause relates to resume, but the essentials for invocation are present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain any parameter. The mission_id parameter is inferable from the tool's purpose, and reason is partially self-explanatory, but repo remains ambiguous. The description fails to compensate for the lack of parameter documentation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states a specific action and resource: 'Pause an active mission.' It also adds a meaningful behavioral qualifier about worker launches. It does not explicitly name sibling tools like mission_resume or mission_cancel, but the pause action is distinct enough that an agent can likely select it correctly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies usage: pause an active mission when you want to stop new workers from launching. It gives no explicit guidance about when not to use it or how it compares to alternatives such as mission_cancel or mission_resume, leaving the agent to infer the context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mission_planA
Generate the structured dependency-aware plan and present it for approval. No implementation happens until mission_approve is called.
If the planner returns an empty plan (goal too vague), retry mission_plan - or pass your OWN decomposition (JSON string or list): features='[{"id":"F001","title":"...","description":"...","milestone":"MS01", "dependencies":[],"acceptance_criteria":["..."],"expected_paths":["..."]}]' milestones='[{"id":"MS01","objective":"...","completion_criteria":["..."]}]'.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | ||
| features | No | ||
| milestones | No | ||
| mission_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It explicitly discloses that the tool does not implement anything, that it presents the plan for approval, and that the planner may return an empty plan when the goal is too vague. This is good disclosure, though it omits any permission or side-effect details.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The primary behavior and approval gating are front-loaded in the first sentence. The JSON examples are long but necessary given the 0% schema coverage, and each paragraph contributes distinct information. Slightly verbose, but efficiently organized.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the core workflow, the approval gate, recovery from empty plans, and the shape of optional decomposition inputs. Since an output schema exists, return-value details are not needed. Minor gaps remain around repo/mission_id semantics and how this relates to mission_replan, but the description is largely complete for invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It provides valuable JSON examples for features and milestones, including nested field names and dependency structures. It does not explicitly explain the required mission_id or optional repo parameter, but their purpose is reasonably inferable and the most complex parameters are well covered.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Generate the structured dependency-aware plan and present it for approval.' It clearly distinguishes the tool from sibling actions by stating no implementation happens until mission_approve is called.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives concrete usage context: use this to produce an approval-ready plan, and use mission_approve for implementation. It also spells out a fallback for empty planner results (retry or supply own decomposition). It does not exhaustively contrast all sibling tools, but the main routing cues are present.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mission_plan_viewC
The current plan for a mission.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | ||
| mission_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It implies a read operation by saying 'current plan' but does not disclose side effects, freshness, error behavior, or whether it modifies anything.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is short and front-loaded, but this is under-specification rather than effective conciseness. One sentence states only the core purpose and omits necessary contextual information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With two parameters, no annotations, an output schema, and many sibling tools, the description is incomplete. It does not provide selection criteria, parameter semantics, or behavioral expectations, though the output schema mitigates the need to document return values.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the description mentions no parameters. It does not clarify the meaning or usage of mission_id or the optional repo parameter, so an agent cannot determine parameter semantics from the description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies the resource ('the current plan for a mission') and the tool name supplies the verb 'view'. It is clear enough to distinguish from related tools like mission_plan and mission_replan, though it does not explicitly contrast them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance about when to use this tool versus mission_status, mission_plan, mission_replan, or other siblings. The intended selection context is left entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mission_readinessC
Readiness/preflight report for a mission.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | ||
| mission_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries the full burden of behavioral disclosure. 'Readiness/preflight report' suggests a non-mutating check, but it does not explicitly state that no mission state is changed, what the report contains, or whether any side effects occur. The description is too thin for a tool with no annotation support.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very short and front-loaded, with no wasted words. However, it is concise partly because it omits necessary detail, so it cannot receive a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value structure may not need to be described. Still, there is no indication of what a readiness report contains, what conditions are checked, when to use this tool, or how the optional repo parameter affects behavior. The definition is too minimal to fully support correct tool selection and invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not compensate by explaining the meaning or purpose of mission_id or repo. While mission_id is somewhat self-explanatory, the optional repo parameter is completely unexplained. The description adds no value beyond the schema's field names and types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies the resource (mission) and the kind of output (readiness/preflight report), which is reasonably clear even though it lacks an explicit verb like 'generate' or 'return'. It is not a tautology and is distinguishable enough from sibling tools like mission_status or mission_validate.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The word 'preflight' implies this tool should be used before a mission run, suggesting a precondition or readiness check. However, it does not explicitly state when to use it versus siblings such as mission_validate, mission_status, or mission_run, nor does it mention exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mission_replanA
Replan the remaining work around a new constraint. Preserves valid completed work, marks invalidated work superseded, and bumps the plan revision.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | ||
| mission_id | Yes | ||
| instruction | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden of disclosing behavior. It does this well by stating that valid completed work is preserved, invalidated work is marked superseded, and the plan revision is bumped. It could add more about side effects or irreversibility, but the key behavioral traits are transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact, front-loaded with the main action, and every sentence adds value. It states the core operation and then key behavioral outcomes without filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Although an output schema exists, the description omits essential parameter semantics and provides no explicit guidance on how to phrase the replan request or which parameters matter. For a tool with three undocumented parameters, this is a meaningful gap. The behavioral details are helpful, but the agent still lacks enough context to invoke it confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must explain the parameters, but it does not. It never clarifies that 'instruction' likely carries the new constraint, what 'repo' is for, or the role of 'mission_id' beyond being required. The agent is left to guess parameter meanings from names alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('replan') and resource ('remaining work'), and indicates the trigger condition ('around a new constraint'). It is clearly distinct from sibling tools like mission_plan or mission_steer by focusing on replanning around a changed constraint.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies a clear usage context: use when a new constraint forces replanning of remaining work. It does not explicitly enumerate alternatives or when not to use the tool, but the condition is specific enough to guide selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mission_resumeA
Resume a paused/blocked mission, reconciling repository and worker state.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | ||
| mission_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the transparency burden. It discloses one meaningful behavior—state reconciliation—but does not say whether this mutates data, what gets overwritten, or what preconditions must hold. This is a notable gap for a resume operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence that leads with the core action and appends a meaningful clause. There is no filler, and every phrase earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema covers return-value expectations, but with no annotations and a large sibling set, the description should provide more context on side effects and preconditions. The purpose and target state are clear, yet the behavioral gaps prevent a higher score.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain the required mission_id or the optional repo parameter. It only alludes to 'repository... state' indirectly, leaving an agent to infer parameter meaning from names alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb, 'Resume,' against a specific resource, 'paused/blocked mission,' and adds a distinguishing mechanism: 'reconciling repository and worker state.' This clearly differentiates it from siblings like mission_run or mission_retry_feature, even without naming them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'paused/blocked mission' gives clear context for when this tool is appropriate. However, it does not explicitly mention exclusions or alternative sibling tools, so the guidance stops short of full routing instructions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mission_retry_featureA
Manually retry a failed/blocked feature. The retry counts toward the feature's attempt budget (attempts + manual_retries >= max_attempts stops further automated retries).
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | ||
| feature_id | Yes | ||
| mission_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the behavioral disclosure burden. It does well by disclosing that a manual retry counts toward the attempt budget and that attempts + manual_retries >= max_attempts stops further automated retries. It does not cover other side effects like reversibility or permission requirements, but the budget impact is a meaningful disclosure.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences: the first states the action and target, the second provides the key behavioral constraint. Every sentence earns its place, and the most important information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is adequately described for basic selection, and the budget side effect is a valuable addition. However, with no annotations, no parameter documentation, and no explicit routing against siblings like mission_run or mission_complete_feature, the description leaves notable gaps for an agent deciding when and how to invoke it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage for its parameters, so the description must compensate, but it does not explain mission_id, feature_id, or repo. The word 'feature' gives minimal context for feature_id, but mission_id and repo remain effectively undocumented beyond their names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('retry') and resource ('failed/blocked feature'), making the tool's core purpose unmistakable. The 'manual' qualifier and failed/blocked scope distinguish it from run/start tools and from automated retry behavior.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly states when to use the tool: for a failed or blocked feature. It also explains how manual retries interact with the automated retry budget, which implies when manual intervention is appropriate. However, it does not explicitly name sibling alternatives or state when not to use it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mission_runA
Advance the mission control loop (schedule, dispatch isolated workers, validate, correct). Resumable - call again to continue an in-progress mission. max_steps counts DISPATCHES; auto-retries inside a dispatch consume the feature's attempt budget, not a step. If your client sandbox times this call out, the loop keeps running server-side: poll mission_events instead of re-issuing. The response ends with a RUN REPORT (dispatches done, queued work) plus a STATE proof line.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | ||
| max_steps | No | ||
| mission_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden, and it covers the key runtime behaviors: server-side continuation after client timeout, step-count semantics (dispatches vs retry budget), resumability, and report/proof output. This is exactly the kind of operational nuance an agent needs before calling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Each sentence conveys a distinct, high-value fact: purpose, resumability, step semantics, timeout behavior, and response shape. The information is front-loaded and dense without being rambling.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex control-loop tool with no annotations and many siblings, the description is unusually complete: it covers the loop's purpose, continuation, timeout safety, step counting, and response structure. It does not elaborate on the optional repo parameter or map every sibling relationship, but the core selection and invocation guidance is present.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description provides essential semantics for max_steps, clarifying that it counts dispatches and does not consume the feature attempt budget. mission_id and repo are recognizable from their titles, so the only genuinely ambiguous parameter is explained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action ('Advance the mission control loop') and enumerates its inner steps, clearly identifying this as the main driving tool rather than a single validation or retry tool. It distinguishes mission_run from siblings like mission_validate and mission_retry_feature by framing them as parts of the loop it advances.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states that the call is resumable and that an in-progress mission should be continued by calling again. It also gives a concrete when-not-to-reissue rule: after a client-side timeout, poll mission_events instead, preventing duplicate dispatch.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mission_statusC
Mission control status for a mission.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | ||
| mission_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. The word 'status' hints at a read-only operation, but the description does not explicitly state side effects, permissions, prerequisites, or any operational behavior beyond the vague notion of status.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no padding, so it is concise and front-loaded. However, it is under-specified to the point of being barely informative, missing a verb and any operational detail.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The output schema covers return values, but with a 0% parameter coverage, no annotations, and no guidance differentiating this tool from many similar mission tools, the description is not complete enough for an agent to confidently select and invoke it. Key semantic gaps remain around what 'status' means and how this tool relates to mission_readiness, mission_events, or mission_list.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not mention either parameter, mission_id or repo. Even the phrase 'for a mission' only weakly implies a mission identifier and provides no meaning for the optional repo parameter or the required mission_id.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description is essentially a noun phrase, 'Mission control status for a mission,' which largely restates the tool name without specifying an action such as retrieving, returning, or querying status. It also does not distinguish this tool from closely related siblings like mission_readiness, mission_events, or mission_list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus any of the 18 sibling tools. The description gives no context about expected use cases, alternatives, or exclusions, leaving the agent to guess whether 'status' means health, progress, readiness, or something else.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mission_steerB
Steer a running mission without rebuilding the plan. Optionally reprioritize a specific feature by id and priority.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | ||
| priority | No | ||
| feature_id | No | ||
| mission_id | Yes | ||
| instruction | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must disclose what changes occur. It only says 'steer' and 'reprioritize', without explaining side effects, persistence, whether instructions are applied, or operational constraints. It provides some behavioral clues but is insufficient for a mutating operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler. The first sentence gives the core behavior and contrast, the second adds the optional parameter relationship. Every word earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool mutates a running mission and has a free-form instruction parameter, yet the description does not explain what 'instruction' does, what 'repo' scopes, or what a successful steer entails. The output schema exists, but key input semantics remain ambiguous.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must carry parameter meaning. It clarifies feature_id and priority together, and implies mission_id, but leaves instruction and repo entirely unexplained. Given five parameters and zero schema-level description, this is a meaningful gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly ties the tool name to a meaningful action: steering a running mission without rebuilding the plan, and explicitly mentions optional reprioritization of a feature. This distinguishes it from mission_replan and from feature-level tools like mission_retry_feature, though 'steer' remains slightly broad.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'without rebuilding the plan' implies a contrast with replanning alternatives, and 'running mission' indicates it applies to an active mission. However, it never explicitly says when to use this tool versus mission_replean, mission_adjust_suggestions, or the feature-completion siblings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mission_validateB
Run a validator now (scrutiny | user_testing | final). Returns a structured verdict.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | scrutiny | |
| repo | No | ||
| mission_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of disclosing side effects and behavior. It says a validator is run and a verdict returned, but it does not disclose whether this mutates mission state, creates validation records, requires a particular mission status, or has blocking/performance implications.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence with no filler. It efficiently packs the action, the variant choices, and the return type, making it easy to scan while remaining appropriately concise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with three parameters, no annotations, and no output schema, the description leaves important gaps: mission_id and repo semantics, which validator kind to choose, what the structured verdict contains, and what state changes may occur. An agent would need external workflow knowledge to invoke this confidently.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% description coverage, so the description must compensate. It partially explains the 'kind' parameter by listing its possible values, but 'mission_id' and 'repo' are left semantically unexplained, including what 'repo' refers to and how it affects validation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific action ('Run a validator now'), identifies the resource (validator), enumerates the supported variants (scrutiny | user_testing | final), and tells the agent what to expect ('structured verdict'). This is specific enough to distinguish it from siblings like mission_approve or mission_run.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is given about when to run a validator versus using alternatives like mission_readiness or mission_approve, nor when to choose scrutiny over user_testing or final. The word 'now' implies on-demand execution, but prerequisites, exclusions, and workflow placement are absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
19 tool updates
v0.2.0- First observed
mission_apply_suggestions - First observed
mission_approve - First observed
mission_cancel - First observed
mission_complete_feature - First observed
mission_create - First observed
mission_events - First observed
mission_gc - First observed
mission_list - First observed
mission_pause - First observed
mission_plan - First observed
mission_plan_view - First observed
mission_readiness - First observed
mission_replan - First observed
mission_resume - First observed
mission_retry_feature - First observed
mission_run - First observed
mission_status - First observed
mission_steer - First observed
mission_validate
TDQS
Scored across 19 tools
The tools map cleanly to lifecycle stages (create/plan/approve/run/pause/resume/steer/replan/cancel) and read-only views (status/events/readiness/plan_view). A few control-loop actions like mission_run, mission_resume, and mission_steer could be confused, but their descriptions clarify the differences.
All tools share the mission_ prefix and use lower_snake_case, but the pattern mixes bare verbs (mission_run, mission_pause), noun-like resources (mission_status, mission_events), and compound verb_noun forms (mission_retry_feature, mission_apply_suggestions). This is readable and predictable, though not as uniform as a strict verb_noun convention.
At 19 tools, the surface is on the heavy side for an MCP server. Most tools are individually justified by mission lifecycle needs, but a few niche helpers like mission_apply_suggestions and mission_gc, plus several status-like queries, could plausibly be consolidated.
The toolset covers the full mission lifecycle: create, plan, approve, run, pause, resume, steer, replan, cancel, plus status/events/readiness/plan_view and manual retry/completion/validation paths. There are no obvious dead ends for the stated purpose, and external-completion and cleanup workflows are included.
Maintenance
Related MCP Connectors
Durable, user-controlled goals and governed plans for AI agents.
AI work orchestration for plans, tasks, teams, and coding-agent dispatch.
Goal and task planning MCP for Codex and AI agents, with evidence-backed completion.
Control plane for autonomous software labor. Agents claim objectives over MCP with audit trail.
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceTurns Claude Code into an engineering project manager by orchestrating feature development with isolated Git worktrees, structured task validation, and approval-gated integration.1MIT
- FlicenseNot gradedqualityBmaintenanceEnables safe, isolated Codex implementation runs with planning approval, verification, and bounded fixes, without merging or pushing code automatically.-
- FlicenseBqualityBmaintenanceEnables AI coding assistants to run a machine-verified DESIGN→PLAN→EXECUTE→VERIFY→COMPLETE workflow with human approval gates, state integrity checks, and DAG task scheduling.7-
- AlicenseNot gradedqualityAmaintenanceEnables bounded, multi-phase coding workflows inside Codex that plan, implement, independently review, repair, and verify repository changes, with an inline dashboard for inspecting runs.MIT