Skip to main content
Glama

🦆 Factory-Droid-style Mission orchestration for Goose

PyPI version Python versions PyPI downloads npm version License: MIT Goose extension Code style: deterministic core

GitHub stars

Type a goal → get a structured plan → approve → watch isolated workers build it, get validated, get corrected — until it's done and proven.

USER GOAL → ANALYSIS → STRUCTURED PLAN → FEATURES + DEPS + MILESTONES → APPROVAL
  → DEPENDENCY-AWARE EXECUTION (isolated workers) → REAL CODE
  → SCRUTINY + USER-FACING VALIDATION → AUTOMATIC CORRECTION
  → FINAL VALIDATION → MISSION COMPLETED ✅

hamgoose is a genuine Goose extension — a standalone stdio MCP server on the official mcp/FastMCP model, not a recipe, todo wrapper, or delegation prompt. Code enforces the orchestration mechanics; models do the semantic reasoning.


📖 Contents


Related MCP server: Codex Orchestrator

✨ Features

🚦 Approval gate

Nothing is implemented until you approve the plan

🏝️ Isolated leaf workers

Each feature runs in its own goose subprocess inside a Git worktree — no nested delegation, crash containment, real diffs

🕸️ Dependency-aware scheduling

A DAG of features with path-overlap conflict detection and a hard concurrency cap (your provider's limits, enforced in code)

🔍 Two validators

Scrutiny distrusts the worker's claims and inspects diff + tests; user-testing exercises the app from the user's perspective

🔁 Automatic correction

Failed validation becomes corrective features; the bounded loop repeats until the milestone passes

🧯 Crash recovery

Atomic JSON state + append-only event log — kill Goose mid-mission, reopen, it reconciles and continues

📡 Live progress

Long calls stream MCP progress notifications, and the guided flow runs in short visible bursts — you always see movement, never a silent wait

🧭 Steering & replanning

Change course mid-mission without losing completed work

🔐 Secrets redacted

Every persisted artifact scrubbed of keys, tokens, credentials

🪶 Per-repo state

Lives in <repo>/.goose/hamgoose/ — nothing global, trivially git-ignored

🚀 Quickstart

Node — via npm:

npm i -g @cooked-ham/hamgoose
hamgoose register

Python — via pip:

pip install git+https://github.com/cooked-ham/hamgoose.git   # Python 3.11+
hamgoose register

Either world, one shot (no install): npx @cooked-ham/hamgoose register

TIP

That's the whole install — two commands, no repo wiring, no config surgery. Pickone channel (npm or pip); both provide the same hamgoose command. Uninstall: hamgoose unregister + npm uninstall -g @cooked-ham/hamgoose (or pip uninstall -y hamgoose).

🎯 The walkthrough

Then, in any repository you're working in:

$ goose
You:   /prompt start_mission
goose: What's the goal?
You:   Migrate the auth module from session cookies to JWT.
goose: Any rules or constraints? (concurrency, provider/model, git, validation)
You:   My provider only allows 3 concurrent agents at a time.
goose: Plan: 2 milestones, 6 features, workers capped at 3 concurrent. Approve?
You:   Approve.
goose: MS01 1/3 … passed scrutiny … MS02 2/3 …
       ✅ Mission COMPLETED — changes on branch mission/base with per-feature commits.

Rules are recorded verbatim on the mission (visible in every status and plan view), translated into execution config ("max 3 concurrent" → max_concurrent_workers: 3), and handed to every worker as context. Mid-mission you can just say "pause", "don't touch config files", or "replan around X" — steering and replanning never lose completed work.

NOTE

No slash command? Just say"start a hamgoose mission" in plain English. /prompt start_mission runs the guided setup; plain English does the same thing.

🏗️ How it works

flowchart LR
    U["👤 You<br/>goal + rules"] --> G["Goose session"]
    G <-->|MCP stdio| H["🦆 hamgoose<br/>orchestrator<br/>(deterministic code)"]
    H -->|isolated goose run| W1["Worker F001<br/>🌳 worktree"]
    H -->|isolated goose run| W2["Worker F002<br/>🌳 worktree"]
    H --> V["🔍 Validators<br/>scrutiny + user-test"]
    W1 -->|merge + commit| R[("repo<br/>mission/base")]
    W2 -->|merge + commit| R
stateDiagram-v2
    [*] --> CREATED
    CREATED --> ANALYZING
    ANALYZING --> PLANNING
    PLANNING --> AWAITING_APPROVAL
    AWAITING_APPROVAL --> RUNNING : approve
    RUNNING --> PAUSED
    RUNNING --> BLOCKED
    PAUSED --> RUNNING : resume
    BLOCKED --> RUNNING : resolve + resume
    RUNNING --> VALIDATING
    VALIDATING --> RUNNING : corrective work
    VALIDATING --> COMPLETED : all pass ✅
    CREATED --> CANCELLED
    AWAITING_APPROVAL --> CANCELLED
    RUNNING --> FAILED
    COMPLETED --> [*]
    FAILED --> [*]
    CANCELLED --> [*]

Why it's different from "just let the agent do it": the orchestrator is deterministic code — state machines, DAG scheduling, retries, Git bookkeeping, persistence are enforced, not hoped for. The LLM only does what LLMs are good at: understanding intent and writing code. A confused model can't corrupt the mission state, skip the approval gate, or double-dispatch a feature. See ARCHITECTURE_REPORT.md for the full design analysis.

📦 Install

Requires goose (≥ 1.40) on your PATH and git (for Git missions).

Option

For

Command

1. From GitHub ⭐

Everyone

pip install git+https://github.com/cooked-ham/hamgoose.git then hamgoose register

2. Goose's own menu

No extra commands

goose configure → Extensions → Add Extension → STDIO / hamgoose / hamgoose

3. From a clone

Contributors

git clone … && cd hamgoose && uv venv .venv && uv pip install -p .venv -e . then hamgoose register

4. Per-run

One-off experiments

goose run -t "..." --with-extension "hamgoose:python -m hamgoose"

5. npm (Node world)

Node-first machines

npm i -g @cooked-ham/hamgoose then hamgoose register — or one-shot: npx @cooked-ham/hamgoose register

IMPORTANT

Pin a release once tags exist:pip install "git+https://github.com/cooked-ham/hamgoose.git@v0.1.2".

# config.yaml — path printed by `goose info`
extensions:
  hamgoose:
    enabled: true
    type: stdio
    name: hamgoose
    description: Mission orchestration for Goose
    cmd: hamgoose
    args: []

🛠️ The lifecycle (tools)

Operation

Tool

Create + analyze repo (guided setup)

mission_create(goal, rules?, config?)

Generate the plan (approval gate)

mission_plan(mission_id)

Approve & start

mission_approve(mission_id)

Execute the control loop (resumable)

mission_run(mission_id, max_steps?)

Pause / resume

mission_pause / mission_resume

Steer (priority / guidance)

mission_steer(instruction, feature_id?, priority?)

Replan (new constraint)

mission_replan(instruction)

Retry a feature / validate now

mission_retry_feature / mission_validate(kind)

Cancel

mission_cancel

Read status / plan / events / list

mission_status / mission_plan_view / mission_events / mission_list

Resources (read): mission://{id}/status|plan|events|features|milestones|validation Prompts (list with /prompts, run with /prompt <name>): start_mission (guided setup) · plan_mission · resume_mission · validate_milestone

🗄️ Where state lives

<repo>/.goose/hamgoose/<mission-id>/
├── mission.json     # canonical atomic state
├── mission.yaml     # human-readable mirror
├── plan.md          # plan mirror
├── events.jsonl     # append-only event log
├── workers/         # redacted worker transcripts
├── validation/      # validation reports
├── worktrees_base/  # mission/base worktree (merged result)
└── worktrees/<F>    # per-feature worktrees

Your current branch is never modified — mission/base accumulates the merged, validated result for you to merge. Add /.goose/hamgoose/ to your repo's .gitignore.

🧪 Development

git clone https://github.com/cooked-ham/hamgoose.git && cd hamgoose
uv venv .venv
uv pip install -p .venv -e ".[dev]"
.venv/bin/python -m pytest -m "not realgoose"   # fast, deterministic (no LLM)
.venv/bin/python -m pytest -m "realgoose"       # real Goose + LLM (slower)

📚 Docs

📜 License

MIT — see LICENSE.


Built on Goose · MCP · 🦆

Available Tools

19 tools
mission_apply_suggestionsA

H10: apply the config deltas recorded at mission create when the model preflight flagged the worker model (e.g. worker_timeout>=900 for SMALL-OUTPUT-BUDGET models). One call; the applied values are echoed.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoNo
mission_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Since no annotations are provided, the description carries the burden of behavioral disclosure. It conveys that the tool mutates mission configuration, that it applies pre-recorded deltas, that it is meant to be called once, and that the applied values are echoed. It does not mention side effects, reversibility, or the no-delta case, but the core behavior is transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and action-first, with no filler. The cryptic 'H10:' prefix is unexplained and adds little value, which prevents a perfect score, but the rest is tightly written and informative.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the trigger, the action, the one-call behavior, and the echoed output, which is good for a focused mutation tool. However, it omits any explanation of the repo parameter and does not situate this tool against the many mission_* siblings, leaving some context incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0% description coverage, and the description does not explain the purpose or format of mission_id or repo. The mission_id is inferable from context and the sibling tool naming, but repo is entirely undocumented. With such low schema coverage, the description needed to compensate and did not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('apply the config deltas recorded at mission create') on a concrete resource, with a triggering condition (model preflight flagged the worker model). It is clearly distinct from the sibling mission_* tools, which handle creation, planning, running, and validation rather than applying recorded suggestions.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives a clear when-to-use condition, including a concrete example (worker_timeout>=900 for SMALL-OUTPUT-BUDGET models). It does not explicitly name alternatives or state when not to use the tool, so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mission_approveA

Approve the plan and begin dependency-aware execution. Safe to call only once, from AWAITING_APPROVAL.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoNo
mission_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It discloses that the tool is non-idempotent ('Safe to call only once') and state-dependent ('from AWAITING_APPROVAL'), and that it triggers dependency-aware execution. It does not detail all side effects, but the key transition behavior is transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, tight sentence with no filler. The action is stated first, and the critical safety constraint is included in the same sentence. Every word contributes.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a state-transition tool with an output schema and a simple input schema, the description covers the core semantics and the key precondition. The main gap is the unexplained optional repo parameter, but the tool is otherwise sufficiently specified for an agent in the AWAITING_APPROVAL state.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not mention either parameter. mission_id is reasonably inferable from the tool name, but repo is left entirely unexplained, and the description does not clarify how either parameter affects the approval or execution behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Approve') and resource ('the plan'), and clearly states the resulting behavior ('begin dependency-aware execution'). It also distinguishes this tool from siblings like mission_plan or mission_validate by tying it to the approval transition.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states when it is safe to call: 'only once, from AWAITING_APPROVAL.' This gives a clear precondition and cautions against repeated calls, though it does not explicitly name alternative tools or the action to take when not in that state.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mission_cancelC

Cancel a mission and clean up its worktrees.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoNo
mission_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries full responsibility for behavioral disclosure. It does reveal a key side effect: cancellation plus worktree cleanup, which signals destructiveness. However, it does not state whether the operation is irreversible, what happens to running processes, or what cleanup entails in practice.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single efficient sentence with no filler. The core action is front-loaded, and the cleanup behavior is mentioned immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the destructive nature of cancellation and an undocumented optional repo parameter, this minimal description leaves important context missing. The output schema exists, so return values need not be described, but the tool's effect on mission state and worktrees is only vaguely conveyed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description adds no meaning to the parameters. It does not explain that mission_id identifies the mission to cancel, nor does it clarify the optional repo parameter's role in locating or cleaning up worktrees. The agent must guess at the intended semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Cancel') and a clear resource ('mission'), and adds a distinctive behavioral detail ('clean up its worktrees'). It is reasonably distinguishable from sibling tools like mission_pause or mission_resume, though it does not explicitly differentiate itself.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance on when to use this tool versus alternatives such as mission_pause or mission_complete_feature. There is no mention of prerequisites, when cancellation is appropriate, or when a different mission lifecycle tool should be chosen.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mission_complete_featureA

Record work on a feature that was implemented OUTSIDE the worker pipeline (by you, the lead agent, or a human). Use this instead of editing mission state files by hand. Verifies the commit exists, runs the feature's validation_commands, runs a real scrutiny validation on the diff, appends proper events, and continues the normal milestone flow.

Args: summary: what was implemented and why it meets the acceptance criteria. commit: the git hash of the implementation (required for git missions). changed_files: JSON list of paths (auto-derived from the commit if omitted). tests: JSON list of verification commands you ran and their outcome.

NOTE (H1): this runs a REAL scrutiny validation and can take minutes. If your client sandbox times the call out, the completion still proceeds server-side - verify via mission_events / mission_status instead of re-calling blindly.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoNo
testsNo
commitNo
summaryYes
feature_idYes
mission_idYes
changed_filesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden, and it does so excellently. It discloses that the tool verifies the commit, runs validation_commands, executes a real scrutiny validation, appends events, and continues milestone flow. It also warns about long runtime, sandbox timeouts, and server-side completion, which is high-value behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured: purpose first, then parameter guidance, then a critical timeout note. Every sentence earns its place, and the most important warnings are front-loaded. The note about not re-calling blindly is especially useful and not redundant.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations and no schema-level descriptions, this is unusually complete. It covers when to use it, what side effects occur, parameter semantics, and failure-mode behavior. The output schema exists, so not describing return values is acceptable.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate, and it adds meaningful detail for summary, commit, changed_files, and tests. It explains formats, defaults, and the special requirement that commit is needed for git missions. It does not explicitly explain repo or the required mission_id/feature_id, but those are largely self-explanatory from their names and schema types.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: record work on a feature that was implemented outside the worker pipeline, and it contrasts this with editing mission state files by hand. It is specific about the verb and resource, and the context makes it distinguishable from generic mission tools. However, it does not explicitly differentiate itself from sibling tools such as mission_retry_feature by name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives clear context for when to use the tool: when a feature was implemented outside the worker pipeline, and it explicitly advises using this instead of hand-editing mission state files. It does not, however, mention exclusions or directly compare against sibling tools like mission_validate or mission_retry_feature.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mission_createA

Create a hamgoose mission. Use this to START a mission for the user.

Guided setup protocol:

  1. If the user has not stated a clear goal, ask for it (one short question).

  2. Ask ONCE whether they have rules/constraints worth recording (e.g. concurrency limits, provider/model for workers, git on/off, validation toggles). If they say none, proceed with defaults - do not interrogate.

  3. Pass their rules VERBATIM in rules (persisted on the mission; shown in plan and status; given to every worker as context).

  4. Translate rules into config overrides with this map:

    • "max N concurrent agents/workers/subagents" -> {"execution": {"max_concurrent_workers": N}}

    • "one worker at a time" / "sequential" -> {"execution": {"max_concurrent_workers": 1}}

    • "same for validators" -> {"validator": {"provider": ..., "model": ...}}

    • "use / for planning" -> {"planner": {"provider": ..., "model": ...}}

    • "no git / no worktrees" -> {"git": {"enabled": false, "use_worktrees": false}}

    • "skip user-facing testing" -> {"validation": {"user_testing": false}}

    • "no scrutiny validation" -> {"validation": {"scrutiny": false}} Note: workers are always isolated goose run leaf processes (never nested delegation); max_concurrent_workers caps how many run simultaneously.

Returns the mission id, a readiness report and next steps. Next: mission_plan, present the plan, get approval, mission_approve, mission_run.

ParametersJSON Schema
NameRequiredDescriptionDefault
goalYes
repoNo
rulesNo
configNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are present, so the description carries the burden. It discloses that rules are persisted, shown in plan/status, and given to every worker, and that workers are isolated 'goose run' leaf processes capped by max_concurrent_workers. It also states that the response includes mission id, readiness report, and next steps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but organized into a numbered setup protocol with a config override map and a workflow handoff. It front-loads the core purpose, and nearly every sentence carries operational value, though it is denser than necessary.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is nearly complete for a mission-creation tool: it covers goal discovery, rule handling, config translation, worker behavior, and next steps, and mentions return contents even though an output schema exists. The main omissions are `repo` semantics and explicit defaults, leaving minor ambiguity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

With 0% schema coverage, the description compensates substantially for `rules` and `config` by requiring verbatim rules and providing a clear natural-language-to-config map. However, `repo` is never explained, and `goal` is only implied, leaving two of four parameters with little semantic guidance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States 'Create a hamgoose mission' and explicitly says 'Use this to START a mission for the user,' making the resource and verb concrete. The 'Next: mission_plan...' handoff distinguishes it from sibling tools like mission_plan and mission_run.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The guided setup protocol gives clear when-to-do steps: ask for the goal if absent, ask once about rules, proceed with defaults if none. It also routes to the follow-up workflow. It does not explicitly state when not to use this tool in favor of an alternative, but the workflow context is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mission_eventsC

Recent mission events (append-only event log).

ParametersJSON Schema
NameRequiredDescriptionDefault
repoNo
tailNo
mission_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the burden of behavioral disclosure. 'Append-only event log' usefully suggests that events are immutable and reading them is safe, but it does not disclose ordering, tail semantics, repo filtering, or whether the tool is strictly read-only.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short, front-loaded, and contains no filler words. It is efficiently phrased, though its brevity means it omits operational details that other dimensions must compensate for.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations, 0% schema description coverage, and a large sibling set, the description is too minimal. It does not explain how tail works, what repo does, what event ordering to expect, or how this differs from mission_status or mission_plan_view.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description mentions none of the three parameters. mission_id, tail, and repo remain semantically undocumented beyond their names and types, so the description adds no value for parameter understanding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The phrase 'Recent mission events' clearly identifies the resource being accessed, and 'append-only event log' distinguishes this from mission planning, execution, and retry siblings. It lacks an explicit verb like 'retrieve' or 'list,' but the meaning is strongly implied.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies this tool is for reading recent event history for a mission, especially given the 'append-only event log' phrasing. However, it provides no explicit guidance about when to use this tool versus alternatives or when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mission_gcA

H11 housekeeping: list terminal and long-stale missions that clutter mission_list. With archive=true, non-terminal stale missions older than max_age_days are cancelled (their data and event history are kept). Terminal missions are never re-touched.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoNo
archiveNo
max_age_daysNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden. It transparently discloses that archive=true cancels non-terminal stale missions older than max_age_days while preserving data and event history, and that terminal missions are never modified. Core destructive behavior is clear, though ancillary details like permissions or error behavior are absent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three compact sentences with the main action front-loaded in the first clause. It contains no filler, and each sentence adds distinct information: the listing purpose, the conditional archive behavior, and the terminal-mission guard.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that an output schema exists and no annotations are present, the description covers the core behavior, side effects, and the terminal-mission exception well enough for typical calls. The main gap is the undocumented repo parameter, which would require an agent to infer its meaning from the name alone. Overall, it is a fairly complete definition for a tool of this complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for parameter meaning. It does explain archive and max_age_days via 'archive=true' and 'older than max_age_days', but it never explains the repo parameter. This is partial compensation: two of three parameters receive meaningful semantic grounding.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description immediately identifies the tool as 'H11 housekeeping' and states a specific action: list terminal and long-stale missions cluttering mission_list, with archive mode cancelling eligible non-terminal missions. It clearly pairs a verb with a resource and distinguishes itself from normal mission listing by the cleanup/housekeeping framing. The terminal-mission exclusion further sharpens the scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The 'H11 housekeeping' prefix gives useful context, and the archive=true conditional plus 'Terminal missions are never re-touched' clarify when mutation happens and what is excluded. However, the description does not explicitly name sibling alternatives such as mission_list or mission_cancel, nor does it say when to prefer them. Usage guidance is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mission_listB

List all missions in the repository.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. 'List' strongly implies a read-only operation and 'all missions in the repository' conveys scope, but it does not mention optional repo behavior, pagination, permissions, or any other side-effect details. The output schema covers return structure, so that omission is less critical.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, clear, front-loaded sentence with no filler. It is appropriately concise for a simple list operation, though it relies on the schema and output schema for additional detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description plus the available output schema provide a minimally usable picture for a simple list tool. However, the optional repo parameter is unexplained, and the lack of sibling differentiation means an agent cannot fully determine when mission_list is the correct choice among the many mission_* tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage for the optional repo parameter, and the description does not explain how to use it, what values are valid, or what the default of null means. The word 'repository' provides a loose hint that the repo parameter identifies the repository, but it does not compensate for the missing parameter documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the operation ('List') and the resource ('missions'), and adds the scoping phrase 'all ... in the repository'. It is distinguishable from the sibling tools by the list verb, though it does not explicitly contrast with related tools like mission_status or mission_plan_view.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no guidance about when to choose this tool over the many sibling mission_* tools, and it mentions no alternatives or exclusions. The imperative 'List' implies a use case, but the agent is left to infer when this is the right tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mission_pauseB

Pause an active mission. No new workers are launched while paused.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoNo
reasonNo
mission_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral burden. It discloses one concrete effect: 'No new workers are launched while paused.' However, it does not clarify whether currently running workers are stopped, whether the pause is reversible, or what state changes occur beyond the pause flag.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences provide the core action and a key behavioral consequence without any filler. The description is front-loaded and every sentence earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a relatively simple pause operation, the description covers the main purpose and the most important behavioral nuance. An output schema exists, so return-value details are not required. It is slightly incomplete in that it leaves the repo parameter unexplained and does not mention how pause relates to resume, but the essentials for invocation are present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not explain any parameter. The mission_id parameter is inferable from the tool's purpose, and reason is partially self-explanatory, but repo remains ambiguous. The description fails to compensate for the lack of parameter documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific action and resource: 'Pause an active mission.' It also adds a meaningful behavioral qualifier about worker launches. It does not explicitly name sibling tools like mission_resume or mission_cancel, but the pause action is distinct enough that an agent can likely select it correctly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage: pause an active mission when you want to stop new workers from launching. It gives no explicit guidance about when not to use it or how it compares to alternatives such as mission_cancel or mission_resume, leaving the agent to infer the context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mission_planA

Generate the structured dependency-aware plan and present it for approval. No implementation happens until mission_approve is called.

If the planner returns an empty plan (goal too vague), retry mission_plan - or pass your OWN decomposition (JSON string or list): features='[{"id":"F001","title":"...","description":"...","milestone":"MS01", "dependencies":[],"acceptance_criteria":["..."],"expected_paths":["..."]}]' milestones='[{"id":"MS01","objective":"...","completion_criteria":["..."]}]'.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoNo
featuresNo
milestonesNo
mission_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It explicitly discloses that the tool does not implement anything, that it presents the plan for approval, and that the planner may return an empty plan when the goal is too vague. This is good disclosure, though it omits any permission or side-effect details.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The primary behavior and approval gating are front-loaded in the first sentence. The JSON examples are long but necessary given the 0% schema coverage, and each paragraph contributes distinct information. Slightly verbose, but efficiently organized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the core workflow, the approval gate, recovery from empty plans, and the shape of optional decomposition inputs. Since an output schema exists, return-value details are not needed. Minor gaps remain around repo/mission_id semantics and how this relates to mission_replan, but the description is largely complete for invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It provides valuable JSON examples for features and milestones, including nested field names and dependency structures. It does not explicitly explain the required mission_id or optional repo parameter, but their purpose is reasonably inferable and the most complex parameters are well covered.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Generate the structured dependency-aware plan and present it for approval.' It clearly distinguishes the tool from sibling actions by stating no implementation happens until mission_approve is called.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives concrete usage context: use this to produce an approval-ready plan, and use mission_approve for implementation. It also spells out a fallback for empty planner results (retry or supply own decomposition). It does not exhaustively contrast all sibling tools, but the main routing cues are present.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mission_plan_viewC

The current plan for a mission.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoNo
mission_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It implies a read operation by saying 'current plan' but does not disclose side effects, freshness, error behavior, or whether it modifies anything.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness2/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is short and front-loaded, but this is under-specification rather than effective conciseness. One sentence states only the core purpose and omits necessary contextual information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With two parameters, no annotations, an output schema, and many sibling tools, the description is incomplete. It does not provide selection criteria, parameter semantics, or behavioral expectations, though the output schema mitigates the need to document return values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the description mentions no parameters. It does not clarify the meaning or usage of mission_id or the optional repo parameter, so an agent cannot determine parameter semantics from the description.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies the resource ('the current plan for a mission') and the tool name supplies the verb 'view'. It is clear enough to distinguish from related tools like mission_plan and mission_replan, though it does not explicitly contrast them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance about when to use this tool versus mission_status, mission_plan, mission_replan, or other siblings. The intended selection context is left entirely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mission_readinessC

Readiness/preflight report for a mission.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoNo
mission_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description carries the full burden of behavioral disclosure. 'Readiness/preflight report' suggests a non-mutating check, but it does not explicitly state that no mission state is changed, what the report contains, or whether any side effects occur. The description is too thin for a tool with no annotation support.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very short and front-loaded, with no wasted words. However, it is concise partly because it omits necessary detail, so it cannot receive a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value structure may not need to be described. Still, there is no indication of what a readiness report contains, what conditions are checked, when to use this tool, or how the optional repo parameter affects behavior. The definition is too minimal to fully support correct tool selection and invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not compensate by explaining the meaning or purpose of mission_id or repo. While mission_id is somewhat self-explanatory, the optional repo parameter is completely unexplained. The description adds no value beyond the schema's field names and types.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies the resource (mission) and the kind of output (readiness/preflight report), which is reasonably clear even though it lacks an explicit verb like 'generate' or 'return'. It is not a tautology and is distinguishable enough from sibling tools like mission_status or mission_validate.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The word 'preflight' implies this tool should be used before a mission run, suggesting a precondition or readiness check. However, it does not explicitly state when to use it versus siblings such as mission_validate, mission_status, or mission_run, nor does it mention exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mission_replanA

Replan the remaining work around a new constraint. Preserves valid completed work, marks invalidated work superseded, and bumps the plan revision.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoNo
mission_idYes
instructionNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden of disclosing behavior. It does this well by stating that valid completed work is preserved, invalidated work is marked superseded, and the plan revision is bumped. It could add more about side effects or irreversibility, but the key behavioral traits are transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact, front-loaded with the main action, and every sentence adds value. It states the core operation and then key behavioral outcomes without filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Although an output schema exists, the description omits essential parameter semantics and provides no explicit guidance on how to phrase the replan request or which parameters matter. For a tool with three undocumented parameters, this is a meaningful gap. The behavioral details are helpful, but the agent still lacks enough context to invoke it confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must explain the parameters, but it does not. It never clarifies that 'instruction' likely carries the new constraint, what 'repo' is for, or the role of 'mission_id' beyond being required. The agent is left to guess parameter meanings from names alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('replan') and resource ('remaining work'), and indicates the trigger condition ('around a new constraint'). It is clearly distinct from sibling tools like mission_plan or mission_steer by focusing on replanning around a changed constraint.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies a clear usage context: use when a new constraint forces replanning of remaining work. It does not explicitly enumerate alternatives or when not to use the tool, but the condition is specific enough to guide selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mission_resumeA

Resume a paused/blocked mission, reconciling repository and worker state.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoNo
mission_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the transparency burden. It discloses one meaningful behavior—state reconciliation—but does not say whether this mutates data, what gets overwritten, or what preconditions must hold. This is a notable gap for a resume operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence that leads with the core action and appends a meaningful clause. There is no filler, and every phrase earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema covers return-value expectations, but with no annotations and a large sibling set, the description should provide more context on side effects and preconditions. The purpose and target state are clear, yet the behavioral gaps prevent a higher score.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not explain the required mission_id or the optional repo parameter. It only alludes to 'repository... state' indirectly, leaving an agent to infer parameter meaning from names alone.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb, 'Resume,' against a specific resource, 'paused/blocked mission,' and adds a distinguishing mechanism: 'reconciling repository and worker state.' This clearly differentiates it from siblings like mission_run or mission_retry_feature, even without naming them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'paused/blocked mission' gives clear context for when this tool is appropriate. However, it does not explicitly mention exclusions or alternative sibling tools, so the guidance stops short of full routing instructions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mission_retry_featureA

Manually retry a failed/blocked feature. The retry counts toward the feature's attempt budget (attempts + manual_retries >= max_attempts stops further automated retries).

ParametersJSON Schema
NameRequiredDescriptionDefault
repoNo
feature_idYes
mission_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the behavioral disclosure burden. It does well by disclosing that a manual retry counts toward the attempt budget and that attempts + manual_retries >= max_attempts stops further automated retries. It does not cover other side effects like reversibility or permission requirements, but the budget impact is a meaningful disclosure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two concise sentences: the first states the action and target, the second provides the key behavioral constraint. Every sentence earns its place, and the most important information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is adequately described for basic selection, and the budget side effect is a valuable addition. However, with no annotations, no parameter documentation, and no explicit routing against siblings like mission_run or mission_complete_feature, the description leaves notable gaps for an agent deciding when and how to invoke it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0% description coverage for its parameters, so the description must compensate, but it does not explain mission_id, feature_id, or repo. The word 'feature' gives minimal context for feature_id, but mission_id and repo remain effectively undocumented beyond their names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('retry') and resource ('failed/blocked feature'), making the tool's core purpose unmistakable. The 'manual' qualifier and failed/blocked scope distinguish it from run/start tools and from automated retry behavior.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly states when to use the tool: for a failed or blocked feature. It also explains how manual retries interact with the automated retry budget, which implies when manual intervention is appropriate. However, it does not explicitly name sibling alternatives or state when not to use it.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mission_runA

Advance the mission control loop (schedule, dispatch isolated workers, validate, correct). Resumable - call again to continue an in-progress mission. max_steps counts DISPATCHES; auto-retries inside a dispatch consume the feature's attempt budget, not a step. If your client sandbox times this call out, the loop keeps running server-side: poll mission_events instead of re-issuing. The response ends with a RUN REPORT (dispatches done, queued work) plus a STATE proof line.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoNo
max_stepsNo
mission_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden, and it covers the key runtime behaviors: server-side continuation after client timeout, step-count semantics (dispatches vs retry budget), resumability, and report/proof output. This is exactly the kind of operational nuance an agent needs before calling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Each sentence conveys a distinct, high-value fact: purpose, resumability, step semantics, timeout behavior, and response shape. The information is front-loaded and dense without being rambling.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex control-loop tool with no annotations and many siblings, the description is unusually complete: it covers the loop's purpose, continuation, timeout safety, step counting, and response structure. It does not elaborate on the optional repo parameter or map every sibling relationship, but the core selection and invocation guidance is present.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, but the description provides essential semantics for max_steps, clarifying that it counts dispatches and does not consume the feature attempt budget. mission_id and repo are recognizable from their titles, so the only genuinely ambiguous parameter is explained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific action ('Advance the mission control loop') and enumerates its inner steps, clearly identifying this as the main driving tool rather than a single validation or retry tool. It distinguishes mission_run from siblings like mission_validate and mission_retry_feature by framing them as parts of the loop it advances.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states that the call is resumable and that an in-progress mission should be continued by calling again. It also gives a concrete when-not-to-reissue rule: after a client-side timeout, poll mission_events instead, preventing duplicate dispatch.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mission_statusC

Mission control status for a mission.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoNo
mission_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. The word 'status' hints at a read-only operation, but the description does not explicitly state side effects, permissions, prerequisites, or any operational behavior beyond the vague notion of status.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single sentence with no padding, so it is concise and front-loaded. However, it is under-specified to the point of being barely informative, missing a verb and any operational detail.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The output schema covers return values, but with a 0% parameter coverage, no annotations, and no guidance differentiating this tool from many similar mission tools, the description is not complete enough for an agent to confidently select and invoke it. Key semantic gaps remain around what 'status' means and how this tool relates to mission_readiness, mission_events, or mission_list.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not mention either parameter, mission_id or repo. Even the phrase 'for a mission' only weakly implies a mission identifier and provides no meaning for the optional repo parameter or the required mission_id.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose2/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description is essentially a noun phrase, 'Mission control status for a mission,' which largely restates the tool name without specifying an action such as retrieving, returning, or querying status. It also does not distinguish this tool from closely related siblings like mission_readiness, mission_events, or mission_list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus any of the 18 sibling tools. The description gives no context about expected use cases, alternatives, or exclusions, leaving the agent to guess whether 'status' means health, progress, readiness, or something else.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mission_steerB

Steer a running mission without rebuilding the plan. Optionally reprioritize a specific feature by id and priority.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoNo
priorityNo
feature_idNo
mission_idYes
instructionNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must disclose what changes occur. It only says 'steer' and 'reprioritize', without explaining side effects, persistence, whether instructions are applied, or operational constraints. It provides some behavioral clues but is insufficient for a mutating operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences with no filler. The first sentence gives the core behavior and contrast, the second adds the optional parameter relationship. Every word earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool mutates a running mission and has a free-form instruction parameter, yet the description does not explain what 'instruction' does, what 'repo' scopes, or what a successful steer entails. The output schema exists, but key input semantics remain ambiguous.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry parameter meaning. It clarifies feature_id and priority together, and implies mission_id, but leaves instruction and repo entirely unexplained. Given five parameters and zero schema-level description, this is a meaningful gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly ties the tool name to a meaningful action: steering a running mission without rebuilding the plan, and explicitly mentions optional reprioritization of a feature. This distinguishes it from mission_replan and from feature-level tools like mission_retry_feature, though 'steer' remains slightly broad.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The phrase 'without rebuilding the plan' implies a contrast with replanning alternatives, and 'running mission' indicates it applies to an active mission. However, it never explicitly says when to use this tool versus mission_replean, mission_adjust_suggestions, or the feature-completion siblings.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

mission_validateB

Run a validator now (scrutiny | user_testing | final). Returns a structured verdict.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNoscrutiny
repoNo
mission_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of disclosing side effects and behavior. It says a validator is run and a verdict returned, but it does not disclose whether this mutates mission state, creates validation records, requires a particular mission status, or has blocking/performance implications.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence with no filler. It efficiently packs the action, the variant choices, and the return type, making it easy to scan while remaining appropriately concise.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with three parameters, no annotations, and no output schema, the description leaves important gaps: mission_id and repo semantics, which validator kind to choose, what the structured verdict contains, and what state changes may occur. An agent would need external workflow knowledge to invoke this confidently.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% description coverage, so the description must compensate. It partially explains the 'kind' parameter by listing its possible values, but 'mission_id' and 'repo' are left semantically unexplained, including what 'repo' refers to and how it affects validation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific action ('Run a validator now'), identifies the resource (validator), enumerates the supported variants (scrutiny | user_testing | final), and tells the agent what to expect ('structured verdict'). This is specific enough to distinguish it from siblings like mission_approve or mission_run.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is given about when to run a validator versus using alternatives like mission_readiness or mission_approve, nor when to choose scrutiny over user_testing or final. The word 'now' implies on-demand execution, but prerequisites, exclusions, and workflow placement are absent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 19 tool updatesv0.2.0
    • First observedmission_apply_suggestions
    • First observedmission_approve
    • First observedmission_cancel
    • First observedmission_complete_feature
    • First observedmission_create
    • First observedmission_events
    • First observedmission_gc
    • First observedmission_list
    • First observedmission_pause
    • First observedmission_plan
    • First observedmission_plan_view
    • First observedmission_readiness
    • First observedmission_replan
    • First observedmission_resume
    • First observedmission_retry_feature
    • First observedmission_run
    • First observedmission_status
    • First observedmission_steer
    • First observedmission_validate

TDQS

B3.2/5.0

Scored across 19 tools

Disambiguation4/5

The tools map cleanly to lifecycle stages (create/plan/approve/run/pause/resume/steer/replan/cancel) and read-only views (status/events/readiness/plan_view). A few control-loop actions like mission_run, mission_resume, and mission_steer could be confused, but their descriptions clarify the differences.

Naming Consistency4/5

All tools share the mission_ prefix and use lower_snake_case, but the pattern mixes bare verbs (mission_run, mission_pause), noun-like resources (mission_status, mission_events), and compound verb_noun forms (mission_retry_feature, mission_apply_suggestions). This is readable and predictable, though not as uniform as a strict verb_noun convention.

Tool Count3/5

At 19 tools, the surface is on the heavy side for an MCP server. Most tools are individually justified by mission lifecycle needs, but a few niche helpers like mission_apply_suggestions and mission_gc, plus several status-like queries, could plausibly be consolidated.

Completeness5/5

The toolset covers the full mission lifecycle: create, plan, approve, run, pause, resume, steer, replan, cancel, plus status/events/readiness/plan_view and manual retry/completion/validation paths. There are no obvious dead ends for the stated purpose, and external-completion and cleanup workflows are included.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers