Skip to main content
Glama

PowerSwarm

Fan work out to many headless coding agents — Grok, Codex or Claude Code — each in its own git worktree, and accept only what passes its kill check.

The PowerSwarm viewer showing a real run: two Grok team lanes, both green after a bug sweep

You (or the agent you work with) split a build into independent targets. PowerSwarm gives each target its own branch and worktree, starts an agent on it, and runs that target's kill check: a command, run without a shell, that has to print one exact line. A lane is green only when the check passes and the agent stayed inside its files. Nothing is merged or pushed. You review the branches and keep what you accept.

It is the open core of the swarm engine I use to build my own tools: same request format, same rules. In the run pictured above, PowerSwarm used Grok to build two pieces of itself: the viewer you are looking at and the quickstart example. Watch that run →

Install

pip install "git+https://github.com/willykeenan/powerswarm"

Python 3.9+, git, and at least one agent CLI: Grok Build (grok), Codex (codex) or Claude Code (claude). No other dependencies.

Related MCP server: Agent Squad Bridge

Quickstart

Three lanes finish a tiny library, textstats, one file each:

git clone https://github.com/willykeenan/powerswarm && cd powerswarm && pip install -e .
cd examples/quickstart
repo="$(mktemp -d)/textstats" && mkdir -p "$repo" && cp -R project/. "$repo/"
git -C "$repo" init -q -b main && git -C "$repo" add -A && git -C "$repo" commit -qm start
powerswarm run request.json --root "$repo"

The quickstart guide explains each step, how to watch the run and how to review each lane's branch.

How a run works

  1. Validate. The request is checked before anything starts: independent scopes, shell-free kill checks that are committed in HEAD and live outside the lane's own scope, sane budgets. Every problem is reported at once.

  2. Worktrees. Each target gets branch powerswarm/<run>/<target> and its own worktree, starting from HEAD. Your checkout is never touched.

  3. Earned waves. The first wave runs up to 3 lanes. A wave at least 75% green doubles the width; under 25% halves it. The cap is min(requested_concurrency, half your cores, 32).

  4. Attempts. A worker gets a compact brief: its aim, owned scope, definition of done and the kill check as the authoritative acceptance test. If attempt 1 fails, attempt 2 gets the evidence: the check's output, the files changed outside scope, the worktree status. Attempt 1 always leaves room for attempt 2.

  5. Acceptance. Green means: the kill check exited 0, printed the exact line, and left the worktree unchanged; and every changed file is inside the lane's scope.

  6. Bug sweep. A fresh worker hunts for defects in the green lane. If its change breaks the check, it is discarded and the earlier green is kept (the discarded commit stays under refs/powerswarm/).

  7. Receipts. ~/.powerswarm/runs/<run>/ holds run.json, an event log, and every prompt, worker log and check output. Deadlines are hard; lanes that no longer fit are skipped, not started.

Request

{
  "objective": "Add JSON and CSV exports to the reports module",
  "targets": [
    {
      "id": "json-export",
      "aim": "Implement reports/export_json.py with export_json(rows) -> str.",
      "scope": ["reports/export_json.py", "tests/test_export_json.py"],
      "kill_check": {"argv": ["python3", "checks/lane.py", "json-export"], "expected_output": "LANE_OK json-export"}
    },
    {
      "id": "csv-export",
      "aim": "Implement reports/export_csv.py with export_csv(rows) -> str (RFC 4180 quoting).",
      "scope": ["reports/export_csv.py", "tests/test_export_csv.py"],
      "kill_check": {"argv": ["python3", "checks/lane.py", "csv-export"], "expected_output": "LANE_OK csv-export"}
    }
  ]
}

powerswarm spec prints every field and rule; powerswarm example prints a complete request.

Runtimes

--runtime

Worker

Notes

grok (default)

Grok Build CLI

Each lane is a team: explorer, skeptic, implementer and reviewer subagents under Grok's own limits. --solo for one worker per lane.

codex

codex exec

workspace-write sandbox.

claude

Claude Code (claude -p)

Edit permission plus its Bash tool, inside the worktree.

command

Your agent

--command-json '["my-agent","--cwd","{worktree}","--prompt-file","{prompt_file}"]'

fake

Scripted

Deterministic, for tests and demos.

--model passes a model id to the runtime.

Commands

powerswarm run REQUEST --root REPO [--runtime ...] [--detach]   start (and follow) a run
powerswarm status [RUN]                                         lanes, attempts, reasons, waves
powerswarm view [RUN] [--open]                                  live viewer on 127.0.0.1
powerswarm report [RUN]                                         markdown summary with review commands
powerswarm cancel RUN | recover RUN | clean RUN                 stop, resume after a crash, remove worktrees
powerswarm validate REQUEST [--root REPO] | spec | example | list
powerswarm mcp                                                  MCP server over stdio

Add --json before the command for machine-readable output.

From your agent

PowerSwarm is built to be conducted by an agent. Register the MCP server:

claude mcp add powerswarm -- python3 -m powerswarm mcp
# ~/.codex/config.toml
[mcp_servers.powerswarm]
command = "python3"
args = ["-m", "powerswarm", "mcp"]

Tools: powerswarm_spec, powerswarm_validate, powerswarm_run, powerswarm_status, powerswarm_cancel, powerswarm_report. The skill teaches when to use it and how to conduct a run: observe, challenge, synthesize, advance, verify. A finished swarm is the start of integration, not the end of the task.

Safety

  • PowerSwarm never merges, pushes, deploys or edits your checkout. Every lane ends as a branch.

  • Kill checks run without a shell. bash, sh, env and friends are refused at validation.

  • Lanes cannot start PowerSwarm.

  • Workers are real agents with your permissions inside their worktrees. The worktree, scope fence and kill check decide what is accepted; they do not sandbox what a worker runs. Use a VM or container for untrusted code. See SECURITY.md.

Status

0.1.0. Tested on macOS and Linux with Python 3.9 and 3.12. Windows is untested.

License

Apache-2.0. Not affiliated with xAI, OpenAI or Anthropic.

Available Tools

6 tools
powerswarm_cancelB

Stop a run's workers. Finished lanes keep their branches.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden, and it does disclose one meaningful trait: finished lanes retain their branches, so cancellation is partial rather than a full rollback. However, it omits whether the action is reversible, what happens to in-flight work, whether it errors on an already-terminated run, or whether it is idempotent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action and followed by the key side effect, with no filler or redundancy. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter operation with no output schema and no annotations, the description covers the core action and one important side effect, which is adequate. It still lacks return/status behavior (does it confirm cancellation, return remaining lanes?) and precondition or error handling that an agent would need before invoking blindly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, but there is only one parameter (run_id) and its meaning is self-evident from the name and tool context. The description adds no format, source, or lookup details for obtaining a valid run_id, so it neither compensates for the coverage gap nor introduces confusion.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Stop a run's workers,' which unambiguously identifies a cancellation operation keyed by run_id. It does not explicitly contrast itself with siblings like powerswarm_status or powerswarm_run, but the action is clear enough that an agent can distinguish it from the other lifecycle tools by intent.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this versus alternatives such as powerswarm_status (to inspect before cancelling) or what preconditions must hold. The only usage signal is implied by the verb 'Stop,' leaving the agent to infer context for invocation.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

powerswarm_reportC

Markdown report of a run: per-lane results, red reasons and review commands.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idNo

TDQS

C2.8/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, yet it only names the output sections. It never states that this is a read-only operation, whether it works on completed vs in-progress runs, how large the report may be, or what happens if run_id is omitted. Useful output detail, but the safety/behavior profile is undisclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with a colon-delimited content list and zero filler; the format is efficient. It is terse almost to a fault, but nothing in it is wasted.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Because there is no output schema, the description's enumeration of report contents is genuinely useful coverage of the return value. However, for a tool with one undocumented and optional parameter plus no annotations, it should also clarify the run_id default and the read-only nature; the definition is minimally viable rather than complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and the single run_id parameter is never mentioned in the description. Since run_id is not required, the most important semantic — whether omitting it selects the most recent run or is an error — is left completely undefined.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource (a per-run Markdown report) and enumerates its contents (per-lane results, red reasons, review commands), which distinguishes it from siblings like powerswarm_status and powerswarm_run. It falls short of a 5 because it uses a noun phrase with no verb (generate/render) and never explicitly contrasts itself with the other powerswarm_* tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no when-to-use guidance at all: nothing says when to prefer this over powerswarm_status, nor whether run_id is required or defaults to the latest run. The agent must infer the trigger condition entirely.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

powerswarm_runB

Start a run: each target gets its own git worktree and headless agent, gated by its kill check. Returns launch truth (worker-live only when a worker process exists). Nothing is merged or pushed.

ParametersJSON Schema
NameRequiredDescriptionDefault
rootYesclean git repository (absolute path)
soloNo
modelNo
requestYes
runtimeNogrok
max_concurrencyNo

TDQS

B3.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does well: it discloses side effects (creates a git worktree and a headless agent per target), a safety gate (kill check), and an important non-effect ('Nothing is merged or pushed'). It still omits failure behavior, permission requirements, and what happens to worktrees afterwards.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences, front-loaded with the action and its mechanism, and the parenthetical on 'launch truth' adds precision without padding. No wasted prose, though the phrasing is terse enough to be slightly cryptic.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex launch tool with a nested request object, 6 parameters, no annotations, no output schema, and low schema coverage, the description covers behavior but not the inputs an agent must supply. An agent can understand what happens but not how to configure it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 17% (only 'root' is documented), so the description is expected to compensate and does not. Key parameters ('request' nested object, 'solo', 'model', 'runtime', 'max_concurrency') are never explained, leaving half the input surface opaque.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Start a run') and immediately clarifies the mechanism (per-target git worktree + headless agent). It is distinguishable from siblings like powerswarm_spec/validate/cancel/status, but it does not explicitly contrast itself with any of them.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit guidance on when to call this versus powerswarm_spec, powerswarm_validate or powerswarm_status, nor any statement of prerequisites or ordering. The intended lifecycle position (validate first, then run) is left entirely to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

powerswarm_specA

The PowerSwarm request format, rules and an example. Read this before writing a request.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses that this tool returns a format, rules, and an example, which is behaviorally useful. But it doesn't state whether the response is static, cached, or resource-intensive, or whether subsequent calls differ.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with what the tool provides and immediately followed by the call-to-action. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a documentation-oriented tool with no parameters, no output schema, and no annotations, the description is adequate but thin. It doesn't describe the shape of the spec (JSON schema, text rules, size) or clarify whether it must be called every session or is a one-time reference.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There are zero parameters, so the baseline is 4. The description correctly implies a parameterless retrieval of the spec without contradicting the empty schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description identifies this as a spec/documentation tool for the PowerSwarm request format, which is a clear verb+resource in context. However, it doesn't explicitly distinguish itself from siblings like powerswarm_validate or powerswarm_run beyond implying it's a prerequisite reading step.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear conditional guidance: 'Read this before writing a request,' which tells the agent when to invoke it relative to the other PowerSwarm tools. It doesn't name specific alternatives or exclusions, but the sequencing instruction is actionable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

powerswarm_statusB

Current state of a run (newest if run_id is omitted): lanes, attempts, reasons, waves.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idNo

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses that this retrieves current state and defaults to newest when run_id is omitted, but it does not state read-only/safety properties, auth requirements, or error behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single compact sentence, front-loaded with the core action and followed by the scoped return fields. There is no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple status tool, it covers the default behavior and high-level return contents. With no output schema and no annotations, however, it stops short of specifying output shape, safety characteristics, or edge cases.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The only parameter, run_id, has 0% schema description coverage. The description adds important default semantics—omitting run_id returns the newest run—but does not explain run_id format, expected identifier type, or accepted values.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific resource ('a run') and action ('current state'), and enumerates what the state contains (lanes, attempts, reasons, waves). It separates status from start/cancel siblings implicitly, though it does not explicitly contrast with powerswarm_report.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives one usage rule: omit run_id to get the newest run. However, there is no explicit when-to-use or when-not-to-use guidance versus siblings such as powerswarm_report, and no prerequisites are stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

powerswarm_validateB

Check a request (and optionally that its kill checks are committed in root) without running anything.

ParametersJSON Schema
NameRequiredDescriptionDefault
rootNo
requestYes

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses that nothing is executed (a non-mutating, side-effect-free operation), but it says nothing about authentication, error reporting, or what a successful vs failed validation returns.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single sentence with the core constraint ('without running anything') front-loaded and no padding. It loses a point only because the parenthetical jargon is dense and hard to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations, no output schema, a nested request object, and 0% parameter description coverage, the description is too thin. It omits what a valid request looks like, what 'kill checks in root' means, and what the validation result contains.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, so the description must compensate for two undocumented parameters. It adds some meaning: 'root' is optional and relates to 'kill checks committed in root', but the required 'request' object's structure and contents are entirely unexplained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (check/validate) and resource (a request), and the phrase 'without running anything' implicitly contrasts it with powerswarm_run. However, it never names a sibling explicitly and the jargon 'kill checks committed in root' blurs what is actually being validated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Without running anything' hints that this is the dry-run alternative to powerswarm_run, which is implied usage guidance. But there is no explicit statement of when to use this versus powerswarm_spec or powerswarm_run, and no prerequisites are given.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 6 tool updatesv0.1.0
    • First observedpowerswarm_cancel
    • First observedpowerswarm_report
    • First observedpowerswarm_run
    • First observedpowerswarm_spec
    • First observedpowerswarm_status
    • First observedpowerswarm_validate

TDQS

B3.4/5.0

Scored across 6 tools

Disambiguation5/5

Each tool targets a distinct lifecycle stage: spec (reference), validate (pre-flight check), run (execution), status (monitoring), cancel (halt), and report (summary). There is no overlap in purpose, so an agent can easily select the right tool.

Naming Consistency4/5

All names use the powerswarm_ prefix and snake_case, which is highly predictable. However, the suffixes mix verbs (validate, run, cancel) with nouns (spec, status, report), a minor deviation from a pure verb_noun pattern.

Tool Count5/5

Six tools are well-scoped for an orchestration workflow. Each tool covers a necessary part of the run lifecycle without redundancy or bloat.

Completeness4/5

The surface covers reference, validation, execution, monitoring, cancellation, and reporting. Minor gaps exist, such as no explicit list-all-runs or worktree cleanup operation, though status defaults to the newest run and report provides review commands.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    A
    maintenance
    Coordination protocol for parallel coding agents, built above Git. The MCP server exposes the full 17-tool lifecycle over stdio: register agents, publish intents with semantic scopes and declared operations, claim work, check for conflicts before code is written, publish ChangeSets, run trusted named verification checks, and record accepted work. Deterministic rules raise findings when two agents'
    18
    507
    Apache 2.0
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables multiple AI coding CLIs (Claude Code, Gemini/Antigravity, Codex, and OpenCode) to collaborate as a coordinated team by routing cross-agent prompts, sharing messages and review tickets, tracking tasks on a shared store, and isolating each agent in its own Git worktree with turn-budget safeguards—all inspectable and steerable from a local web dashboard.
    MIT
  • F
    license
    Not graded
    quality
    C
    maintenance
    Enables multiple AI coding agents to coordinate on a shared software project by registering, claiming tasks, declaring file intents, publishing structured change reports, and handing off context, with a local dashboard showing state in near real time.
    -
  • A
    license
    Not graded
    quality
    B
    maintenance
    Turns a build or audit target into a graph of claims that a coding agent must decompose and verify bottom-up against evidence, with judgments pinned to git commits and file hashes and findings recorded as issues on the claims they refute. It keeps an append-only, replayable record of every operation, exposes the standing and frontier of the work, and flags stale judgments when cited code changes.
    81 npm
    4
    MIT