Skip to main content
Glama
README.md
# PowerSwarm

**Fan work out to many headless coding agents — Grok, Codex or Claude Code — each in its own git worktree, and accept only what passes its kill check.**

![The PowerSwarm viewer showing a real run: two Grok team lanes, both green after a bug sweep](docs/images/viewer.png)

You (or the agent you work with) split a build into independent targets. PowerSwarm gives each target its own branch
and worktree, starts an agent on it, and runs that target's **kill check**: a command, run without a shell, that has to
print one exact line. A lane is green only when the check passes and the agent stayed inside its files. Nothing is
merged or pushed. You review the branches and keep what you accept.

It is the open core of the swarm engine I use to build my own tools: same request format, same rules. In the run
pictured above, PowerSwarm used Grok to build two pieces of itself: the viewer you are looking at and the
[quickstart example](examples/quickstart/). [Watch that run →](https://huggingface.co/spaces/willykeenan/powerswarm)

## Install

```sh
pip install "git+https://github.com/willykeenan/powerswarm"
```

Python 3.9+, git, and at least one agent CLI: Grok Build (`grok`), [Codex](https://github.com/openai/codex) (`codex`) or [Claude Code](https://www.anthropic.com/claude-code) (`claude`). No other dependencies.

## Quickstart

Three lanes finish a tiny library, `textstats`, one file each:

```sh
git clone https://github.com/willykeenan/powerswarm && cd powerswarm && pip install -e .
cd examples/quickstart
repo="$(mktemp -d)/textstats" && mkdir -p "$repo" && cp -R project/. "$repo/"
git -C "$repo" init -q -b main && git -C "$repo" add -A && git -C "$repo" commit -qm start
powerswarm run request.json --root "$repo"
```

The [quickstart guide](examples/quickstart/README.md) explains each step, how to watch the run and how to review each
lane's branch.

## How a run works

1. **Validate.** The request is checked before anything starts: independent scopes, shell-free kill checks that are
   committed in `HEAD` and live outside the lane's own scope, sane budgets. Every problem is reported at once.
2. **Worktrees.** Each target gets branch `powerswarm/<run>/<target>` and its own worktree, starting from `HEAD`.
   Your checkout is never touched.
3. **Earned waves.** The first wave runs up to 3 lanes. A wave at least 75% green doubles the width; under 25% halves
   it. The cap is `min(requested_concurrency, half your cores, 32)`.
4. **Attempts.** A worker gets a compact brief: its aim, owned scope, definition of done and the kill check as the
   authoritative acceptance test. If attempt 1 fails, attempt 2 gets the evidence: the check's output, the files
   changed outside scope, the worktree status. Attempt 1 always leaves room for attempt 2.
5. **Acceptance.** Green means: the kill check exited 0, printed the exact line, and left the worktree unchanged; and
   every changed file is inside the lane's scope.
6. **Bug sweep.** A fresh worker hunts for defects in the green lane. If its change breaks the check, it is discarded
   and the earlier green is kept (the discarded commit stays under `refs/powerswarm/`).
7. **Receipts.** `~/.powerswarm/runs/<run>/` holds `run.json`, an event log, and every prompt, worker log and check
   output. Deadlines are hard; lanes that no longer fit are skipped, not started.

## Request

```json
{
  "objective": "Add JSON and CSV exports to the reports module",
  "targets": [
    {
      "id": "json-export",
      "aim": "Implement reports/export_json.py with export_json(rows) -> str.",
      "scope": ["reports/export_json.py", "tests/test_export_json.py"],
      "kill_check": {"argv": ["python3", "checks/lane.py", "json-export"], "expected_output": "LANE_OK json-export"}
    },
    {
      "id": "csv-export",
      "aim": "Implement reports/export_csv.py with export_csv(rows) -> str (RFC 4180 quoting).",
      "scope": ["reports/export_csv.py", "tests/test_export_csv.py"],
      "kill_check": {"argv": ["python3", "checks/lane.py", "csv-export"], "expected_output": "LANE_OK csv-export"}
    }
  ]
}
```

`powerswarm spec` prints every field and rule; `powerswarm example` prints a complete request.

## Runtimes

| `--runtime` | Worker | Notes |
| --- | --- | --- |
| `grok` (default) | Grok Build CLI | Each lane is a **team**: explorer, skeptic, implementer and reviewer subagents under Grok's own limits. `--solo` for one worker per lane. |
| `codex` | `codex exec` | `workspace-write` sandbox. |
| `claude` | Claude Code (`claude -p`) | Edit permission plus its Bash tool, inside the worktree. |
| `command` | Your agent | `--command-json '["my-agent","--cwd","{worktree}","--prompt-file","{prompt_file}"]'` |
| `fake` | Scripted | Deterministic, for tests and demos. |

`--model` passes a model id to the runtime.

## Commands

```text
powerswarm run REQUEST --root REPO [--runtime ...] [--detach]   start (and follow) a run
powerswarm status [RUN]                                         lanes, attempts, reasons, waves
powerswarm view [RUN] [--open]                                  live viewer on 127.0.0.1
powerswarm report [RUN]                                         markdown summary with review commands
powerswarm cancel RUN | recover RUN | clean RUN                 stop, resume after a crash, remove worktrees
powerswarm validate REQUEST [--root REPO] | spec | example | list
powerswarm mcp                                                  MCP server over stdio
```

Add `--json` before the command for machine-readable output.

## From your agent

PowerSwarm is built to be conducted by an agent. Register the MCP server:

```sh
claude mcp add powerswarm -- python3 -m powerswarm mcp
```

```toml
# ~/.codex/config.toml
[mcp_servers.powerswarm]
command = "python3"
args = ["-m", "powerswarm", "mcp"]
```

Tools: `powerswarm_spec`, `powerswarm_validate`, `powerswarm_run`, `powerswarm_status`, `powerswarm_cancel`,
`powerswarm_report`. The [skill](skills/powerswarm/SKILL.md) teaches when to use it and how to conduct a run: observe,
challenge, synthesize, advance, verify. A finished swarm is the start of integration, not the end of the task.

## Safety

- PowerSwarm never merges, pushes, deploys or edits your checkout. Every lane ends as a branch.
- Kill checks run without a shell. `bash`, `sh`, `env` and friends are refused at validation.
- Lanes cannot start PowerSwarm.
- **Workers are real agents with your permissions inside their worktrees.** The worktree, scope fence and kill check
  decide what is accepted; they do not sandbox what a worker runs. Use a VM or container for untrusted code.
  See [SECURITY.md](SECURITY.md).

## Status

0.1.0. Tested on macOS and Linux with Python 3.9 and 3.12. Windows is untested.

## License

Apache-2.0. Not affiliated with xAI, OpenAI or Anthropic.

TDQS

B3.4/5.0

Scored across 6 tools

Disambiguation5/5

Each tool targets a distinct lifecycle stage: spec (reference), validate (pre-flight check), run (execution), status (monitoring), cancel (halt), and report (summary). There is no overlap in purpose, so an agent can easily select the right tool.

Naming Consistency4/5

All names use the powerswarm_ prefix and snake_case, which is highly predictable. However, the suffixes mix verbs (validate, run, cancel) with nouns (spec, status, report), a minor deviation from a pure verb_noun pattern.

Tool Count5/5

Six tools are well-scoped for an orchestration workflow. Each tool covers a necessary part of the run lifecycle without redundancy or bloat.

Completeness4/5

The surface covers reference, validation, execution, monitoring, cancellation, and reporting. Minor gaps exist, such as no explicit list-all-runs or worktree cleanup operation, though status defaults to the newest run and report provides review commands.

Maintenance

ActivityMaintained
ResponsivenessNo issues