PowerSwarm
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@PowerSwarmfan out request.json into lanes and report which are green"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
PowerSwarm
Fan work out to many headless coding agents — Grok, Codex or Claude Code — each in its own git worktree, and accept only what passes its kill check.

You (or the agent you work with) split a build into independent targets. PowerSwarm gives each target its own branch and worktree, starts an agent on it, and runs that target's kill check: a command, run without a shell, that has to print one exact line. A lane is green only when the check passes and the agent stayed inside its files. Nothing is merged or pushed. You review the branches and keep what you accept.
It is the open core of the swarm engine I use to build my own tools: same request format, same rules. In the run pictured above, PowerSwarm used Grok to build two pieces of itself: the viewer you are looking at and the quickstart example. Watch that run →
Install
pip install "git+https://github.com/willykeenan/powerswarm"Python 3.9+, git, and at least one agent CLI: Grok Build (grok), Codex (codex) or Claude Code (claude). No other dependencies.
Related MCP server: Agent Squad Bridge
Quickstart
Three lanes finish a tiny library, textstats, one file each:
git clone https://github.com/willykeenan/powerswarm && cd powerswarm && pip install -e .
cd examples/quickstart
repo="$(mktemp -d)/textstats" && mkdir -p "$repo" && cp -R project/. "$repo/"
git -C "$repo" init -q -b main && git -C "$repo" add -A && git -C "$repo" commit -qm start
powerswarm run request.json --root "$repo"The quickstart guide explains each step, how to watch the run and how to review each lane's branch.
How a run works
Validate. The request is checked before anything starts: independent scopes, shell-free kill checks that are committed in
HEADand live outside the lane's own scope, sane budgets. Every problem is reported at once.Worktrees. Each target gets branch
powerswarm/<run>/<target>and its own worktree, starting fromHEAD. Your checkout is never touched.Earned waves. The first wave runs up to 3 lanes. A wave at least 75% green doubles the width; under 25% halves it. The cap is
min(requested_concurrency, half your cores, 32).Attempts. A worker gets a compact brief: its aim, owned scope, definition of done and the kill check as the authoritative acceptance test. If attempt 1 fails, attempt 2 gets the evidence: the check's output, the files changed outside scope, the worktree status. Attempt 1 always leaves room for attempt 2.
Acceptance. Green means: the kill check exited 0, printed the exact line, and left the worktree unchanged; and every changed file is inside the lane's scope.
Bug sweep. A fresh worker hunts for defects in the green lane. If its change breaks the check, it is discarded and the earlier green is kept (the discarded commit stays under
refs/powerswarm/).Receipts.
~/.powerswarm/runs/<run>/holdsrun.json, an event log, and every prompt, worker log and check output. Deadlines are hard; lanes that no longer fit are skipped, not started.
Request
{
"objective": "Add JSON and CSV exports to the reports module",
"targets": [
{
"id": "json-export",
"aim": "Implement reports/export_json.py with export_json(rows) -> str.",
"scope": ["reports/export_json.py", "tests/test_export_json.py"],
"kill_check": {"argv": ["python3", "checks/lane.py", "json-export"], "expected_output": "LANE_OK json-export"}
},
{
"id": "csv-export",
"aim": "Implement reports/export_csv.py with export_csv(rows) -> str (RFC 4180 quoting).",
"scope": ["reports/export_csv.py", "tests/test_export_csv.py"],
"kill_check": {"argv": ["python3", "checks/lane.py", "csv-export"], "expected_output": "LANE_OK csv-export"}
}
]
}powerswarm spec prints every field and rule; powerswarm example prints a complete request.
Runtimes
| Worker | Notes |
| Grok Build CLI | Each lane is a team: explorer, skeptic, implementer and reviewer subagents under Grok's own limits. |
|
|
|
| Claude Code ( | Edit permission plus its Bash tool, inside the worktree. |
| Your agent |
|
| Scripted | Deterministic, for tests and demos. |
--model passes a model id to the runtime.
Commands
powerswarm run REQUEST --root REPO [--runtime ...] [--detach] start (and follow) a run
powerswarm status [RUN] lanes, attempts, reasons, waves
powerswarm view [RUN] [--open] live viewer on 127.0.0.1
powerswarm report [RUN] markdown summary with review commands
powerswarm cancel RUN | recover RUN | clean RUN stop, resume after a crash, remove worktrees
powerswarm validate REQUEST [--root REPO] | spec | example | list
powerswarm mcp MCP server over stdioAdd --json before the command for machine-readable output.
From your agent
PowerSwarm is built to be conducted by an agent. Register the MCP server:
claude mcp add powerswarm -- python3 -m powerswarm mcp# ~/.codex/config.toml
[mcp_servers.powerswarm]
command = "python3"
args = ["-m", "powerswarm", "mcp"]Tools: powerswarm_spec, powerswarm_validate, powerswarm_run, powerswarm_status, powerswarm_cancel,
powerswarm_report. The skill teaches when to use it and how to conduct a run: observe,
challenge, synthesize, advance, verify. A finished swarm is the start of integration, not the end of the task.
Safety
PowerSwarm never merges, pushes, deploys or edits your checkout. Every lane ends as a branch.
Kill checks run without a shell.
bash,sh,envand friends are refused at validation.Lanes cannot start PowerSwarm.
Workers are real agents with your permissions inside their worktrees. The worktree, scope fence and kill check decide what is accepted; they do not sandbox what a worker runs. Use a VM or container for untrusted code. See SECURITY.md.
Status
0.1.0. Tested on macOS and Linux with Python 3.9 and 3.12. Windows is untested.
License
Apache-2.0. Not affiliated with xAI, OpenAI or Anthropic.
Available Tools
6 toolspowerswarm_cancelB
Stop a run's workers. Finished lanes keep their branches.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full behavioral burden, and it does disclose one meaningful trait: finished lanes retain their branches, so cancellation is partial rather than a full rollback. However, it omits whether the action is reversible, what happens to in-flight work, whether it errors on an already-terminated run, or whether it is idempotent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and followed by the key side effect, with no filler or redundancy. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter operation with no output schema and no annotations, the description covers the core action and one important side effect, which is adequate. It still lacks return/status behavior (does it confirm cancellation, return remaining lanes?) and precondition or error handling that an agent would need before invoking blindly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but there is only one parameter (run_id) and its meaning is self-evident from the name and tool context. The description adds no format, source, or lookup details for obtaining a valid run_id, so it neither compensates for the coverage gap nor introduces confusion.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Stop a run's workers,' which unambiguously identifies a cancellation operation keyed by run_id. It does not explicitly contrast itself with siblings like powerswarm_status or powerswarm_run, but the action is clear enough that an agent can distinguish it from the other lifecycle tools by intent.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this versus alternatives such as powerswarm_status (to inspect before cancelling) or what preconditions must hold. The only usage signal is implied by the verb 'Stop,' leaving the agent to infer context for invocation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
powerswarm_reportC
Markdown report of a run: per-lane results, red reasons and review commands.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, yet it only names the output sections. It never states that this is a read-only operation, whether it works on completed vs in-progress runs, how large the report may be, or what happens if run_id is omitted. Useful output detail, but the safety/behavior profile is undisclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with a colon-delimited content list and zero filler; the format is efficient. It is terse almost to a fault, but nothing in it is wasted.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Because there is no output schema, the description's enumeration of report contents is genuinely useful coverage of the return value. However, for a tool with one undocumented and optional parameter plus no annotations, it should also clarify the run_id default and the read-only nature; the definition is minimally viable rather than complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the single run_id parameter is never mentioned in the description. Since run_id is not required, the most important semantic — whether omitting it selects the most recent run or is an error — is left completely undefined.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource (a per-run Markdown report) and enumerates its contents (per-lane results, red reasons, review commands), which distinguishes it from siblings like powerswarm_status and powerswarm_run. It falls short of a 5 because it uses a noun phrase with no verb (generate/render) and never explicitly contrasts itself with the other powerswarm_* tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance at all: nothing says when to prefer this over powerswarm_status, nor whether run_id is required or defaults to the latest run. The agent must infer the trigger condition entirely.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
powerswarm_runB
Start a run: each target gets its own git worktree and headless agent, gated by its kill check. Returns launch truth (worker-live only when a worker process exists). Nothing is merged or pushed.
| Name | Required | Description | Default |
|---|---|---|---|
| root | Yes | clean git repository (absolute path) | |
| solo | No | ||
| model | No | ||
| request | Yes | ||
| runtime | No | grok | |
| max_concurrency | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does well: it discloses side effects (creates a git worktree and a headless agent per target), a safety gate (kill check), and an important non-effect ('Nothing is merged or pushed'). It still omits failure behavior, permission requirements, and what happens to worktrees afterwards.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, front-loaded with the action and its mechanism, and the parenthetical on 'launch truth' adds precision without padding. No wasted prose, though the phrasing is terse enough to be slightly cryptic.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex launch tool with a nested request object, 6 parameters, no annotations, no output schema, and low schema coverage, the description covers behavior but not the inputs an agent must supply. An agent can understand what happens but not how to configure it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 17% (only 'root' is documented), so the description is expected to compensate and does not. Key parameters ('request' nested object, 'solo', 'model', 'runtime', 'max_concurrency') are never explained, leaving half the input surface opaque.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Start a run') and immediately clarifies the mechanism (per-target git worktree + headless agent). It is distinguishable from siblings like powerswarm_spec/validate/cancel/status, but it does not explicitly contrast itself with any of them.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit guidance on when to call this versus powerswarm_spec, powerswarm_validate or powerswarm_status, nor any statement of prerequisites or ordering. The intended lifecycle position (validate first, then run) is left entirely to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
powerswarm_specA
The PowerSwarm request format, rules and an example. Read this before writing a request.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses that this tool returns a format, rules, and an example, which is behaviorally useful. But it doesn't state whether the response is static, cached, or resource-intensive, or whether subsequent calls differ.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with what the tool provides and immediately followed by the call-to-action. No wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a documentation-oriented tool with no parameters, no output schema, and no annotations, the description is adequate but thin. It doesn't describe the shape of the spec (JSON schema, text rules, size) or clarify whether it must be called every session or is a one-time reference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters, so the baseline is 4. The description correctly implies a parameterless retrieval of the spec without contradicting the empty schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies this as a spec/documentation tool for the PowerSwarm request format, which is a clear verb+resource in context. However, it doesn't explicitly distinguish itself from siblings like powerswarm_validate or powerswarm_run beyond implying it's a prerequisite reading step.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives clear conditional guidance: 'Read this before writing a request,' which tells the agent when to invoke it relative to the other PowerSwarm tools. It doesn't name specific alternatives or exclusions, but the sequencing instruction is actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
powerswarm_statusB
Current state of a run (newest if run_id is omitted): lanes, attempts, reasons, waves.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses that this retrieves current state and defaults to newest when run_id is omitted, but it does not state read-only/safety properties, auth requirements, or error behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single compact sentence, front-loaded with the core action and followed by the scoped return fields. There is no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple status tool, it covers the default behavior and high-level return contents. With no output schema and no annotations, however, it stops short of specifying output shape, safety characteristics, or edge cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The only parameter, run_id, has 0% schema description coverage. The description adds important default semantics—omitting run_id returns the newest run—but does not explain run_id format, expected identifier type, or accepted values.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource ('a run') and action ('current state'), and enumerates what the state contains (lanes, attempts, reasons, waves). It separates status from start/cancel siblings implicitly, though it does not explicitly contrast with powerswarm_report.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives one usage rule: omit run_id to get the newest run. However, there is no explicit when-to-use or when-not-to-use guidance versus siblings such as powerswarm_report, and no prerequisites are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
powerswarm_validateB
Check a request (and optionally that its kill checks are committed in root) without running anything.
| Name | Required | Description | Default |
|---|---|---|---|
| root | No | ||
| request | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It usefully discloses that nothing is executed (a non-mutating, side-effect-free operation), but it says nothing about authentication, error reporting, or what a successful vs failed validation returns.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single sentence with the core constraint ('without running anything') front-loaded and no padding. It loses a point only because the parenthetical jargon is dense and hard to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no annotations, no output schema, a nested request object, and 0% parameter description coverage, the description is too thin. It omits what a valid request looks like, what 'kill checks in root' means, and what the validation result contains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, so the description must compensate for two undocumented parameters. It adds some meaning: 'root' is optional and relates to 'kill checks committed in root', but the required 'request' object's structure and contents are entirely unexplained.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb (check/validate) and resource (a request), and the phrase 'without running anything' implicitly contrasts it with powerswarm_run. However, it never names a sibling explicitly and the jargon 'kill checks committed in root' blurs what is actually being validated.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
'Without running anything' hints that this is the dry-run alternative to powerswarm_run, which is implied usage guidance. But there is no explicit statement of when to use this versus powerswarm_spec or powerswarm_run, and no prerequisites are given.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v0.1.0- First observed
powerswarm_cancel - First observed
powerswarm_report - First observed
powerswarm_run - First observed
powerswarm_spec - First observed
powerswarm_status - First observed
powerswarm_validate
TDQS
Scored across 6 tools
Each tool targets a distinct lifecycle stage: spec (reference), validate (pre-flight check), run (execution), status (monitoring), cancel (halt), and report (summary). There is no overlap in purpose, so an agent can easily select the right tool.
All names use the powerswarm_ prefix and snake_case, which is highly predictable. However, the suffixes mix verbs (validate, run, cancel) with nouns (spec, status, report), a minor deviation from a pure verb_noun pattern.
Six tools are well-scoped for an orchestration workflow. Each tool covers a necessary part of the run lifecycle without redundancy or bloat.
The surface covers reference, validation, execution, monitoring, cancellation, and reporting. Minor gaps exist, such as no explicit list-all-runs or worktree cleanup operation, though status defaults to the newest run and report provides review commands.
Maintenance
Related MCP Connectors
Inspect Depot builds, CI runs, job logs and Actions runners; retry or cancel CI runs.
Build, test, deploy, and operate Connic agents.
AI-native git hosting — repos, PRs, issues, CI gates, and AI code review over MCP (60 tools).
- HydrantOAuthdev.hydrant
Track issues, projects and dependencies with your agent in a workspace you control.
Related MCP Servers
- AlicenseAqualityAmaintenanceCoordination protocol for parallel coding agents, built above Git. The MCP server exposes the full 17-tool lifecycle over stdio: register agents, publish intents with semantic scopes and declared operations, claim work, check for conflicts before code is written, publish ChangeSets, run trusted named verification checks, and record accepted work. Deterministic rules raise findings when two agents'18507Apache 2.0
- AlicenseNot gradedqualityCmaintenanceEnables multiple AI coding CLIs (Claude Code, Gemini/Antigravity, Codex, and OpenCode) to collaborate as a coordinated team by routing cross-agent prompts, sharing messages and review tickets, tracking tasks on a shared store, and isolating each agent in its own Git worktree with turn-budget safeguards—all inspectable and steerable from a local web dashboard.MIT
- FlicenseNot gradedqualityCmaintenanceEnables multiple AI coding agents to coordinate on a shared software project by registering, claiming tasks, declaring file intents, publishing structured change reports, and handing off context, with a local dashboard showing state in near real time.-
- AlicenseNot gradedqualityBmaintenanceTurns a build or audit target into a graph of claims that a coding agent must decompose and verify bottom-up against evidence, with judgments pinned to git commits and file hashes and findings recorded as issues on the claims they refute. It keeps an append-only, replayable record of every operation, exposes the standing and frontier of the work, and flags stale judgments when cited code changes.81 npm4MIT