ZergSwarm
Allows using Hugging Face's model inference as a provider, routing swarm tasks to a wide range of open models with a small free monthly credit.
Allows connecting a local Ollama server as an OpenAI-compatible provider, letting the swarm run tasks on local models without a cloud API key.
Provides access to Perplexity's paid per-use models as a swarm provider for routing tasks to Perplexity.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@ZergSwarmSummarize every file in ./logs and return JSON with filename and summary."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Hand your coding agent a swarm. Claude Code, Claude Desktop and Codex send the busywork to dozens of cheap AI models at once, and get the answers back as data.
A real run: 12 review jobs on one cheap model (Qwen3.8 27B on Groq), all at once. 69 seconds and $0.30 at list price, none of it on a Claude plan.
Your main agent is expensive and does one thing at a time. Most of what it spends its day on is wide and repetitive: read sixty files and report on each, check every call site, summarize every log, grade every answer. Today it grinds through that alone, or spins up subagents that still run on your Claude plan. ZergSwarm is for people who use Claude Code, Claude Desktop or Codex every day and would rather spend that plan on the thinking.
ZergSwarm takes that part. It is an MCP server, a CLI and a local web console. Your agent hands it a batch of tasks; it runs them all at once on the cheapest models that are good enough for the job (NVIDIA, Gemini, Groq, Cerebras, DeepSeek, OpenRouter, Mistral, Hugging Face, or any OpenAI-compatible endpoint), and hands back each answer, checked against the JSON schema you asked for. Your agent stays the one that plans, decides and verifies.
The easiest install is to tell your agent (Claude Code, Codex, or the Code tab in Claude Desktop):
Install ZergSwarm for me with its one-line installer: https://github.com/Lunarwerx/ZergSwarm
Or run the installer yourself:
irm https://raw.githubusercontent.com/Lunarwerx/ZergSwarm/main/install.ps1 | iex # Windows (PowerShell)
curl -fsSL https://raw.githubusercontent.com/Lunarwerx/ZergSwarm/main/install.sh | sh # macOS and LinuxEither way, that one command installs it, connects it to every assistant it finds (Claude Code, Claude Desktop, Codex) and opens the console. Paste one free key there (a Gemini or Groq key is the quickest; one is enough), then ask your agent in plain words, in a new chat: "Use zswarm to review every file in src/ for bugs." You never write the batch yourself.
✨ At a glance
Many workers, one call |
|
It picks the model | Leave the model on |
Starts on a free key | NVIDIA, Gemini, Groq and Cerebras have free tiers that AUTO can use. Paste one key and it works. NVIDIA's calls cost nothing, so AUTO tries its models first. |
Keys that look after themselves | A rate-limited key rests, a key out of credit is set aside, a key the provider refuses is never kept. When a provider runs dry, tasks move to the next capable model instead of stopping. |
Every call costed | Each task records its model, time and price. The console charts spend per day, per provider and per model. There is no spending limit unless you set one; set a daily limit and new work stops when a day reaches it. |
Workers get only what you allow | Each task names its tools: |
One console for all of it |
|
Related MCP server: Rutherford MCP Server
📦 Install
Each installer puts the zswarm command on your PATH from the latest release, in its own environment: with
uv when you have it, else with pipx on a Python 3.11 or newer.
With neither, it installs uv, which brings its own Python, so there is nothing to install first. Then it runs
zswarm setup.
An agent installing this for someone: run the one-liner for their system, below. It asks nothing and ends by opening the console in their browser. The key goes into that console, never into the chat. Then tell them to open a new chat (restart Claude Desktop) so the zswarm tools load.
Windows (PowerShell)
irm https://raw.githubusercontent.com/Lunarwerx/ZergSwarm/main/install.ps1 | iexmacOS and Linux
curl -fsSL https://raw.githubusercontent.com/Lunarwerx/ZergSwarm/main/install.sh | shWith pip or pipx directly
pipx install "git+https://github.com/Lunarwerx/ZergSwarm" # or: pip install "git+https://..."
uv tool install "git+https://github.com/Lunarwerx/ZergSwarm"Then run zswarm setup.
Every release also carries the wheel and the source archive,
if you would rather download and install a file (pipx install zergswarm-<version>-py3-none-any.whl).
New terminals see the zswarm command; check it with zswarm --version. To install without connecting anything,
set ZSWARM_NO_SETUP=1 first; then zswarm setup --client claude-code connects just that one (or codex,
claude-desktop).
🚀 Quick start
Install. Tell your agent to, or run the one-line installer above; it ends by running
zswarm setup, which registers the MCP server with every assistant it finds and opens the console athttp://127.0.0.1:7790/ui. Runzswarm setupagain any time;--clientpicks the assistants, and--instructionsalso adds a short "when to use the swarm" note to~/.claude/CLAUDE.mdand~/.codex/AGENTS.md.Add a key. One is enough. In the console, pick a provider (Gemini or Groq is the quickest, both free), follow its Get a key link, and paste the key on its page. The providers marked picks automatically (NVIDIA, Gemini, Groq, Cerebras, DeepSeek and OpenRouter) are the ones ZergSwarm can choose models from by itself; the others you name when you want them. The key is checked with the provider straight away and kept only if it works. From the terminal:
zswarm keys add gemini(it asks for the key, hidden, and checks it the same way).Ask your agent to use it, in a new chat:
"Use zswarm to read every file under src/ and list the functions that do network I/O."
"Use zswarm to review my uncommitted changes."
zswarm doctor says what is ready: keys, models, binaries.
No agent? zswarm run tasks.json runs the same kind of batch from a terminal, and any program on this machine can
send one to the local HTTP API (docs/API.md).
🖥️ The console
A tree of everything on the left, the selected item on the right. Each provider opens onto its models; each provider and model page charts its own use.
Every button in it is a JSON call you can make yourself: docs/API.md.
🧠 How it picks a model
zswarm/data/published-models.json holds published benchmark results. Every model whose provider file names a
benchmark_slug takes part in AUTO. For each task, AUTO:
keeps the configurations that meet every score floor of the task's profile (
general,code,decision,research,critical,routine),drops those whose provider has no ready key,
puts your starred models first (star a model in the console, or give it a
priority),orders the rest by what the published test run cost, cheapest first.
If a provider runs out mid-task, the task carries on with the next configuration, its finished tool calls kept.
Preview the choice without a model call: zswarm_select, or Routing & roles › Preview AUTO in the console.
Details: docs/RUNTIME-SELECTION.md.
Everything else that needs a model picks the same way: the judge, the doubt and review roles, the blind panel
(two different makers' models) and a bare auto. So one key is enough: every part of ZergSwarm runs on the
providers you have a key for, never on one you do not.
🧰 Tools your agent gets
tool | what it does |
| a batch of tasks, run concurrently; each answer comes back, plus |
| one tool-free question: classify, summarize, rewrite, a second opinion |
| follow a long batch started with |
| which models AUTO would use right now, without a model call |
| review a change: a worker reads the diff (or your uncommitted work) and cites the line of each finding, then a second worker tries to disprove each one |
| two to five different models answer the same question, see each other's answers without names, and show where they agree |
| a fresh second look at an answer, by a reviewer that sees only the answer and what it must meet |
| typed decisions (pick one, yes or no, a score) answered in bulk |
| what is configured, ready and spent |
A task looks like this, a real one run over this repository:
{"id": "review-redaction",
"prompt": "Review zswarm/redaction.py for bugs a user could hit. Report only real problems, each with its line, the exact code on that line, and one sentence saying what goes wrong. If there are none, return an empty list.",
"cwd": "/path/to/ZergSwarm", "tools": "read",
"schema": {"type": "object", "required": ["findings"], "properties": {"findings": {"type": "array", "items": {
"type": "object", "required": ["line", "code", "problem"], "properties": {
"line": {"type": "integer"}, "code": {"type": "string"}, "problem": {"type": "string"}}}}}}}And what came back, shortened to fit (the real answer also quotes the code on each line):
{"id": "review-redaction", "status": "ok", "model": "rank:qwen3-8-27b:groq", "seconds": 40.1, "cost_usd": 0.0378,
"data": {"findings": [
{"line": 159, "problem": "A task pattern named like a built-in detector ('secret') replaces that detector instead of adding to it, so real keys go out unredacted."},
{"line": 49, "problem": "conf['password'] = '012345678901' is not caught, while password = '012345678901' is: the same password leaks in one form."},
{"line": 52, "problem": "The card pattern lets 19 digits through, so a 19-digit order number is masked as a card."}]}}Workers are cheap, not infallible. The first two findings above were real bugs, and both are fixed; the third was
wrong (13 to 19 digits is what card numbers have). A finding that citesfile:line should be checked at that line
before anything depends on it.
🔑 Providers
You need only one key to start: a free Gemini or Groq key is the quickest. Every provider is one TOML file in
zswarm/providers/. Yours live in ~/.zswarm/providers/
and change only what they say; a new file name is a new provider. The console writes these files for you and
keeps your comments.
provider | free tier | picks automatically | notes |
NVIDIA | ✅ | ✅ | build.nvidia.com trial keys, about 40 requests a minute each: GLM 5.3 and 5.3 Flash, Kimi K3, DeepSeek V4.1 Flash, Nemotron 3 Ultra; testing only, and NVIDIA may keep what it is sent. Tried first, since its calls cost nothing; several keys paste in at once |
Gemini | ✅ | ✅ | Google's models; they can read images |
Groq | ✅ | ✅ | very fast open models, daily limits |
Cerebras | ✅ | ✅ | very fast open models, daily limits |
Mistral | ✅ | Mistral's own models, reached by name | |
DeepSeek | ✅ | direct, and prices halve off-peak | |
OpenRouter | ✅ | one account, hundreds of models; a few are free | |
Hugging Face | a router to many open models, reached by name; a small free monthly credit | ||
Cohere · Moonshot · DashScope · Zhipu · Perplexity | paid per use, reached by name | ||
anything OpenAI-compatible | Ollama, vLLM, LM Studio, Together, Azure, your own gateway: + Add provider |
A provider picks automatically when its models have published benchmark scores ZergSwarm ships with. The others work too: name a model (model: "command-a") or point a role at it; ZergSwarm just will not choose it on its own.
# ~/.zswarm/providers/ollama.toml: a local server as a provider
base_url = "http://127.0.0.1:11434/v1"
keys = ["ollama"] # a local server needs no key: any placeholder works
[models.llama-local]
api_id = "llama3.2"
ctx = 131072
price = { hit = 0, miss = 0, out = 0 }you want to | console | in the provider's file |
add a key | Providers › name: paste it |
|
use some keys first | give them a priority number (1 first; the same number takes turns) |
|
park one key | its switch in the key table | kept in |
cap what a day can cost | Routing & roles › Daily cap |
|
stop using a provider, keep its keys | its switch |
|
never use a model | the model's switch |
|
try some models first | the star, or a priority number |
|
point a role at a model | Routing & roles |
|
add a model | + Add model |
|
Priority reorders; it never lowers the bar. A model you add yourself has no published scores, so reach it by name, by a role, or through a route. A running server picks up file changes on its next call. The full field list is in docs/PROVIDERS.md.
⌨️ Command line
command | what it does |
| connect every assistant found here and open the console (what the installers run) |
| open the console (starts the local server if it is not running) |
| register with Claude Code only, or the clients |
| add a key, typed hidden; |
| what is configured and ready |
| one tool-free question from the terminal |
| run a batch from a file |
| follow and manage jobs |
| what the ledger says was spent |
| every command, with whether it reads, writes or spends |
🔧 Build from source
git clone https://github.com/Lunarwerx/ZergSwarm
cd ZergSwarm
pip install -e ".[test]" # an editable install with the test tools
pytest # the unit tests: offline, nothing touches your real ~/.zswarm
python -m build # the wheel and source archive, into dist/ (pip install build first)
python scripts/console_dev.py # the console against a scratch home, on port 7815python zswarm.py <command> also runs straight from a clone once httpx, jsonschema, mcp and tomlkit are
installed. A release is a tag: push v<version> and the release workflow tests,
builds, installs the wheel on Windows, macOS and Linux, and publishes it. Layout and conventions for
contributors, human or agent, are in AGENTS.md.
🔒 Security and privacy
The server listens on
127.0.0.1only, and refuses any request whose Host is not this machine, so a web page cannot drive it. API calls need the token in~/.zswarm/console-token. The console page opens without a login; on a machine other people share, setZSWARM_UI_SIGN_IN=1for a one-time sign-in link instead.Keys you add live in plain text in
~/.zswarm/providers/<name>.toml(owner-only on macOS and Linux), like most CLI tools' credentials. They are never logged, printed or returned: everything shows a fingerprint. Prefer environment variables? Leavekeysout and set<PROVIDER>_API_KEY.Prompts, and the files a worker reads, go to the provider that serves the task. Switch off any provider you do not want your code sent to. A provider's free tier may keep what it is sent under its own terms: set
ZSWARM_REDACT_FREE_TIER=onto mask secrets, email addresses and card numbers in tasks sent to free tiers, or use a paid key for private code.A worker with
editchanges files only inside its task's folder (plus any extrarootsthe task lists);allalso gives it a shell, and shell commands are not limited to that folder. The optionalccbackend (Claude Code as the worker) needsconfirm_writefor either and is not held to the folder, so treat it likeall. Give each task the narrowest set that does the job.ZergSwarm sends LunarWerx anonymous usage statistics: which command started, how many tasks a finished job had and how many came back ok, the version, OS and Python version, and a random install id. That is how we see what people use and where to spend our time.
ZSWARM_NO_PING=1switches it off.
❓ FAQ
What does it cost? ZergSwarm itself is free. You pay each provider directly, or nothing on a free tier. A small task uses about 3,000 tokens, so at $0.40 per million tokens it costs about a tenth of a cent; a review that reads a whole file, like the one above, costs a few cents. Every task's cost is in the console, and a daily cap you set (Routing & roles › Daily cap; there is none until you do) stops new work once it is reached.
How is this different from Claude Code's own subagents? Subagents run on Claude, so they draw on your Claude plan or bill. ZergSwarm workers run on other companies' cheaper models, several with free tiers, so fanning out over fifty files costs cents and leaves your Claude limits for the hard parts. Use subagents when a job needs Claude-level reasoning, and ZergSwarm for the wide, repetitive parts.
Which assistants does it work with? Claude Code (the CLI, the IDE extensions and the desktop app's Code tab), Claude Desktop, and Codex (CLI, IDE extension and desktop app). Anything else can use the local HTTP API. Exact client configs and the timeouts that matter are in docs/CLIENTS.md.
Do I have to pick models?
No. Leave everything on auto. Star a model only if you want it tried first.
Can it change my files?
Only when a task asks for it. With edit, a worker changes files only inside that task's folder (plus any extra
roots it lists). With all it also gets a shell, and a shell command can reach anything your user account can,
so use all only for tasks that must run commands. The optional cc backend needs confirm_write for either and
is not held to the folder, so treat it like all.
Where does my data go?
Your tasks go to the provider serving each one. The console and the ledger stay on your machine, and LunarWerx gets
the anonymous usage statistics described under Security and privacy. A free tier may keep what it is sent, so for
private code use a paid key or set ZSWARM_REDACT_FREE_TIER=on.
📄 License
Website: zergswarm.lunarwerx.com. Made by LunarWerx Studios. Check out sibling projects AgentHydra, RepoYeti, SageThumbs, and QuickDictate.
This server cannot be deployed
Maintenance
Related MCP Connectors
Real-time chat hub for AI agents — Claude Code, Cursor, Cline, Codex over MCP or REST.
Build and supervise fleets of agents from Claude Code, Codex or Cursor. Connects over OAuth.
AI routing, memory, guardrails, and governance. Routes across Claude, GPT, Gemini.
One MCP endpoint for Claude, GPT & Gemini: 100+ tools + no-code connectors + agent workers.
Related MCP Servers
- AlicenseNot gradedqualityDmaintenanceConnects Claude Code with multiple AI models (Gemini, Grok-3, ChatGPT, DeepSeek) simultaneously, allowing users to get diverse AI perspectives, conduct AI debates, and leverage each model's unique strengths.153MIT
- AlicenseAqualityAmaintenanceEnables one AI coding agent to delegate tasks to, and build consensus across, multiple other coding CLIs (Claude Code, Codex, etc.) by orchestrating them as headless subprocesses.1829 PyPI6MIT
- AlicenseNot gradedqualityDmaintenanceLets Claude Code query multiple AI models (Gemini, Grok, ChatGPT, DeepSeek) for diverse perspectives, code reviews, debates, and more.MIT
- AlicenseCqualityAmaintenanceRoutes coding tasks across multiple AI CLIs (Copilot, Claude Code, Gemini, etc.) with cost-aware tier routing and parallel wave orchestration.552Apache 2.0