Skip to main content
Glama

agentbus

A small, single-machine coordinator for Claude Code and Codex. Keep the coding harnesses you already use, give each write candidate a Git worktree, compare results, and carry a task to the other harness without losing its working files.

MCP stdio server, local CLI, and an authenticated loopback web cockpit. No JS framework, build step, hosted service, or new model subscription.

agentbus demo: one bug, two harnesses, keep the better fix

Full video (23 s, 1080p, with sound). It shows one real, low-effort run: the same bug fanned out to Codex and Claude Code, both fixes passing the test, a handoff, and pick making a branch without touching the checkout.

The real agentbus cockpit, populated with fictional demo jobs

What ships

  • Fan-out and compare: run the same task against several agent/model choices. Candidates share one pinned base commit and get separate branches and worktrees. Optional tests run after each candidate, with timeout, output, exit code and duration.

  • Human merge: pick creates a separate committed branch and prints a merge command. It never switches, stages, commits in, or merges into your checkout.

  • Handoff: continue in the same worktree on either harness. A fresh session gets the original task, capped previous answer, diff versus base, touched files and relevant board notes. Parent links make a visible thread.

  • Private context: longest resolved path prefix matches a project block in a private YAML org map. Personal rules and project boundaries enter every job prompt.

  • Cockpit: jobs, SSE log tails, diffs, candidate comparison, board, org map and usage. The UI can start read jobs, cancel, pick, discard and post local notes.

  • Usage and routing: CLI-reported token/cost facts, plus deterministic model suggestions that never start a job themselves.

Related MCP server: session-coord-mcp

How it compares

This is a personal bridge with a cockpit, not a desktop ADE replacement. Comparison checked against Orca's repository and Spotify's Xirp introduction on 26 September 2026. It is a scope comparison, not a performance benchmark.

Capability

agentbus v0.2

Orca

Xirp + Portal

Parallel isolated work

Git worktree per candidate, queued caps

Parallel agents in worktrees

Worktree per session, 50+ parallel sessions described

Compare and keep changes

Tests, diff previews, pick branch, human merge

Compare and merge workflow

Concurrent cross-harness workstreams

Change harness mid-task

New session with a capped handoff packet and same files

Multiple agent integrations

Context and working state decoupled from the harness

Organizational context

Private hand-maintained project YAML and board

Project/agent workspace tooling

Catalog, ownership, dependencies and architectural decisions

Interface

CLI, MCP, local responsive web UI

Desktop ADE and mobile companion

ADE with Portal context integration

Machines

One Linux/POSIX machine

Desktop and remote runtimes

Local and remote sessions

Usage/routing

Reported job usage, keyword suggestions

Account/usage tooling

Model flexibility and price-performance routing

No desktop IDE, SSH worktrees, remote execution, mobile app, notifications, inline diff annotations, Gemini adapter, account quota polling, or automatic model selection. The responsive UI does not make this a phone-accessible service: it binds to this machine's loopback only. No tunnel or remote-access mechanism is included or enabled.

Install

Python 3.11+ on Linux, Git, and locally authenticated codex and claude CLIs. The current Codex adapter requires --ignore-user-config, --ignore-rules, --ephemeral, and --json; check codex exec --help on older installations.

From your local clone:

python3 -m venv .venv
.venv/bin/python -m pip install -r requirements-dev.txt
mkdir -p ~/.config/agentbus
cp PERSONAL.example.md ~/.config/agentbus/PERSONAL.md
cp orgmap.example.yaml ~/.config/agentbus/orgmap.yaml
chmod 600 ~/.config/agentbus/PERSONAL.md ~/.config/agentbus/orgmap.yaml
.venv/bin/python -m pytest -q
bin/agentbus jobs

Edit those private copies for your machine. Do not overwrite existing private context when upgrading. Runtime-only installs can use requirements.txt. bin/agentbus resolves its repository path, including through a future symlink; it is not installed or symlinked by this project.

Register MCP manually if needed, substituting the absolute clone path:

claude mcp add --transport stdio --scope user agentbus -- /absolute/path/agentbus/.venv/bin/python /absolute/path/agentbus/agentbus.py
codex mcp add agentbus -- /absolute/path/agentbus/.venv/bin/python /absolute/path/agentbus/agentbus.py

The server still starts in MCP stdio mode with no arguments. Existing registrations pointing to agentbus.py work after restarting the harness/MCP connection; no configuration-file edits are needed for the upgrade.

CLI

bin/agentbus run codex "Review the cancellation path" --cwd /path/to/repo --model gpt-5.6-luna --effort low
bin/agentbus run claude "Fix the parser" --write --cwd /path/to/repo
bin/agentbus fanout "Fix the parser" --cwd /path/to/repo --agents codex:gpt-6-astra,claude:opus --test "python -m pytest -q"
bin/agentbus jobs
bin/agentbus show JOB_ID
bin/agentbus compare FANOUT_ID
bin/agentbus continue JOB_ID "Review the edge cases" --agent claude --model haiku --effort low
bin/agentbus pick FANOUT_ID 1
bin/agentbus discard FANOUT_ID
bin/agentbus usage --days 7
bin/agentbus route "Quick explanation of this function"
bin/agentbus serve

run returns immediately; --wait 50 optionally waits up to 55 seconds. Jobs survive CLI and MCP disconnection. Candidate numbers start at 1. fanout defaults to write mode and requires a committed Git repository. --read still creates independent worktrees. A supplied test command is trusted shell code; only supply commands you are willing to run locally.

serve prints one random bootstrap URL. Open it once; it redirects to / and uses an HttpOnly SameSite=Strict session cookie. Every request, including fonts, icons and SSE, requires that cookie. Restarting the server invalidates it. Port defaults to 8765; --port can change it. All accepted loopback host arguments are normalized to 127.0.0.1, and non-loopback arguments are refused. Access logging is disabled so bootstrap tokens do not enter request logs.

MCP tools

The original tool names and parameters remain available.

Tool

Parameters / purpose

delegate

agent, task, cwd="", model="", effort="high", mode="read", isolate=True

job_result

job_id, wait_seconds=50; waits at most 55 seconds

job_list

limit=10; latest records with thread and parent links

job_cancel

job_id; queued cancellation or verified process termination

board_post

note, author, topic=""; local shared notes

board_read

topic="", limit=20

personal_context

Private standing rules

fanout

task, agents, cwd, test_cmd="", mode="write"

compare

fanout_id; table, reported facts and first 40 diff lines

pick

fanout_id, n; creates agentbus/pick-<id> with author agentbus

discard

fanout_id; removes unpicked worktrees/branches, preserves records

continue_job

job_id, instruction, agent="", model="", effort=""

org_context

path; matching private project block and global rules

usage

days=7; UTC day, agent and model aggregates

route

task; transparent suggestion, rule table and queue load

Agent specs look like [{"agent":"codex","model":"gpt-6-astra","effort":"high"},{"agent":"claude","model":"opus"}]. Supported model names are explicit in agentbus_core/harness.py; availability still depends on the installed CLI and subscription. There is no silent model fallback.

State, queue and continuation

By default, runtime state lives beside agentbus.py under ignored jobs/, worktrees/, fanouts/ and board.jsonl. AGENTBUS_HOME relocates it. AGENTBUS_CONFIG_HOME relocates private config (default ~/.config/agentbus). Personal context loads from the config directory first, then the old AGENTBUS_HOME/PERSONAL.md path for compatibility.

Environment variable

Default

Meaning

AGENTBUS_MAX_CLAUDE

1

Concurrent Claude jobs

AGENTBUS_MAX_CODEX

2

Concurrent Codex jobs

AGENTBUS_MAX_GLOBAL

3

All concurrent jobs, including post-job tests

AGENTBUS_JOB_TIMEOUT

1800

Seconds allowed for the harness process

AGENTBUS_TEST_TIMEOUT

120

Seconds allowed for each test command

AGENTBUS_CALLER

detected or unknown

Caller harness for review routing

AGENTBUS_EXTRA_PATH

empty

Colon-separated directories appended to PATH for spawned CLIs, e.g. where node lives

AGENTBUS_PASS_ENV

empty

Comma-separated credential-looking variable names that delegates may still receive

AGENTBUS_SCRATCH

${TMPDIR:-/tmp}

Root that scripts/demo.py and scripts/smoke.py must write inside

If AGENTBUS_EXTRA_PATH is unset, agentbus reads the same colon-separated list from the private file extra_path in the config directory. That lets an MCP registration without an env block still find a node or CLI that sits outside the harness PATH.

Use the same cap settings in all entrypoints sharing a home. Limits are at least one. The queue is FIFO among eligible jobs: a saturated Claude slot does not block Codex. A detached lightweight supervisor waits for each queued job; fan-out requests accept at most 16 candidates. API instances share an advisory file lock and atomic JSON records. A dead supervisor is marked lost, not silently rerun.

Completion captures tracked and non-ignored untracked files through a temporary Git index. The candidate's index stays unchanged. Snapshots are retained under refs/agentbus/snapshots/<job_id>; discard removes worktrees and candidate branches, not these snapshots or job history. Large artifacts should be Git-ignored. There is no automatic history-retention policy.

A continuation requires a terminal parent and no competing directory user. Continue the latest job in a thread; forks of one mutable worktree are refused. Pick freezes the selected chain, and compare follows the latest continuation of each candidate. Changes made to a worktree after completion cause pick to refuse until a continuation captures and tests them. Test status is separate from harness status: done can have failing tests. Pick is permitted after reviewing a failed candidate too.

Handoff is a new harness session, not a transfer of internal conversation state. Caps: original task 12,000 characters, prior answer 12,000, diff 24,000, touched files 8,000 and board notes 8,000. Complete saved records/diffs remain on disk. Read jobs and non-isolated/non-Git jobs do not have an isolated Git diff. Historical v0.1 jobs remain readable but lack the base snapshots needed for continuation; active legacy wrappers are counted toward caps and can be cancelled only when their unique output paths verify their process identity.

Usage and routing rules

Codex JSON events report turn.completed.usage; the adapter sums numeric fields from those events. Claude's final JSON provides usage and total_cost_usd when available. The real CLI smoke confirmed token counts from both; Codex did not report a cost.

Unknown usage is null in JSON and not reported in the UI. Aggregates expose how many jobs reported each metric, including partial totals. Token totals add input and output tokens; Claude cache-read/cache-creation inputs are separate and are added, while Codex cached-input counts are already included in its input total. Reasoning/output detail fields are not counted twice. Reported dollar values are API-equivalent CLI values, not subscription bills or remaining quota.

Routing uses the first matching case-insensitive keyword rule, in this order:

Match

Suggestion

review, second opinion, second-opinion, red-team

Other harness: Claude opus/high or Codex gpt-6-astra/high; unknown caller defaults to Claude

refactor, migration, migrate, architecture, large change

Codex gpt-6-astra/xhigh

quick, small question, one line, one-line, explain, summarize

Codex gpt-5.6-luna/low; Claude haiku/low if Codex has more queued jobs

anything else

Codex gpt-5.6-sol/high

The caller comes from AGENTBUS_CALLER, otherwise CLAUDECODE or CODEX_THREAD_ID when present. Queue load accompanies every suggestion. No LLM call, automatic execution, pricing lookup or claimed model-quality benchmark is involved.

Safety model and limits

  • For trusted local use. A Git worktree protects checkouts from ordinary collisions; it is not an OS security boundary against hostile generated code.

  • Codex runs with user config/rules ignored, MCP and app integrations disabled, and a read-only or workspace-write sandbox. Claude uses strict empty MCP config, no settings sources, no slash commands/Chrome integration, and an explicit tool allowlist. Claude write mode uses bypassPermissions and has no OS sandbox.

  • Every environment variable whose name looks like a credential (a KEY, TOKEN, SECRET, PASSWORD, PASSWD, PASSPHRASE, CREDENTIAL(S) or PAT name part, or one of those suffixes) is removed from delegates, supervisors and test commands. CLIs use their on-disk subscription login. AGENTBUS_PASS_ENV is the explicit escape hatch, for example for CLAUDE_CODE_OAUTH_TOKEN on a headless machine.

  • Recursion: delegates get no MCP servers, AGENTBUS_DEPTH=1, and every delegate/fanout/continue/pick/discard call also refuses when any ancestor process is an agentbus job supervisor, so clearing the variable is not enough. A write delegate that deliberately daemonizes out of its process tree can still escape; that is the same trust limit as any unsandboxed Claude write job.

  • pick and discard only touch the worktree worktrees/<job> and branch agentbus/<fanout>/<n> that agentbus created for that candidate. Records naming anything else are refused before any removal. Bash in a write job can still reach the network or invoke other programs; prompt rules and disabled connectors do not make this an egress firewall.

  • Read jobs prohibit edits/commands in the prompt; Claude only gets Read/Glob/Grep. Legacy explicit isolate=False and write jobs outside Git preserve v0.1 behavior and operate in place. Use default isolation inside a committed repository.

  • Paperclip editing is prohibited in every prompt and direct write starts in its data directory are rejected. Private org-map boundaries are injected instructions, not a filesystem allowlist or an automatically discovered dependency graph.

  • Cancellation and timeout terminate owned process groups, including normal child processes. Linux boot/start identities guard against reused PIDs. This does not contain deliberately escaped daemons or hostile local processes.

  • UI forms use CSRF tokens, reject cross-site requests and foreign Origins, and accept Origin: null only with Sec-Fetch-Site: same-origin. Output is HTML-escaped, with restrictive CSP, no framing and Referrer-Policy: same-origin to prevent cross-origin Referer leakage; cockpit actions were verified in real headless Chrome. Known credential patterns are scrubbed in capture logs and UI text, including answers, diffs and board notes. The scrubber is heuristic, not a guarantee that arbitrary secrets can be detected. Raw prompts and Git snapshots are private local state and must not be shared blindly.

  • The local HTTP cookie deliberately has no Secure flag because this is an HTTP loopback server. Do not proxy or expose the cockpit. It has no multi-user model.

  • Browsers do not scope cookies by port, so the session cookie is also sent to any other service you open on 127.0.0.1. Only open local services you trust in the same browser profile while the cockpit is running, and restart serve to rotate it.

PERSONAL.md and orgmap.yaml live in private config and are Git-ignored. Keep your own copies out of any repository you publish: deleting a file in a later commit does not remove it from Git history.

Design and verification

Native server-rendered Jinja HTML, one CSS file and a small vanilla JS file. Geist and Geist Mono WOFF2 files are vendored from the official Vercel v1.7.2 release, with their SIL OFL license. Phosphor SVGs retain their MIT license. No font download fallback was needed. No runtime CDNs, remote assets, analytics or tracking.

Dark by default, automatic system light preference and a persistent manual theme switch. One signal-orange accent, zinc neutrals, semantic pass/fail colors, 6px corners, mono tabular figures, reduced-motion support and single-column mobile views.

Verification and real smoke outputs. Screenshots come from the real server running fictional jobs in a separate scratch home: desktop, phone, comparison, phone comparison, light mode. No private jobs or board notes are in them.

To create your own throwaway demo (no real harness calls). The demo runs with a scratch HOME, so every path it shows is ~/projects/demo or ~/agentbus/...:

export AGENTBUS_SCRATCH="$(mktemp -d "${TMPDIR:-/tmp}/agentbus-demo.XXXXXX")"
export HOME="$AGENTBUS_SCRATCH/demo-home" AGENTBUS_HOME="$AGENTBUS_SCRATCH/demo-home/agentbus"
export AGENTBUS_CONFIG_HOME="$HOME/.config/agentbus"
.venv/bin/python scripts/demo.py
bin/agentbus serve

Run those in a subshell or a separate terminal, because they change HOME. Both scripts refuse any AGENTBUS_HOME outside the scratch root: $AGENTBUS_SCRATCH, or a directory named agentbus-* directly under ${TMPDIR:-/tmp}. The demo marker visibly labels its fictional data. scripts/smoke.py is separate, explicitly invokes real subscription CLIs (keep your real HOME), and is never run by pytest or the demo:

AGENTBUS_HOME="$(mktemp -d "${TMPDIR:-/tmp}/agentbus-smoke.XXXXXX")/state" .venv/bin/python scripts/smoke.py

Code: Apache-2.0. Vendored fonts/icons retain their respective licenses.

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    F
    maintenance
    Enables orchestrating multiple AI CLI agents (Claude Code, Codex, Gemini CLI, Copilot CLI) through a unified MCP interface for task delegation, cross-agent comparison, and specialized tools like code review and debugging.
    14
    7 npm
    14
    MIT
  • A
    license
    Not graded
    quality
    A
    maintenance
    An MCP server for orchestrating a fleet of CLI coding agents in isolated git worktrees. It exposes tools for spawning workers, sending instructions, reviewing diffs, and merging changes, with full terminal visibility.
    6 npm
    4
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    An MCP server and HTTP gateway that dispatches coding tasks to Claude Code, Codex, and Antigravity CLI sessions, keeps track of those sessions, and optionally runs an implement → review → verify loop against a real repository.
    6
    MIT