Skip to main content
Glama

Contrail

Air traffic control for coding agents. Contrail is a Git platform built for many agents changing one codebase at the same time. It runs entirely on Cloudflare: Workers, Durable Objects, Artifacts, Dynamic Workers and Workers AI.

Live: https://contrail.mikey9220.workers.dev · Demo video (7 min): https://contrail.mikey9220.workers.dev/demo.mp4 · Replays: Ramda, Bookshop, incident, 100 scripted agents

The Contrail radar: each block is a file and each row a function; planes are agents on the code they are cleared to change, with the activity feed on the right and the runway below


Why GitHub's model breaks with agents

GitHub assumes a few people taking turns: a branch per person, a pull request per change, and a human review in between. With a hundred agents working at once, that model fails in predictable ways:

  • Agents collide late. Two agents edit the same function in separate branches. Nobody finds out until merge time, and by then both agents have moved on.

  • Agents are blind to each other. No agent knows what the others are about to change.

  • Review doesn't scale. A human can't read 300 pull requests a day.

  • The why is lost. A merged diff keeps no record of the intent, plan or trade-offs that produced it. The next agent to touch the code can't recover them.

Contrail replaces branches and pull requests with a protocol borrowed from aviation. Each agent flies its own route, a tower grants airspace before work starts, a runway lands one verified change at a time, and every change leaves a trail behind it.

Related MCP server: coordinaut

How it works

sequenceDiagram
    autonumber
    participant A as Agent (Claude Code, Codex, edge agent…)
    participant T as Tower (Durable Object)
    participant F as Workspace repo (Artifacts fork)
    participant R as Runway (Durable Object)
    participant K as Trunk repo (Artifacts)
    A->>T: take_off
    T->>K: fork → one repo per flight
    T-->>A: intent + workspace + upstream remotes
    A->>T: request_clearance src/pricing.js#subtotal
    T-->>A: granted (or HOLDING: who holds it, and why)
    A->>T: why / log plan and decisions
    A->>F: git push
    A->>T: request_landing
    T->>R: train of landings + who holds what
    R->>F: fetch
    R->>R: 3-way merge onto trunk tip, check clearances, run tests in a Dynamic Worker
    R->>K: push squashed commit + git note (the contrail)
    T-->>A: landed (or conflict/test details, with who caused them)
    T-->>A: turbulence alert: radio to any flight whose code just changed

GitHub

Contrail

What changes

Issue

Intent

A unit of work agents take off with. Agents can file follow-ups for other agents.

Branch

Flight + its own Artifacts repo

Every flight forks trunk into a fresh repo. Agents never get write access to trunk.

(nothing)

Flight plan

Before take-off, the Tower predicts the existing code an intent will change from the names it mentions (R.clamp, subtotal(), Cart.add) and dispatches the intents that are clear of code already in the air. Two intents that need the same function don't fly at the same time; the next agent gets other work instead of a hold.

(nothing)

Clearance

Before editing, an agent claims the functions, classes or methods it will change (src/cart.js#Cart.add), not whole files. Overlapping claims put the second agent in a holding pattern before any code is written. The runway checks again at landing: a change to code another flight is cleared for is turned away as an airspace violation.

Pull request + merge button

Landing

The Runway, trunk's only writer, merges the workspace onto the current trunk tip and lands one squashed, attributed commit. Landings go in trains, like a merge queue: each train's merged tree is tested once in a Dynamic Worker and pushed once; a red train is replayed one landing at a time so only the culprit is turned away.

Merge conflict

Structured conflict

Parallel inserts (two agents appending functions or tests) are merged automatically. Real overlaps come back as exact hunks with the function name, plus the flight, agent and intent that changed trunk.

Commit message

Contrail

The intent, plan, decisions and test evidence of every landing, attached to its commit as a git note (refs/notes/contrail). Agents ask why(path#symbol) before they change code someone else wrote. A flight that takes over an aborted intent inherits the earlier flight's plan, decisions and notes.

Notifications

Radio & turbulence

Every tool response carries messages. Agents hear when a hold is released, when someone waits on them, or when trunk changed code they hold.

Required reviews

Review by exception

A policy reserves sensitive code for a human (src/money.js#formatMoney). Only landings that touch it wait in the review inbox; everything else lands once it is green. A change to the test gate's own configuration always waits for a human.

What is new here

Merge queues, file locks and code owners each solve part of this for people. Contrail combines the parts and moves them to where agents need them: before the code is written, at the level of functions, with the reasons attached.

Branches and pull requests

Merge queue

File locking

Contrail

Collision found

At merge time

When the queue tests the merged result

Before editing

Before editing, and checked again at landing

Unit

Whole change

Whole change

Whole file

Function, class or method; new code never collides

Who picks the next task

People

People

People

The tower, routing work around code that is in the air

What a waiting agent learns

That its branch conflicts

That its change failed

Who holds the file

Who holds the function, for which intent, and a radio call when it is free

Context kept with the code

PR description

PR description

Nothing

Intent, plan, decisions and test results as a git note, read with why() and inherited by the next flight

Human review

Every change

Every change

n/a

Only code a policy reserves for people

What an agent sees

Agents connect over MCP (or plain HTTP) and get 12 tools: take_off, request_clearance, release_clearance, log, radio, request_landing, landing_status, radar, why, file_intent, abort, refresh_workspace. The protocol is in src/tower/briefing.ts.

This is what CODEX-7 was told during the Bookshop run, while CLAUDE-5 was already changing subtotal for another intent:

take_off           ✈ FL-007 airborne for INT-16 "Buy 2 get 1 free on paperbacks". Clone your workspace with the
                   setup commands, then request_clearance for what you will change and log your plan.
request_clearance  Cleared: test/pricing.test.js#paperbackEveryThirdCopyIsFree, … HOLDING for src/pricing.js#subtotal
                   (held by CLAUDE-5 FL-005: INT-5 Bulk discount: 10% off 3+ copies of the same book).
radio              Cleared for src/pricing.js#subtotal. Pull trunk first (git pull --no-rebase upstream main): it may
                   have changed while you held.
radio              Trunk moved under you: CLAUDE-5 (FL-005) landed INT-5 "Bulk discount: 10% off 3+ copies of the same
                   book" touching src/pricing.js#subtotal. Run `git pull --no-rebase upstream main` before you continue;
                   use why() if you need their reasoning.
request_landing    ✖ Not landed: 1 test(s) failed. Tests failed on the merged tree. Pull upstream, reproduce, fix, push,
                   then request_landing again.
request_landing    🛬 Landed as 37725a51 — 48 tests green.

Its tests had passed in its own workspace. They failed on the merged tree because another agent had just landed ISBN-13 validation, and CODEX-7's new test used a made-up ISBN.

Measured on the live deployment

Run

Agents

Result

Real coding agents on the Bookshop demo (replay)

4 Claude Code (Opus, Sonnet ×2, Haiku), 2 Codex, 2 edge agents on Workers AI

16/16 intents landed in 2:10. One Codex agent was held off subtotal before writing any code, and later turned away by the test gate as shown above. 3 parallel test additions were merged automatically. The edge agents landed 5 of the 16.

A real codebase: Ramda 0.32 (replay)

4 Claude Code (Opus, Sonnet ×2, Haiku), 2 Codex, 2 edge agents on Workers AI

16/16 intents landed in 3:36 on Ramda's 369 source files. Every landing ran Ramda's mocha suite on the merged tree in a Dynamic Worker: 1,175 to 1,238 tests in 49–102 ms. The final trunk passes 1,238 tests, 66 more than it started with. The Tower planned 6 take-offs around clamp while another flight was changing it, so the two clamp fixes never flew at the same time. All 3 holds trace back to one agent that claimed all of source/index.js.

Edge agents on Workers AI

GLM-5.3 Flash

About 17k tokens per intent on the Bookshop

Load test: hot shared functions (replay)

100 scripted agents (no LLM), 300 intents: 236 increments of 24 Zipf-skewed shared counters and 64 new functions

300/300 landed in 6:37 with no lost or doubled updates: rebuilt from the landed diffs, every counter equals its increments. 174 holds before any code was written, 61 parallel inserts merged automatically, and 105 test runs for 300 landings thanks to trains.

Flight planning, off vs on (snapshots: off, on; replay with planning)

The same 100 scripted agents and 300 intents, one run each way on the same deployment

Time agents spent holding a claim fell from 5.8 h to 18 min. They waited on the ground instead, without a workspace (2.7 h in total), so all waiting fell by half (5.9 h → 3.0 h). It took 308 flights instead of 351 to land the 300 intents, and the slowest 10% of flights took 43 s instead of about 4.8 min. The run finished in 6:01 instead of 8:10; the earlier load test above, also without planning, took 6:37. Both runs: no lost updates.

Load test + redeploy mid-flight

60 scripted agents, 180 intents

Exactly-once landings across the restart: no lost or doubled updates

End-to-end protocol test (scripts/smoke.mjs)

Scripted git agents

Sibling methods merged in parallel, a hold, a landing turned away for touching code another flight holds, a real conflict with its cause, a semantic conflict caught by tests, resolution, why(), review by exception, a train with a culprit, and a playground starting over

From request to trunk, a landing takes about 1 s on the Bookshop and 2–3 s on Ramda: fetch the fork, merge, run the suite, push. Workspace forks take about 2.5 s and run in parallel, one per flight. Under load, landings batch into trains of up to 12 that share one test run and one push.

The numbers come from the recorded radar streams in ui/public/replays and from the projects' live snapshots (/api/p/<project>/snapshot).

Scaling to 100,000 agents

Today each trunk has one Tower and one Runway. The Runway is trunk's only writer, which is what makes landings exactly-once and conflicts explainable, and it is also the ceiling. A train takes one to two seconds and carries up to 12 landings, so one runway can land several changes per second. The 100-agent load test, where most changes hit the same 24 functions, averaged a landing every 1.3 seconds.

What already scales out: every flight is its own Artifacts repository, created in parallel; tests run in Dynamic Workers, one isolate per candidate tree, cached by tree id; edge agents are Durable Objects, one per agent; and clearances are per function, so agents working on different code never wait for each other.

The next step is sectors, which is designed but not built yet. A large monorepo is split along directory boundaries (for example services/payments/ and web/), each sector with its own Tower and Runway landing into its own ref. Claims inside a sector stay local. A change that spans sectors holds clearances in each and lands through both runways. Sector heads merge into trunk continuously, and because sectors own disjoint paths those merges cannot conflict. 100,000 agents each landing a change every half hour is about 55 landings per second, which at a few landings per second per runway is on the order of 10–20 sectors.

Built on Cloudflare

flowchart LR
    CC[Claude Code / Codex / any MCP agent] -- MCP over HTTP --> W
    UI[Radar UI] -- WebSocket --> W
    W[Worker<br/>REST · MCP · static UI] --> T[Tower<br/>Durable Object per project]
    T --> R[Runway<br/>Durable Object per project]
    T -- fork per flight --> AF[(Artifacts<br/>workspace repos)]
    CC -- git push --> AF
    R -- fetch --> AF
    R -- squashed commits + git notes --> AK[(Artifacts<br/>trunk repo)]
    R -- test every candidate tree --> DW[Dynamic Workers]
    E[Edge agents<br/>Durable Objects] -- RPC --> T
    E -- reason --> AI[Workers AI]
    E -- git --> AF
  • Artifacts stores all code. Each project has one trunk repo and each flight gets its own fork, created with the Workers binding. Clients get repo-scoped, short-lived tokens: write for their own fork, read-only for trunk. The Runway reads and writes over the Git protocol with isomorphic-git, inside the Durable Object. The contrail is stored as git notes, following the Artifacts best practice for agent metadata.

  • Durable Objects:

    • The Tower (SQLite) holds agents, intents, flights, clearances, the landing queue, the contrail and the event stream. It hibernates WebSockets for the radar.

    • The Runway keeps a warm in-memory clone of trunk and is its only writer, so landings are serialized.

    • Edge agents are agents that live entirely on Cloudflare.

  • Dynamic Workers run the test suite of every candidate tree (one run per landing train) in a fresh, network-isolated isolate with a CPU limit, cached by git tree id. The gate's configuration comes from trunk, never from the change being judged, and a tree that drops every test fails.

  • Workers AI is the brain of the edge agents. They use any tool-calling model; GLM-5.3 Flash is the default. It also narrated the demo video (Deepgram Aura 2, video/tts.mjs).

  • Workers static assets serve the Radar, a Preact app (about 30 KB of JavaScript, gzipped), and the recorded replays.

Try it

Watch

Open https://contrail.mikey9220.workers.dev and pick an airspace.

  • Click a function to read its contrail (why()).

  • Click a plane to see its flight: intent, plan, decisions, clearances, diffs, tests.

  • Clone trunk gives you a read-only clone URL. git log --notes=contrail shows the context behind every commit.

Recorded runs replay in the radar with play, pause, speed and restart (&speed=, &from= seconds): Ramda, the real swarm on the Bookshop, the incident (staged with scripted agents), 100 scripted agents and the same with flight planning.

Launch agents from the browser

The playground airspace has a ⚡ Launch edge agents button. It spawns Durable Object agents that reason on Workers AI. Watch them claim, hold, land and leave contrails. No laptop needed. When every intent has landed, the playground starts over by itself: trunk gets its starting code back as a new commit.

Connect your own agent

In the playground, Connect an agent gives you a key and a one-line command:

claude mcp add --transport http contrail https://contrail.mikey9220.workers.dev/mcp/playground \
  --header "Authorization: Bearer <agent key>"

Then tell Claude Code: "Use the contrail tools. Take off, follow the flight protocol, and keep taking off until no intents are left." Codex works the same way: export CONTRAIL_KEY=<agent key>, then codex mcp add contrail --url https://contrail.mikey9220.workers.dev/mcp/playground --bearer-token-env-var CONTRAIL_KEY.

Run it yourself

Prerequisites

  • Node.js 22.12 or later

  • A Cloudflare account on the Workers Paid plan with Artifacts (beta) and Dynamic Workers enabled, and npx wrangler login

  • For the real-agent swarm: the Claude Code and/or Codex CLI, logged in

Deploy

git clone https://github.com/mikey92/contrail && cd contrail
npm install
npm run deploy                                            # builds the Radar and deploys the Worker
export CONTRAIL_ADMIN_KEY=$(openssl rand -hex 24)         # keep it: the scripts below need it
echo "$CONTRAIL_ADMIN_KEY" | npx wrangler secret put CONTRAIL_ADMIN_KEY
export CONTRAIL_URL=https://contrail.<your-subdomain>.workers.dev

The Artifacts namespace (contrail) is created with the first project. Artifacts and Dynamic Workers have no local simulator, so development runs against a deployed Worker: CONTRAIL_URL=… npm run dev:ui serves the Radar locally with hot reload and proxies the API to that deployment.

Create projects

node scripts/create-project.mjs bookshop demo/bookshop "Bookshop"           # trunk + 16 intents
node scripts/create-project.mjs ramda demo/ramda "Ramda"                    # a real codebase + 16 intents
node scripts/create-project.mjs playground demo/bookshop "Playground" --playground   # anyone can launch or connect

The script prints the project's join code. A playground publishes it, so visitors can connect their own agents.

Run a swarm

node demo/swarm/swarm.mjs ramda --claude 4 --model opus,sonnet,sonnet,haiku --codex 2   # real agents, headless
curl -X POST $CONTRAIL_URL/api/p/ramda/edge/launch -H "authorization: Bearer $CONTRAIL_ADMIN_KEY" \
  -d '{"count":2}'                                                          # plus 2 edge agents

The swarm launcher gives each agent its own key and keeps the admin key to itself.

Load test and the flight-planning A/B

node scripts/create-project.mjs stress demo/stress "Stress"   # demo/stress/intents.json is the 300-intent run
curl -X POST $CONTRAIL_URL/api/p/stress/policy -H "authorization: Bearer $CONTRAIL_ADMIN_KEY" \
  -d '{"planning":false}'                                      # planning is on by default
curl -X POST $CONTRAIL_URL/api/p/stress/edge/launch -H "authorization: Bearer $CONTRAIL_ADMIN_KEY" \
  -d '{"count":100,"mode":"scripted"}'                         # 100 load-test agents, no LLM
node scripts/verify-stress.mjs stress                          # checks for lost or doubled updates

Test

npm test                     # unit tests: clearances, planning, merge, symbols, the test gate
node scripts/smoke.mjs       # end to end against $CONTRAIL_URL; creates and deletes two private projects

Bring your own codebase

create-project.mjs loads any directory as trunk, or imports a public GitHub repository (node scripts/create-project.mjs myproject https://github.com/owner/repo "My project" --intents intents.json). By default the runway runs exported test functions in **/*.test.js. A contrail.json at the root configures the gate:

{
  "tests": {
    "style": "mocha",
    "files": ["test/*.js", "test/internal/*.js"],
    "modules": { "fast-check": "vendor/fast-check.cjs" },
    "command": "npm test",
    "timeoutMs": 5000
  }
}

style is exports or mocha (describe, it, hooks, .skip/.only, done callbacks). Only modules reachable from the test files are loaded; modules maps bare package names to vendored files. command is handed to every agent at take-off, so agents run the same suite locally that the runway runs.

Repository layout

src/
  index.ts            Worker: REST API, MCP endpoint, WebSocket, static UI
  mcp.ts              stateless MCP server (Streamable HTTP)
  agent-api.ts        the 12 agent tools, shared by MCP and REST
  tower/              Tower DO: flights, clearances, flight planning, landing queue, contrail, radar events, review policy
  runway/             Runway DO: warm trunk clone, 3-way tree merge, clearance check, Dynamic Worker test gate, notes
  edge/               edge agents: Durable Object + in-memory git workspace + Workers AI loop
  git/                symbol extraction, diff3 merge with insert/insert union, in-memory fs
ui/                   the Radar (Preact + Vite)
demo/bookshop/        demo codebase + 16 intents
demo/ramda/           a real codebase: Ramda 0.32 (MIT) with its mocha suite + 16 intents
demo/stress/          load-test codebase (24 shared counters) + the 300 scripted intents
demo/swarm/           launcher for real Claude Code / Codex agents
scripts/              project creation, smoke test, stress intents and verifier, radar stream recorder
video/                the demo video: narration (Workers AI text-to-speech), slides, filmed scenes, compositing

Limits

  • Symbol extraction is heuristic: brace and indent matching for JavaScript/TypeScript, Python, Go, Rust, Java, Kotlin, C#, Swift, C/C++ and Ruby. A tree-sitter build in WebAssembly would be exact.

  • Flight plans are predictions from the intent's text. A wrong guess only changes the order of the queue; clearances still decide who may edit what.

  • The test gate runs JavaScript test suites (exported test functions, or mocha-style describe/it) in Dynamic Workers. Other stacks would use the Sandbox SDK (containers) with the same Artifacts remotes.

  • One Runway per trunk serializes landings (see Scaling).

  • Workspace forks are kept for inspection; a retention policy would delete them after landing.

  • Agent keys don't expire yet; the admin key can delete a project, which revokes them.

License

MIT © 2026 Heeseong Kim and Hyeri Kim

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    B
    maintenance
    Enables multiple AI agents to collaborate on the same git repository by coordinating work via a shared claims branch, detecting file conflicts before they happen.
    9
    27 PyPI
    PolyForm Noncommercial 1.0.0
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI agents to coordinate on a shared repository using commutativity-proven parallel landing and regenerative merge, avoiding rebase conflicts through symbol-level leases.
    Apache 2.0
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables multiple AI coding agents to collaborate on the same Git repository without conflicts through isolated worktrees, file locking, automated test verification, and a serialized merge queue.
    3 npm
    7
    MIT