Skip to main content
Glama
sipyourdrink-ltd

Bernstein - Multi-agent orchestration

"To achieve great things, two things are needed: a plan and not quite enough time." - attributed to Leonard Bernstein

deterministic multi-agent CLI orchestration

CI PyPI GHCR Python 3.12+ License OpenSSF Scorecard CodeQL Open in Codespaces MCP Toplist

website · docs · install · first run · glossary · limitations · name policy · sponsor

简体中文 · 繁體中文


Status: beta. Solo-maintained, under active development. The version number counts releases, not maturity - minor versions may change interfaces. Pin the version for anything you depend on; regressions get fixed fast, file them.

Bernstein is a deterministic orchestrator for CLI coding agents (Claude Code, Codex, Gemini CLI, and 40+ more). It runs them in parallel, gates what they produce, and records enough of the run that you can check it afterwards. Air-gap install profile included. Apache-2.0.

at a glance

Four things set it apart; everything after is detail.

  • No LLM in the coordination loop. Scheduling is plain Python, so a run is reproducible end to end. Replay yesterday's plan and get yesterday's task graph.

  • Checkable after the fact. The replay journal records every run, and the always-on lineage spine records every lineage-bearing step; the opt-in HMAC-chained audit log (BERNSTEIN_AUDIT=1) adds receipts you verify offline. Non-determinism surfaces as a hash mismatch at the exact step, not a flaky re-run. Non-code deliverables get the same treatment. A task can declare an artifact contract — report, dataset, action log, ops result — from a plan step, a backlog entry, or the task CLI, and completes on a signed lineage receipt rather than a git commit.

  • Isolated by construction. Each coding task gets its own git worktree behind merge gates; artifact-mode tasks get working-directory separation under .sdd/workspaces/. Under this default isolation agents share no mutable workspace; coordination state (the task backlog) is shared and claimed atomically. Filesystem enforcement beyond that separation is opt-in, from the sandbox backends. Disable worktrees and every task runs in the shared checkout.

  • Broad and local. 40+ CLI agent adapters plus a generic --prompt wrapper, file-based state, no SaaS hop, no third-party data plane.

The full list is on the capabilities page; the feature matrix is the exhaustive index.

install in 30 seconds

uv tool install bernstein    # or: pipx install bernstein
bernstein init
bernstein doctor             # checks a CLI agent is installed and authenticated
bernstein -g "fix the failing test in tests/test_foo.py"

pipx, pip, brew, dnf, npm, and Docker are covered in the install guide; the air-gapped wheelhouse has its own air-gap guide.

The recording above is a real run, and it ships with its own proof. The cast, the signed run receipt derived from that run's journal, and the public key that pins it all live in docs/assets/demo-run/. Verify the run you just watched, offline:

bernstein verify receipt docs/assets/demo-run/run-receipt.json \
    --public-key docs/assets/demo-run/run-receipt.pub.pem

CI re-verifies the committed receipt on every push to main — and proves a tampered copy fails — so the published evidence cannot rot into a decorative file. scripts/record_demo.sh regenerates the recording, receipt, and key from a fresh real run; nothing inside the terminal is synthesised.

A run in flight is watchable from either operator surface. Both read the same task API, so neither is a lagging mirror of the other.

A two-column terminal dashboard - agents with their live logs on the left, the task board on the right - with a full-width activity feed and a cost line underneath

A browser dashboard listing sixty-two tasks with eleven running, one of them opened to its working-tree diff

bernstein live — the terminal dashboard

bernstein gui serve — the browser dashboard

prove a run

Determinism here is something you check, not something you take on faith. Run once with audit enabled, then verify what was recorded:

BERNSTEIN_AUDIT=1 bernstein -g "fix the failing test in tests/test_foo.py"
bernstein replay list                 # run ids recorded on disk
bernstein replay latest --verify      # recompute the journal head, name the first divergent step
bernstein lineage verify <run_id>     # recompute the always-on lineage spine
bernstein audit verify                # HMAC chain + Merkle seal (written because audit was enabled)
bernstein audit diagnose <run_id> --signal gate --sign-key KEY
                                      # name the exact step a failure entered the run, as a signed receipt
bernstein verify run <run_id> --signing-key-path key.pem   # sign one portable run receipt
bernstein verify receipt .sdd/runs/<run_id>/run-receipt.json  # verify it offline: file only

The journal is written on every run; the lineage spine is always on and gains an entry for each lineage-bearing step, so a short run can finish with a valid, empty spine. bernstein audit verify only has a chain to check when the run was started with BERNSTEIN_AUDIT=1, a compliance preset, or bernstein run --audit. The --audit flag belongs to bernstein run; on the bernstein -g form above, set the environment variable.

One run receipt binds the journal head, the lineage-spine head when the run wrote spine entries, and — opt-in — an audit-chain range, under a single Ed25519-signed subject with the public key embedded. A reviewer holding that file and the operator's public key can confirm the recorded actions are exactly what executed: no HMAC key, no live .sdd/, and exit 2 naming the first divergent step on tamper. With the file alone and no --public-key pin, the check is integrity-only — it proves the receipt is internally consistent, not who signed it, and the verdict says so. Details in deterministic replay.

The same checkability applies to evaluation numbers: bernstein bench run <suite> --reliability k (also spelled bernstein eval --reliability k) runs every task k times under fixed coordination and reports a pass^k floor (all k attempts must pass) alongside the pass@1 ceiling, sealed in a signed receipt that bernstein bench reliability-verify recomputes offline — a fabricated floor fails verification. Details: pass^k reliability floor.

how it works

Each goal moves through four stages:

  1. Decompose. The manager breaks your goal into tasks with roles, owned files, and completion signals. One LLM call, then plain Python from there.

  2. Spawn. Agents start in isolated git worktrees, one per coding task; an artifact-mode task gets a plain working directory instead. Main branch stays clean.

  3. Verify. The janitor checks concrete signals: tests pass, files exist, lint clean, types correct.

  4. Merge. Verified work lands in main. Failed tasks get retried or routed to a different model.

Why the scheduler is plain Python, and what that trades away: why deterministic.

everyday commands

cd your-project
bernstein init                    # creates .sdd/ workspace, bernstein.yaml + templates/
bernstein -g "Add rate limiting"  # agents spawn, work in parallel, verify, exit
bernstein live                    # watch progress in the TUI dashboard
bernstein run plan.yaml           # multi-stage plan: skip LLM planning, execute directly
bernstein stop                    # graceful shutdown with drain

The full operator surface (PR automation, schedules, chat bridges, the autofix daemon) is in operator commands.

Repository hygiene gates: bernstein readme-l10n verify fails a PR whose translated READMEs drifted from the English source (naming the stale section), bernstein readme-l10n sync rebinds them after an English edit. See readme-l10n.

supported agents

Claude Code, Codex CLI, Gemini CLI, GitHub Copilot CLI, Cursor, Aider, Goose, Muse Code, OpenAI Agents SDK, Amp, Cody, Continue, Devin Terminal, Junie, Kilo, Kiro, AWS Q Developer, Ollama, OpenCode, OpenHands, Open Interpreter, gptme, Plandex, AIChat, Letta Code, Qwen, and more. The adapter index carries install commands for 30 of them. bernstein integrations list enumerates all 51 wired-in integrations from src/bernstein/adapters/registry.py, the single source of truth for what resolves. 49 of them are selectable agent adapters; the other two rows are the mock test stub and the self-hosted-endpoints endpoint profile. Per-adapter copy lives in src/bernstein/adapters/use_cases.py. Anything else with a --prompt flag works through the generic wrapper.

Mix agents in the same run: cheap local models for boilerplate, heavier cloud models for architecture. bernstein integrations list --installed shows what is available on your machine.

beyond the front page

Everything deep lives on the docs site:

page

what it covers

capabilities

the full capability list: MCP server mode, signed agent cards, sandbox backends, artifact sinks, regulatory mappings

who this is for

where the value lands, and where Bernstein is the wrong tool

workflows

declarative YAML DAGs of agent / command / loop nodes

web UI

browser dashboard on the same API the TUI uses

cloud execution

experimental: run agents on Cloudflare Workers with R2 workspace sync against your own account. The hosted api.bernstein.run service is not yet available

datasources

read-only query receipts, plus a query driver that binds each result to the schema snapshot it was derived against

security

scorecard, fuzzing, hardening

architecture

how it works under the hood

why the name?

Bernstein is named after Leonard Bernstein, the American conductor and composer. The project orchestrates a crew of CLI coding agents the way Bernstein conducted the New York Philharmonic: every player on cue, the score deterministic, the conductor accountable for the result.

i wrote bernstein because i was paying $400/month in claude bills running three coding agents in parallel and getting nondeterministic merges. Apache 2.0, solo maintained. Live stats: bernstein.run.

mentioned in

Listed in vinta/awesome-python, covered in Augment Code's open-source agent orchestrators roundup, and listed in Python Weekly #742. We also wrote up the approach as the deterministic zero-LLM orchestration pattern in awesome-agentic-patterns.

The full tracked list, including every awesome-list entry, catalog listing, prior-art citation, and newsletter mention, lives in docs/mentions.md. Entries are added as they appear; corrections welcome by issue or PR.

contributing, support, license

PRs welcome; CONTRIBUTING.md has setup and code style. Security reports go through SECURITY.md. If Bernstein saves you time: GitHub Sponsors. Contact: forte@bernstein.run.

Citation metadata lives in CITATION.cff. License: Apache-2.0; the project name is covered separately in TRADEMARKS.md.


Alex Chernysh &middot; GitHub &middot; X &middot; bernstein.run

Install Server
A
license - permissive license
A
quality
A
maintenance

Maintenance

Maintainers
3dResponse time
0dRelease cycle
141Releases (12mo)
Commit activity
Issues opened vs closed

Related MCP Servers

  • -
    license
    B
    quality
    -
    maintenance
    Enables AI-driven orchestration of GitHub development workflows including automated issue analysis, code generation, code review, and PR creation through multiple specialized agents. Integrates with GitHub Actions to automate the complete development process from issue to pull request.
    7
  • F
    license
    -
    quality
    -
    maintenance
    A multi-agent runtime that coordinates six specialized agents through a typed artifact pipeline with 41 RPC methods. It features dynamic autonomy levels and context sufficiency scoring that adjust agent behavior based on the operator's state and task requirements.
  • A
    license
    A
    quality
    C
    maintenance
    Multi-agent orchestration server that enables parallel task delegation, sequential pipelines, cron scheduling, and cross-model peer review via CLI providers like Codex, Antigravity, OpenCode, and Claude Code.
    42
    26
    5
    MIT

View all related MCP servers

Related MCP Connectors

  • Build, validate, and deploy multi-agent AI solutions from any AI environment.

  • Coding agents from Claude Code, Cursor and Codex claim jobs and lock files on one shared board.

  • Create and manage AI agents that collaborate and solve problems through natural language interacti…

View all MCP Connectors

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/sipyourdrink-ltd/bernstein'

If you have feedback or need assistance with the MCP directory API, please join our Discord server