Skip to main content
Glama
sipyourdrink-ltd

Bernstein - Multi-agent orchestration

"To achieve great things, two things are needed: a plan and not quite enough time." — Leonard Bernstein

orchestrate any AI coding agent. any model. one command.

CI PyPI GHCR Python 3.12+ License CodeTrendy

website · docs · install · first run · enterprise eval · glossary · limitations · sponsor


Bernstein is a deterministic Python scheduler that runs a crew of CLI coding agents (Claude Code, Codex, Gemini CLI, and 40 more) against a single goal in parallel git worktrees, with an HMAC-signed audit chain over every step.

at a glance

  • 43 CLI agent adapters ship in v1.10.1 — 40 third-party wrappers, 2 leaf-node delegators, plus a generic --prompt wrapper. Source of truth: the supported agents table below.

  • HMAC-SHA256 audit chain per RFC 2104, one record per scheduling decision, tamper-evident. Operator guide: docs/security/audit-log.md.

  • Signed agent cards use detached JWS (RFC 7515 §A.5) over RFC 8785 (JCS) canonicalization, with Ed25519 / EdDSA keys. Code: src/bernstein/core/security/agent_card_signer.py.

  • Per-artefact lineage records every file write linked back to producer + inputs + prompt SHA + model + cost; customer-key signing for DORA / NIS2 / EU AI Act Article 12 evidence. CLI: bernstein lineage verify <run_id>.

  • Deterministic scheduler: zero LLM in the coordination loop. Plain Python decides who runs, where, with what budget. Replay yesterday's plan, get yesterday's task graph.

why this exists

i wrote bernstein because i was paying $400/month in claude bills running three coding agents in parallel and getting nondeterministic merges.

as of 2026-05-08: 296 stars, 35 forks, ~3,769 pypi downloads/day (mostly bots; ~54k/month), apache 2.0, solo maintained, no funding. numbers will drift; the line above is the source-of-truth date — re-run pip stats / GitHub API to refresh.

install in 30 seconds

pipx install bernstein
bernstein init
bernstein run -g "fix the failing test in tests/test_foo.py"

sponsor

if bernstein routed a model that saved you a claude bill, $25 covers a month of my coffee.

github.com/sponsors/chernistry →

tier ladder, escalation thresholds, and what each tier gets you live at bernstein.run/sponsors.

Related MCP server: agent-orchestration

who this is for

specific shapes where the value lands:

  • engineering teams running ≥3 cli coding agents in parallel — each agent gets its own git worktree, the merge queue serialises landings, no race conditions

  • regulated or on-prem environments — every routing decision is in plain text, the audit log is hmac-signed and tamper-evident, no saas hop, no third-party data plane

  • platform teams that need an audit log of agent decisions — the orchestrator writes one row per scheduling decision, you can grep it

  • anyone burning more than $1k/mo on cursor/aider/claude-max who wants determinism — you can replay yesterday's plan and get yesterday's task graph

  • forward-deployed engineers dropping into a client repo — credentials stay in your env, not the client's; agents you spawn are whichever cli tool the client already trusts

if you nodded at two of those bullets, this fits.

who this is NOT for

equally specific. these are the cases where you should pick something else:

  • "i want one pair-programmer to chat with about my code" — claude code or cursor alone. bernstein adds orchestration overhead you don't need

  • prototypes where merge gates are overkill — the lint/types/tests/cross-model-review pipeline is value when the cost of a bad merge is real, friction when you're throwing the repo away on friday

  • non-coding tasks (research, writing, data analysis pipelines) — bernstein wraps cli coding agents specifically, not generic llm workflows. crewai or autogen are the right shape there

  • anyone who wants a saas wrapper with a credit card form — bernstein is on-prem only by design. if you want managed, this is the wrong project, not the wrong fit

  • teams that need a vendor with a support sla and a contract — solo open-source project. github issues are how support happens

  • research-shape "let the agents collaborate emergently" use cases — the deterministic scheduler is a hard wall there

how it compares

Feature

Bernstein

Archon

LangGraph

Deterministic scheduler (no LLM in loop)

yes

no

no

Multi-agent crew (parallel adapters)

yes

one

yes

Signed lineage / audit chain

yes

no

no

Air-gap / sovereign deploy

yes

partial

no

Visual workflow YAML

yes yaml

yes

no

Hosted dashboard / SaaS

no

partial

no

A longer feature matrix against CrewAI, AutoGen, LangGraph, and the four CLI-agent orchestrators that share Bernstein's category lives in the Detailed comparison section below.


what is this, in one paragraph

You tell Bernstein what you want built. It splits the work across several AI coding agents, runs them in parallel inside isolated git worktrees, records every handoff in an HMAC-SHA256-chained audit log (RFC 2104), runs the tests, and merges the code that actually passes. You come back to a green PR.

Forward-deployed engineering, on a swarm. Drop Bernstein into a client repo and you get a multi-agent crew with file-based state (.sdd/), per-agent credential scoping, and a signed audit trail running on whichever CLI agents the client already trusts.

Cited as the "deterministic zero-LLM orchestration" pattern reference implementation in nibzard/awesome-agentic-patterns and "the most architecturally interesting tool" by Augment Code's open-source agent orchestrators roundup (2026).

other install methods

curl -fsSL https://bernstein.run/install.sh | sh        # macOS / Linux one-liner
irm https://bernstein.run/install.ps1 | iex             # Windows PowerShell
pip install bernstein                                   # pip
uv tool install bernstein                               # uv
brew tap chernistry/tap && brew install bernstein       # Homebrew

See the full install matrix for dnf copr, npx, optional extras, and the wheelhouse path for air-gapped sites.

why the scheduler is plain Python

Most agent orchestrators use an LLM to decide who does what. That is non-deterministic and burns tokens on scheduling instead of code. Bernstein does one LLM call to break down your goal, then the rest (running agents in parallel, isolating their git branches, running tests, routing retries) is plain Python. Every run is reproducible. Every step is logged and replayable.

No framework to learn. No vendor lock-in. Swap any agent, any model, any provider.

What you see while it runs:

$ bernstein -g "Add JWT auth"
[manager] decomposed into 4 tasks
[agent-1] claude-sonnet: src/auth/middleware.py  (done, 2m 14s)
[agent-2] codex:         tests/test_auth.py      (done, 1m 58s)
[verify]  all gates pass. merging to main.

YAML workflow manifests (optional)

When the open-ended bernstein run -g "<goal>" is too coarse-grained, the bernstein workflow family runs a declarative DAG of agent / command / loop nodes. Manifests are plain YAML, validated up-front, and dispatched through the same AgentSpawner the rest of Bernstein uses. No parallel spawn path, no LLM in the scheduler.

bernstein workflow list                       # bundled + user-installed
bernstein workflow run idea-to-pr -g "Add JWT auth"
bernstein workflow init my-flow               # scaffold a starter manifest
bernstein workflow validate path/to/flow.yaml

Stock workflows that ship with the wheel:

Name

What it does

idea-to-pr

research → plan → implement → tests → PR

refactor-with-tests

find target → propose → implement → loop until green

security-review

scan → triage → patch → adversary review

doc-update

audit → update → docs build

dependency-bump

bump → install → tests-loop → smoke

hot-fix

reproduce → fix → regression loop → changelog

Loop nodes re-fire until a bash predicate exits 0 (pytest -x is a typical one). fresh_context: true mints a new agent session per iteration. The interactive: true flag is reserved for the approval-gate work tracked in ticket #1110 and currently raises a clear NotImplementedError.

use cases

  • forward-deployed engineering — drop the swarm onto a client repo when you arrive, take it with you when you leave.

  • self-evolving projects — point Bernstein at its own repo and let it execute the backlog (this codebase is one).

  • CI fleets — run a swarm of agents in parallel on PRs, with per-agent credential scoping and signed audit trail.

  • air-gapped / regulated deployment — install from a signed wheelhouse, run with --profile airgap to deny outbound by default, allow-list specific destinations as needed. See Air-gap installation.

supported agents

Bernstein auto-discovers installed CLI agents. Mix them in the same run. Cheap local models for boilerplate, heavier cloud models for architecture.

43 CLI agent adapters: 40 third-party wrappers, 2 leaf-node delegators (Composio, Ralphex), plus a generic wrapper for anything with --prompt.

Agent

Models

Install

Claude Code

Opus 4, Sonnet 4.6, Haiku 4.5

npm install -g @anthropic-ai/claude-code

Codex CLI

GPT-5, GPT-5 mini

npm install -g @openai/codex

OpenAI Agents SDK v2

GPT-5, GPT-5 mini, o4

pip install 'bernstein[openai]'

GitHub Copilot CLI

Copilot-managed (GPT-5, Sonnet 4.6)

npm install -g @github/copilot

Gemini CLI

Gemini 2.5 Pro, Gemini Flash

npm install -g @google/gemini-cli

Cursor

Sonnet 4.6, Opus 4, GPT-5

Cursor app

Devin Terminal (Cognition)

Devin-managed

curl -fsSL https://cli.devin.ai/install.sh | bash then devin auth login

Aider

Any OpenAI/Anthropic-compatible

pip install aider-chat

Amp

Amp-managed

npm install -g @sourcegraph/amp

CLM gateway (sovereign / on-prem LLM)

Any OpenAI-compatible CLM endpoint

pip install aider-chat, then set CLM_ENDPOINT / CLM_TOKEN

Cody

Sourcegraph-hosted

npm install -g @sourcegraph/cody

Continue

Any OpenAI/Anthropic-compatible

npm install -g @continuedev/cli (binary: cn)

Goose

Any provider Goose supports

See Goose docs

IaC (Terraform/Pulumi)

Any provider the base agent uses

Built-in

Junie

BYOK (Anthropic, OpenAI, Google, xAI, OpenRouter, Copilot)

curl -fsSL https://junie.jetbrains.com/install.sh | bash

Kilo

Kilo-hosted

See Kilo docs

Kiro

Kiro-hosted

See Kiro docs

AWS Q Developer

Amazon Q-managed (Claude-backed)

brew install --cask amazon-q then q login

Ollama + Aider

Local models (offline)

brew install ollama

OpenCode

Any provider OpenCode supports

See OpenCode docs

Qwen

Qwen Code models

npm install -g @qwen-code/qwen-code

Cloudflare Agents

Workers AI models

bernstein cloud login

OpenHands

Any LiteLLM-supported (Anthropic, OpenAI, ...)

uv tool install openhands --python 3.12

Open Interpreter

Any (LiteLLM-backed)

pip install open-interpreter

gptme

Anthropic, OpenAI, OpenRouter

pipx install gptme

Plandex

Plandex Cloud or self-hosted models

curl -sL https://plandex.ai/install.sh | bash

AIChat

OpenAI, Anthropic, OpenRouter, Groq, Gemini

cargo install aichat

Letta Code

Letta-routed (Anthropic, OpenAI)

npm install -g @letta-ai/letta-code

Generic

Any CLI with --prompt

Built-in

orchestrator delegation (leaf-node)

A separate, smaller class of adapters that wrap other CLI orchestrators as if they were single agents. Bernstein hands the wrapped tool a prompt or plan and only sees the final exit code; sub-agent costs and quality gates inside the wrapped orchestrator are not visible to Bernstein. Useful when you want to drop an existing workflow built on one of these tools into a step of a larger Bernstein plan.

Orchestrator

Wrapped as

Install

Composio Agent Orchestrator (@aoagents/ao)

composio

npm install -g @aoagents/ao

umputun/ralphex

ralphex

go install github.com/umputun/ralphex/cmd/ralphex@latest

Any adapter also works as the internal scheduler LLM. Run the entire stack without any specific provider:

internal_llm_provider: gemini            # or qwen, ollama, codex, goose, ...
internal_llm_model: gemini-3.1-pro
TIP

Runbernstein --headless for CI pipelines. No TUI, structured JSON output, non-zero exit on failure.

quick start

cd your-project
bernstein init                    # creates .sdd/ workspace + bernstein.yaml
bernstein -g "Add rate limiting"  # agents spawn, work in parallel, verify, exit
bernstein live                    # watch progress in the TUI dashboard
bernstein stop                    # graceful shutdown with drain

For multi-stage projects, define a YAML plan:

bernstein run plan.yaml           # skips LLM planning, goes straight to execution
bernstein run --dry-run plan.yaml # preview tasks and estimated cost

how it works

Bernstein runs a four-stage pipeline per goal:

  1. Decompose. The manager breaks your goal into tasks with roles, owned files, and completion signals. One LLM call, then plain Python from there.

  2. Spawn. Agents start in isolated git worktrees, one per task. Main branch stays clean.

  3. Verify. The janitor checks concrete signals: tests pass, files exist, lint clean, types correct.

  4. Merge. Verified work lands in main. Failed tasks get retried or routed to a different model.

The orchestrator is a Python scheduler, not an LLM. Scheduling decisions are deterministic, auditable, and reproducible. Every step writes a record to the HMAC-chained audit log (.sdd/audit/YYYY-MM-DD.jsonl) per RFC 2104 — see docs/security/audit-log.md.

cloud execution (Cloudflare)

Bernstein can run agents on Cloudflare Workers instead of locally. The bernstein cloud CLI handles deployment and lifecycle.

  • Workers. Agent execution on Cloudflare's edge, with Durable Workflows for multi-step tasks and automatic retry.

  • V8 sandbox isolation. Each agent runs in its own isolate, no container overhead.

  • R2 workspace sync. Local worktree state syncs to R2 object storage so cloud agents see the same files.

  • Workers AI (experimental). Use Cloudflare-hosted models as the LLM provider, no external API keys required.

  • D1 analytics. Task metrics and cost data stored in D1 for querying.

  • Browser rendering. Headless Chrome on Workers for agents that need to inspect web output.

  • MCP remote transport. Expose or consume MCP servers over Cloudflare's network.

bernstein cloud login      # authenticate with Bernstein Cloud
bernstein cloud deploy     # push agent workers
bernstein cloud run plan.yaml  # execute a plan on Cloudflare

capabilities

Core orchestration. Parallel execution, git worktree isolation, janitor verification, quality gates (lint, types, PII scan), cross-model code review, circuit breaker for misbehaving agents, token growth monitoring with auto-intervention.

Intelligence. Contextual bandit router for model/effort selection. Knowledge graph for codebase impact analysis. Semantic caching saves tokens on repeated patterns. Cost anomaly detection (burn-rate alerts). Behavior anomaly detection with Z-score flagging.

Sandboxing. Pluggable SandboxBackend protocol; run agents in local git worktrees (default), Docker containers, E2B Firecracker microVMs, or Modal serverless containers (with optional GPU). Plugin authors can register custom backends through the bernstein.sandbox_backends entry-point group. Inspect installed backends with bernstein agents sandbox-backends.

Artifact storage. .sdd/ state can stream to pluggable ArtifactSink backends: local filesystem (default), S3, Google Cloud Storage, Azure Blob, or Cloudflare R2. BufferedSink keeps the WAL crash-safety contract by writing locally with fsync first and mirroring to the remote asynchronously.

Skill packs. Progressive-disclosure skills (OpenAI Agents SDK pattern): only a compact skill index ships in every spawn's system prompt, agents pull full bodies via the load_skill MCP tool on demand. 17 built-in role packs plus third-party bernstein.skill_sources entry-points.

Controls. HMAC-SHA256 audit chain (RFC 2104), policy engine, lethal-trifecta capability gate (refuses spawns whose tool chain combines private data + untrusted input + external comm — Simon Willison's framing, June 2025: "if your AI agent combines all three of these, an attacker can trick it into stealing your data"), PII output gating, WAL-backed crash recovery (experimental, multi-worker safety), OAuth 2.0 with PKCE (RFC 7636) and RFC 8707 resource-indicator binding, per-artefact lineage with customer-key Ed25519 signing (RFC 8037) and regulator export.

Observability. Prometheus /metrics, OTel exporter presets, Grafana dashboards. Per-model cost tracking (bernstein cost) plus a run savings summary on every bernstein run. Terminal TUI and web dashboard. Agent process visibility in ps.

Ecosystem. MCP server mode, A2A protocol support, GitHub App integration, pluggy-based plugin system, multi-repo workspaces, cluster mode for distributed execution, self-evolution via --evolve (experimental).

Full feature matrix: FEATURE_MATRIX.md &middot; Recent features: What's New

regulatory anchors (as of 2026-05-09)

For compliance reviewers asking "which regulation does Bernstein actually map to":

Regulation

Mapping

Bernstein surface

EU AI Act Article 12 (logging)

Automatic record-keeping for high-risk AI systems

bernstein audit export --article-12 --since … --until … → deterministic, retention-pinned bundle with audit slice + governance catalog. See docs/compliance/.

SOC 2 Trust Service Criteria

CC4 / CC7 (audit + monitoring)

bernstein audit pack --soc2 → per-control evidence checklist with sha-256 pointers.

DORA / NIS2

Per-artefact lineage with customer-key Ed25519 signature

bernstein lineage export <run_id> --format jsonld → schema v2 records.

OWASP Agent Security Initiative (ASI06 — memory poisoning, 2026)

Memory provenance audit

bernstein verify --memory-audit walks the lesson-memory chain.

RFC 2104 (HMAC)

Audit chain integrity

.sdd/audit/*.jsonl HMAC-SHA256 with secret outside the audit volume.

RFC 7515 §A.5 (detached JWS) + RFC 8785 (JCS) + RFC 8037 (EdDSA)

Signed agent cards + lineage signatures

src/bernstein/core/security/agent_card_signer.py, src/bernstein/core/security/lineage_kms.py.

RFC 7636 (PKCE) + RFC 8707 (resource indicators)

Web dashboard auth + MCP audience binding

src/bernstein/core/security/oauth_pkce.py, auth.py.

These are mappings, not certifications. Production accreditation (SOC 2 Type II, ISO 27001) is out of scope for a solo-maintained OSS project; the surfaces exist to make a customer's accreditation path shorter.

what's new in v1.9

ACP bridge. bernstein acp serve --stdio exposes Bernstein to any editor that speaks the Agent Communication Protocol (Zed, etc.). No plugin code needed on the editor side.

Autonomous CI repair. bernstein autofix watches open Bernstein PRs and, when CI turns red, spawns a fixer agent automatically. Once green, it pushes the fix and re-requests review.

Credential vault. bernstein connect <provider> writes API keys to the OS keychain; bernstein creds lists and rotates them. Agents inherit scoped credentials without touching environment variables.

Preview tunnels. bernstein preview start boots a sandboxed dev server and prints a public URL. Useful for sharing a running branch with a reviewer without deploying to staging.

Full changelog: docs/whats-new.md

operator commands

Commands that eliminate the glue code most teams end up writing around their runs.

Command

What it does

bernstein pr

Auto-creates a GitHub PR from a completed session; body carries the janitor's gate results and token/USD cost breakdown.

bernstein from-ticket <url>

Imports a Linear / GitHub Issues / Jira ticket as a Bernstein task. Label-based role + scope inference. Supports --dry-run and --run.

bernstein ticket import <url>

Alias / group form of from-ticket for scripting.

bernstein remote

SSH sandbox backend. remote test <host>, remote run <host> <path>, remote forget <host>. ControlMaster socket reuse for fast repeat calls.

bernstein hooks

Lifecycle hooks for pre_task, post_task, pre_merge, post_merge, pre_spawn, post_spawn; shell scripts or pluggy @hookimpls. hooks list, hooks run <event>, hooks check.

bernstein chat serve --platform=telegram|discord|slack

Drive runs from chat with /run, /status, /approve, /reject, /switch, /stop.

bernstein approve-tool / bernstein reject-tool

Interactive mid-run tool-call approval. --latest, --id, --always.

bernstein tunnel start <port> [--provider auto|cloudflared|ngrok|bore|tailscale]

One wrapper around four tunnel providers. Also tunnel list, tunnel stop <name>|--all. ControlMaster-style process reuse.

bernstein daemon install [--user|--system] [--command="..."] [--env KEY=VAL]...

Installs a systemd (Linux) or launchd (macOS) unit for auto-start. Also daemon start/stop/restart/status/uninstall.

bernstein connect <provider> / bernstein creds

Stores and rotates API credentials in the OS keychain. Agents inherit scoped keys per-run.

bernstein autofix

Daemon that monitors open Bernstein PRs; spawns a fixer agent when CI fails and pushes the repair automatically.

bernstein preview start

Starts a sandboxed dev server for the current branch and prints a shareable public tunnel URL.

bernstein agents-md

Generates a canonical AAIF AGENTS.md for the repo and rewrites it into each CLI's native shape. generate (preview), write (single file), sync (canonical + Cursor .cursor/rules/*.mdc + Claude CLAUDE.md + Aider CONVENTIONS.md + Goose .goosehints), verify (CI gate), diff (shows drift between canonical IR and on-disk files).

bernstein scaffold "<prompt>"

Bootstraps a project skeleton from a single goal prompt. --template auto|python-cli|..., --output <dir>, --force.

bernstein wiki build

Renders WIKI.md for the current repo from the AST symbol graph. Local, no LLM call, no cloud round-trip.

bernstein identity show / decode / verify / disable

Operator-side helpers for the install-rev fingerprint embedded in shared yaml/trace/role-prompt artefacts. No network egress; discovery uses public gh search code.

bernstein security role-adapter-policy

Inspects and edits the per-role adapter allow-list (deny-list enforcement at spawn time).

retrieval & caching: what's actually under the hood

Bernstein deliberately uses no neural embeddings, no vector databases, and no external embedding APIs. There are two retrieval/caching layers, both keyword/lexical:

  • Codebase RAG (core/knowledge/rag.py); SQLite FTS5 with BM25 ranking and AST-aware chunking for Python files. Built incrementally on file mtime; used to enrich agent task context within token budgets.

  • Semantic cache (core/knowledge/semantic_cache.py); despite the name, fuzzy matching is done with TF (term-frequency) cosine similarity over word counts, not learned embeddings. It deduplicates near-identical LLM planning and agent-output requests so we don't re-spawn agents for the same goal.

If you need real semantic retrieval (vector DB, neural embeddings), wire it yourself via the retrieval role/skill in templates/; nothing in core performs vector search.

detailed comparison

Feature

Bernstein

CrewAI

AutoGen autogen

LangGraph

Orchestrator

Deterministic code

LLM-driven (+ code Flows)

LLM-driven

Graph + LLM

Works with

Any CLI agent (43 adapters)

Python SDK classes

Python agents

LangChain nodes

Git isolation

Worktrees per agent

No

No

No

Pluggable sandboxes

Worktree, Docker, E2B, Modal

No

No

No

Verification

Janitor + quality gates

Guardrails + Pydantic output

Termination conditions

Conditional edges

Cost tracking

Built-in

usage_metrics

RequestUsage

Via LangSmith

State model

File-based (.sdd/)

In-memory + SQLite checkpoint

In-memory

Checkpointer

Remote artifact sinks

S3, GCS, Azure Blob, R2

No

No

No

Self-evolution

Built-in (experimental)

No

No

No

Declarative plans (YAML)

Yes

Yes (agents.yaml, tasks.yaml)

No

Partial (langgraph.json)

Model routing per task

Yes

Per-agent LLM

Per-agent model_client

Per-node (manual)

MCP support

Yes (client + server)

Yes

Yes (client + workbench)

Yes (client + server)

Agent-to-agent chat

Bulletin board

Yes (Crew process)

Yes (group chat)

Yes (supervisor, swarm)

Web UI

TUI + web dashboard

CrewAI AMP

AutoGen Studio

LangGraph Studio + LangSmith

Cloud hosted option

Yes (Cloudflare)

Yes (CrewAI AMP)

No

Yes (LangGraph Cloud)

Built-in RAG/retrieval

Yes (codebase FTS5 + BM25)

crewai_tools

autogen_ext retrievers

Via LangChain

Last verified: 2026-04-19. See full comparison pages for detailed feature matrices.

The table above compares Bernstein against LLM-orchestration frameworks (they orchestrate LLM calls). The table below covers the closer category: other tools that orchestrate CLI coding agents:

Feature

Bernstein

awslabs/cli-agent-orchestrator

ComposioHQ/agent-orchestrator

emdash

umputun/ralphex

Shape

Python CLI + library + MCP server

Python CLI + tmux sessions + web UI

TypeScript CLI + local dashboard

Electron desktop app

Go CLI

Primary language

Python

Python

TypeScript

TypeScript

Go

Install

pipx install bernstein

uv tool install cli-agent-orchestrator

npm install -g @aoagents/ao

.dmg / .msi / .AppImage

go install / single binary

Agent adapters

43

5 (Kiro, Claude Code, Codex, Gemini, Kimi)

3 (Claude Code, Codex, Aider)

24

1 (Claude Code only)

Parallel multi-agent execution

Yes

Yes (tmux session per agent)

Yes

Yes

No (single sequential session)

Git worktree per agent

Yes

No (planned, #100)

Yes

Yes

Optional --worktree flag

MCP server mode (exposes self as MCP)

Yes (stdio + HTTP/SSE)

Yes (inter-agent comms)

No

No

No

Coordinator

Deterministic Python scheduler

Hierarchical LLM supervisor

LLM-driven

Not documented

Linear plan executor

HMAC-chained audit replay

Yes

No

No

No

No

Cross-model verifier / quality gates

Yes (multi-stage)

No

No

No

Multi-phase review (Claude only)

Autonomous CI-fix / PR flow

Yes (bernstein autofix)

No

Yes

No

No

Visual dashboard

TUI + web

Web UI + tmux

Web

Desktop app

Web (--serve)

Notification sinks

Telegram/Slack/Discord/Email/Webhook/Shell

—

No

No

Telegram / Email / Slack / Webhook

Backing

Solo OSS

AWS Labs

Funded (Composio.dev)

YC W26

Solo OSS

License

Apache 2.0

Apache 2.0

MIT

Apache 2.0

MIT

Bernstein's wedge in this category: Python-native, MCP-server-first, widest adapter coverage, true multi-agent parallelism, deterministic scheduler with no LLM in the coordination loop. If you want AWS-aligned tmux-session isolation with a hierarchical LLM supervisor, AWS Labs' cao is a closer fit; if your stack is TypeScript and you want a product with a dashboard, Composio's @aoagents/ao is a better fit; if you want a polished desktop ADE, emdash is; if you only use Claude Code and want a single Go binary that walks a plan top-to-bottom, ralphex is. If you want a primitive that imports into Python, exposes itself over MCP to any client, runs many agents in parallel, and covers the full agent breadth (including Qwen, Goose, Ollama, OpenAI Agents SDK, Cloudflare Agents, and more), Bernstein.

what people use it for

These are real workflow patterns from Bernstein's own docs, examples, and project surface, not invented customer quotes.

  • Parallel test generation. Fan out across untested modules with BERNSTEIN_MAX_AGENTS=5 bernstein -g "Generate unit tests for untested modules in src/".

  • CI failure repair. Watch open PRs and dispatch scoped fixers with bernstein autofix start --repo your-org/your-repo --foreground.

  • PR review follow-up. Turn review comments into tracked fix tasks with bernstein review-responder start --repo your-org/your-repo --foreground.

  • Codebase modernization. Run wide refactors like BERNSTEIN_MAX_AGENTS=8 bernstein -g "Migrate callback-based modules in src/ to async/await and update tests".

  • Ticket-to-run workflows. Import GitHub, Jira, or Linear work directly with bernstein from-ticket https://github.com/your-org/your-repo/issues/123 --run.

  • API-change safety checks. Catch downstream breakage before merge with bernstein dep-impact --base main.

See Who Uses Bernstein for the longer version with command examples and notes on when each workflow fits.

monitoring

bernstein live       # TUI dashboard
bernstein dashboard  # web dashboard
bernstein status     # task summary
bernstein ps         # running agents
bernstein cost       # spend by model/task
bernstein doctor     # pre-flight checks
bernstein recap      # post-run summary
bernstein trace <ID> # agent decision trace
bernstein run-changelog --hours 48  # changelog from agent-produced diffs
bernstein explain <cmd>  # detailed help with examples
bernstein dry-run    # preview tasks without executing
bernstein dep-impact # API breakage + downstream caller impact
bernstein aliases    # show command shortcuts
bernstein config-path    # show config file locations
bernstein init-wizard    # interactive project setup
bernstein debug-bundle   # collect logs, config, and state for bug reports
bernstein skills list    # discoverable skill packs (progressive disclosure)
bernstein skills show <name>  # print a skill body with its references
bernstein fingerprint build --corpus-dir ~/oss-corpus  # build local similarity index
bernstein fingerprint check src/foo.py                 # check generated code against the index

install

Method

Command

One-liner (macOS / Linux)

curl -fsSL https://bernstein.run/install.sh | sh

One-liner (Windows)

irm https://bernstein.run/install.ps1 | iex

pip

pip install bernstein

pipx

pipx install bernstein

uv

uv tool install bernstein

Homebrew

brew tap chernistry/tap && brew install bernstein

Fedora / RHEL

sudo dnf copr enable alexchernysh/bernstein && sudo dnf install bernstein

npm (wrapper)

npx bernstein-orchestrator

Docker (GHCR)

docker run --rm -v "$PWD:/work" -w /work -e ANTHROPIC_API_KEY ghcr.io/sipyourdrink-ltd/bernstein:latest run -g "fix tests/test_foo.py"

The one-liner scripts check for Python 3.12+, bootstrap pipx when it's missing, fix PATH for the current session, and install (or upgrade) bernstein. They handle brew-managed macOS environments and the Windows py -3 launcher fallback. Script sources: install.sh · install.ps1.

optional extras

Provider SDKs are optional so the base install stays lean. Pick what you need:

Extra

Enables

bernstein[openai]

OpenAI Agents SDK v2 adapter (openai_agents)

bernstein[docker]

Docker sandbox backend

bernstein[e2b]

E2B microVM sandbox backend (needs E2B_API_KEY)

bernstein[modal]

Modal sandbox backend, optional GPU (needs MODAL_TOKEN_ID / MODAL_TOKEN_SECRET)

bernstein[s3]

S3 artifact sink (via boto3)

bernstein[gcs]

Google Cloud Storage artifact sink

bernstein[azure]

Azure Blob artifact sink

bernstein[r2]

Cloudflare R2 artifact sink (S3-compatible boto3)

bernstein[grpc]

gRPC bridge

bernstein[k8s]

Kubernetes integrations

Combine extras with brackets, e.g. pip install 'bernstein[openai,docker,s3]'.

Editor extensions: VS Marketplace &middot; Open VSX

"powered by bernstein" badge (optional)

If your project ships diffs that bernstein helped land, you can advertise it:

[![signed by bernstein](https://img.shields.io/badge/signed_by-bernstein-FBBF24?logo=githubactions&logoColor=white&style=flat-square)](https://bernstein.run/?utm_source=badge&utm_medium=readme&utm_campaign=powered-by)

bernstein init --add-badge injects it into your README under the existing badge stack. Variants: signed, audited-by, orchestrated-by, crew-managed-by — pass via --badge-variant. Picky maintainers can keep their READMEs untouched: the flag is opt-in.

contributing

PRs welcome. See CONTRIBUTING.md for setup and code style.

support

If Bernstein saves you time: GitHub Sponsors

Contact: forte@bernstein.run

Curated lists, newsletters, and peer projects that picked up Bernstein:

cite

If Bernstein helps your research or industry work, please cite it. Machine-readable metadata lives in CITATION.cff (CFF 1.2.0); GitHub renders the "Cite this repository" button automatically. A Zenodo DOI will be minted on the next release once Zenodo's GitHub integration is enabled — see CITATION.cff for the current canonical citation.

license

Apache License 2.0


Made with love by Alex Chernysh &middot; GitHub &middot; X &middot; bernstein.run

translations

Español &middot; 中文 &middot; العربية &middot; Português &middot; Bahasa Indonesia &middot; Français &middot; 日本語 &middot; Русский &middot; Deutsch &middot; עברית &middot; יידיש

  1. Workflow YAML shipped in PR #1117 (merged 2026-05-08). Plans are authored as YAML and validated by bernstein workflow validate.

    ↩
  2. AutoGen is in maintenance mode; successor is Microsoft Agent Framework 1.0.

    ↩

Available Tools

12 tools
bernstein_approveA

Sign off a finished result only when the task is in pending_approval; every other status is refused. This completes the task on the Bernstein server. Host effects: makes network requests.

ParametersJSON Schema
NameRequiredDescriptionDefault
noteNo
task_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals that the tool refuses non-pending_approval statuses, completes the task on the server, and makes network requests. This gives a clear state-change and safety profile, though it does not mention reversibility or error handling specifics, which might be covered by the output schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences plus a brief 'Host effects' note. It is front-loaded with the core condition and action, contains no filler, and every clause earns its place. This is an exemplar of concise, structured writing.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the essential behavior, the precondition, and the side effects. Since an output schema exists, return values are likely documented elsewhere. Missing details include the role of the note parameter and the relationship to bernstein_complete, but overall the description is sufficiently complete for a simple approval action.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0% description coverage, so the tool description must compensate. It only vaguely refers to 'task' without explaining task_id or note. The note parameter's purpose (e.g., an approval message) is entirely unexplained, and the condition 'pending_approval' does not clarify parameter semantics enough. This is a clear gap.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's specific function: 'Sign off a finished result' only when the task is in pending_approval, and that it 'completes the task.' The explicit status condition ('every other status is refused') adds scope and distinguishes it from sibling tools like bernstein_complete, which may have different completion semantics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear guidance on when to use the tool: only when the task is in pending_approval. It explicitly states that any other status is refused, signaling when not to use it. However, it does not name alternative tools (e.g., bernstein_complete) or explain the exact workflow relationship, leaving some ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bernstein_cancelA

Cancel one task and its subtask tree; the orchestrator keeps running. Cancellable statuses: open, claimed, in_progress, blocked, waiting_for_subtasks, planned. An already-terminal task is reported, not treated as an error. This changes task state on the Bernstein server. Host effects: makes network requests.

ParametersJSON Schema
NameRequiredDescriptionDefault
reasonNo
task_idYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.3/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must explicitly communicate side effects. It states 'changes task state on the Bernstein server,' 'makes network requests,' and describes behavior for terminal tasks. This fully discloses the tool's behavioral footprint.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded. Each sentence contributes unique information: purpose, eligible statuses, terminal-task handling, and side effects. No redundancy or filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists, return values need not be explained. The description covers purpose, status constraints, error behavior, and side effects. Minor gaps: no detail on invalid task_id handling or how the reason parameter influences cancellation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 0% parameter description coverage and the description does not compensate. It never mentions the 'reason' parameter, and task_id is only implicitly tied to 'one task.' No guidance is given on how to obtain valid task_id values beyond the schema pattern.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb+resource: 'Cancel one task and its subtask tree.' It clearly distinguishes from sibling tools like bernstein_shutdown_orchestrator by stating 'orchestrator keeps running.' Listing the cancellable statuses further clarifies scope.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides clear context for when to use: canceling a specific task tree while keeping the orchestrator alive. The enumerating of cancellable statuses and non-error on terminal tasks gives implied usage boundaries, but it does not explicitly name alternatives or exclusion scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bernstein_claimA

Claim the next eligible task for an agent. This changes task assignment state on the Bernstein server. Host effects: makes network requests.

ParametersJSON Schema
NameRequiredDescriptionDefault
roleNo
projectNo
capabilityNo
claimer_idYes
max_attemptsNo
completed_idsNo
claimer_card_fingerprintNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It explicitly states that the tool changes task assignment state and makes network requests ('Host effects: makes network requests'). This goes beyond a simple 'claim' verb and informs the agent of side effects, though it does not cover authorization or reversibility.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences deliver the action, the state change, and the host effect without redundancy. Every sentence adds unique value, and the description is front-loaded with the core purpose.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (7 parameters, state-changing, network requests) and lack of annotations, the description is incomplete. It does not address parameter semantics, eligibility criteria, or usage guidance, and though an output schema exists, the missing parameter context prevents a fully capable agent from using the tool effectively.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema description coverage is 0%, and the tool description does not explain any of the seven parameters (role, project, capability, claimer_id, max_attempts, completed_ids, claimer_card_fingerprint). The agent is left to infer semantics purely from parameter names, which is inadequate for correct invocation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Claim') and resource ('next eligible task for an agent'), clearly distinguishing it from siblings like bernstein_complete or bernstein_cancel. It also clarifies the state-changing nature by stating 'This changes task assignment state on the Bernstein server.'

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use the tool: when an agent needs to claim the next eligible task. It does not explicitly mention alternatives or when-not-to-use, but the purpose is self-evident and distinct from sibling tools. No exclusions are necessary for this action.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bernstein_completeA

Report a task result only from a worker-held state: open, claimed, or in_progress; every other status is refused. This completes the task on the Bernstein server. Host effects: makes network requests.

ParametersJSON Schema
NameRequiredDescriptionDefault
task_idYes
result_summaryYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden. It discloses the refusal behavior for invalid states and the host effect of making network requests. However, it does not mention reversibility, authentication, or other potential side effects beyond completing the task.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise, with three short sentences. It front-loads the purpose, then adds the state constraint and a host effect note. Every sentence is informative and there is no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 2-parameter mutation tool with an output schema, the description covers the core purpose and a key constraint, but parameter semantics are under-specified. An agent may need to infer the expected content of result_summary, making the description adequate but not fully complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It only mentions 'task result' generically, leaving task_id and result_summary to name inference. No additional semantic guidance is given for their format or content beyond the schema's constraints.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: 'Report a task result' and 'completes the task on the Bernstein server.' It also specifies the allowed worker-held states, which distinguishes it from sibling tools like bernstein_claim, bernstein_cancel, and bernstein_approve.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit context for when to use the tool by listing valid task statuses ('open, claimed, or in_progress') and stating that other statuses are refused. It does not name alternative tools, but the state constraint effectively communicates the appropriate usage scenario.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bernstein_post_artifactC

Post a versioned artifact to a task on the Bernstein server. Host effects: makes network requests.

ParametersJSON Schema
NameRequiredDescriptionDefault
keyYes
urlNo
bodyNo
rowsNo
toolNo
posterYes
targetNo
columnsNo
task_idYes
link_kindNo
sarif_resultNo
tool_versionNo
artifact_typeYes
invocation_argv_hashNo
pinned_ruleset_or_feed_digestNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

C2.6/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It adds only 'Host effects: makes network requests,' which is marginal because 'post' already implies a network request. It does not disclose versioning behavior, task state changes, idempotency, or failure semantics.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two short sentences with no filler, and the core action is front-loaded. However, the 'Host effects' sentence is boilerplate and adds little information, preventing a top score.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a complex tool with 15 parameters, conditional requirements, and four artifact_type variants, yet the description provides almost none of that context. Even though an output schema exists and return-value documentation is not required, the missing parameter semantics and usage guidance leave the description incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate for 15 parameters, but it explains none of them. It does not clarify artifact_type variants, required conditional fields, link_kind values, or the meaning of key, poster, target, or invocation_argv_hash.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action and resource: 'Post a versioned artifact to a task on the Bernstein server.' This is a specific verb plus resource and is distinguishable from siblings like bernstein_post_message by the artifact focus, but it does not explicitly differentiate itself from any sibling or mention the artifact_type variants.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no guidance on when to use this tool versus bernstein_post_message or other artifact-related operations. The phrase 'versioned artifact' implies a niche, but no context, exclusions, or alternative tool referrals are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bernstein_post_messageB

Post a progress message to a task mailbox on the Bernstein server. Host effects: makes network requests.

ParametersJSON Schema
NameRequiredDescriptionDefault
bodyYes
kindNo
senderYes
task_idYes
sender_card_fingerprintNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

B3.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description discloses the host effect 'makes network requests', which is a useful side-effect warning in the absence of annotations. However, it does not elaborate on other behavioral traits such as idempotency, required task state, or error outcomes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences with the primary action front-loaded. Every phrase earns its place, making it extremely concise and easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The purpose is clear, but the description lacks usage guidelines, parameter semantics, and differentiation from sibling tools. Given the low schema coverage and absence of annotations, more context is needed for an agent to invoke this tool appropriately.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters1/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, and the description does not explain any of the five parameters. While the schema includes constraints and an enum for 'kind', the description itself adds no meaning beyond the field names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Post') and identifies the resource ('progress message to a task mailbox on the Bernstein server'), clearly distinguishing it from sibling tools like bernstein_post_artifact. It is concise and unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. The description only states the action without indicating prerequisites, exclusions, or comparisons to sibling tools such as bernstein_post_artifact or bernstein_complete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bernstein_runA

Start an orchestration run. A run does real work and takes minutes to hours; the call returns once the run is queued, not when it finishes. Do not re-issue it while waiting, that starts a second run. Follow the run with bernstein_run_status, passing either the returned task_id or the returned run_id, after waiting the returned poll_after_ms. Pass parent_task_id to create the run as a subtask of an existing task. The queued orchestration writes project state and starts agent work. Host effects: writes files; spawns agent processes; makes network requests.

ParametersJSON Schema
NameRequiredDescriptionDefault
goalYes
roleNo
scopeNo
priorityNo
complexityNo
parent_task_idNo
estimated_minutesNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
_meterYes
resultYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and delivers excellently. It discloses async (returns when queued, not finished), side effects (writes files, spawns agent processes, makes network requests), and the returned poll_after_ms. This is far beyond minimal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every sentence earns its place: purpose, async warning, follow-up, subtask usage, and host effects. It is well-structured and not bloated.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex orchestration tool with side effects, the description covers all critical aspects: queuing model, duplicate-run risk, polling strategy, subtask support, and host-level consequences. Output schema exists, so return values are covered structurally.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Despite 7 parameters and 0% schema description coverage, the description only explains parent_task_id ('pass to create the run as a subtask'). The required 'goal' and others like role, scope, priority, complexity, estimated_minutes are left unexplained, relying solely on naming and enums.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Start an orchestration run', a specific verb and resource. It clearly distinguishes from siblings like bernstein_run_status (which monitors) and bernstein_cancel by emphasizing the queuing behavior and follow-up steps.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use and alternatives: tells users not to re-issue while waiting (starts a second run), instructs to follow with bernstein_run_status after waiting poll_after_ms, and mentions parent_task_id for subtasks. This is exemplary usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bernstein_run_statusA

Poll a verifiable handle for a run started with bernstein_run. Accepts either identifier that call returned: the task_id or the run_id. Reads the local run journal and audit evidence without changing them. Host effects: reads files.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYesThe run to project. Either the task_id or the run_id returned by bernstein_run. Resolved journal run id first, then the task id slugified into a journal run id, so both forms reach one journal.
workdirNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
_meterYes
resultYes

TDQS

A3.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and succeeds by stating it reads the local run journal and audit evidence without changing them, plus host effects: reads files. This gives a clear safety profile of a read-only operation. It does not add details on errors or return format, but the output schema likely covers that, so the provided behavioral transparency is strong.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and well-structured: it opens with the primary purpose, then describes accepted identifiers, then discloses the read-only behavior and host effects. Every sentence earns its place without repetition or fluff. It is appropriately sized for the tool's simplicity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has an output schema, so return values are covered elsewhere. The description covers purpose, parameter identity, and behavioral side effects, which is mostly sufficient. However, it leaves workdir unexplained and does not mention the sibling bernstein_status, so an agent might struggle to choose correctly between them. This is a gap given the tool's moderate complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 50%: run_id is well described in the schema (accepts task_id or run_id), but workdir has no schema description and the tool description does not explain it either. The description merely restates run_id semantics already in the schema, adding no new meaning and leaving workdir's purpose ambiguous. With low schema coverage, the description should compensate but does not.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool polls a verifiable handle for a run started with bernstein_run, using a specific verb and resource. It accepts either task_id or run_id, which defines its purpose well. However, it does not differentiate from the sibling tool bernstein_status, leaving some ambiguity about their distinct roles.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use the tool: after calling bernstein_run, with either returned identifier. It provides some guidance on parameter inputs but does not mention bernstein_status as an alternative or specify when to choose this tool over others. The context is clear but lacks explicit exclusions or alternative comparisons.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bernstein_shutdown_orchestratorA

Shut down the ENTIRE Bernstein orchestrator for this project, including every run and worker; use bernstein_cancel to stop one task while the orchestrator keeps running. Writes the local SHUTDOWN signal file. Host effects: writes files.

ParametersJSON Schema
NameRequiredDescriptionDefault
workdirNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the transparency burden. It discloses a clear side effect: 'Writes the local SHUTDOWN signal file. Host effects: writes files.' It also communicates the destructive scope ('ENTIRE', 'including every run and worker'). It does not detail reversibility, permissions, or whether shutdown is graceful, but the disclosed effects go well beyond a vague mutation claim.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is three sentences, front-loaded with the primary action and scope. Every sentence adds value: the main behavior, the alternative tool for narrower cancellation, and the local file side effect. No filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description covers the core purpose, scope, side effect, and alternative, but leaves workdir unexplained and does not describe the return value or post-shutdown state. Given the destructive nature and lack of annotations, a more complete description—especially about the parameter and consequences—would be expected. The output schema may compensate for return details, but the parameter omission remains a gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has one optional parameter, workdir, with no property description (coverage 0%). The description never mentions workdir or how it affects which project is shut down. The agent is left to infer that workdir selects the project context, which is not explicitly clarified.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description starts with a specific verb and resource: 'Shut down the ENTIRE Bernstein orchestrator for this project, including every run and worker.' It clearly distinguishes itself from bernstein_cancel, which stops a single task while the orchestrator continues, making the tool's scope and purpose unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit guidance is provided: 'use bernstein_cancel to stop one task while the orchestrator keeps running.' This tells the agent exactly when to choose this tool versus the alternative, and it also implies when a full shutdown (rather than a cancel) is appropriate.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bernstein_statusA

Liveness, task counts, and cost in one read. Pass status to include the matching tasks; pass detail=true for full per-role and per-task rows. Retrieves data from the Bernstein server without changing it. Host effects: makes network requests.

ParametersJSON Schema
NameRequiredDescriptionDefault
detailNo
statusNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
_meterYes
resultYes

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description must carry the transparency burden. It explicitly states 'Retrieves data from the Bernstein server without changing it' and 'Host effects: makes network requests,' which discloses read-only behavior and side effects. This adds useful context beyond the schema, though it does not cover details like permissions or error cases.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the core purpose ('Liveness, task counts, and cost in one read') and followed by parameter guidance and behavioral notes. Every sentence adds value, and there is no wasted text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has an output schema, return values are already covered. The description sufficiently covers the purpose, parameters, and read-only nature, making it complete for a status tool. Minor gaps exist in not fully explaining what 'liveness' entails, but overall it is well-rounded.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, but the description compensates by explaining both parameters: 'Pass status to include the matching tasks' and 'pass detail=true for full per-role and per-task rows.' This adds meaningful semantics beyond the bare enum and boolean in the schema, clarifying their purpose and defaults.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool as a read-only status endpoint for liveness, task counts, and cost. It uses specific nouns and implies a resource, but it does not explicitly differentiate from the sibling tool bernstein_run_status, so it meets 'clear' but lacks direct sibling differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides usage guidance for parameters (pass status to filter tasks, detail=true for full rows) but does not state when to choose this tool over alternatives like bernstein_run_status. There is no explicit 'use this for server-level status' or mention of exclusions, so it is implied rather than explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

bernstein_task_capsuleA

Read a task capsule together with its local journal and audit evidence. With verify=true, verification may create the install audit key if it is absent. Host effects: reads files; writes files.

ParametersJSON Schema
NameRequiredDescriptionDefault
verifyNo
task_idYes
workdirNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of disclosure. It explicitly mentions conditional side effects (verify=true may create install audit key) and host effects ('reads files; writes files'), which is unusually transparent for a tool named 'read'.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the primary purpose. The side-effect disclosure is compact and informative; no filler or redundant content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be described. The description covers the core action, conditional mutation, and host effects, but lacks usage context around workdir and when to prefer sibling status tools.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema property descriptions are absent (0% coverage), so the description must compensate. It explains the verify parameter's conditional side effect, but task_id and workdir receive no semantic explanation beyond their names and schema types.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Description states a specific action ('Read a task capsule...') with a clear resource (task capsule, local journal, audit evidence). It distinguishes from siblings like bernstein_run and bernstein_status by focusing on reading capsule contents rather than executing or checking status.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No explicit guidance on when to use this tool versus alternatives such as bernstein_status or bernstein_run_status. The description implies a read/inspection use case but does not state exclusions or preferred alternatives.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

load_skillA

List available skills when name is omitted, or load a named skill body, reference, or script file contents. Returns file contents as text; executes nothing. Host effects: reads files.

ParametersJSON Schema
NameRequiredDescriptionDefault
nameNoSkill to load. Omit to return the compact skill index.
scriptNo
referenceNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations available, the description carries full burden. It discloses that it returns file contents as text, executes nothing, and reads files. This clearly signals a read-only, safe operation. It could further mention error behavior (e.g., not found), but the provided info is solid.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, tightly packed with meaningful content: behavior, return type, safety, and host effect. Every word earns its place; no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the simple tool structure, an output schema exists, and the description covers core behavior and safety, it is mostly complete. It could be improved by explaining how missing files are handled, but that is a minor gap. Overall, it supplies enough for an agent to use the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is only 33% (only name has a description), so the description must compensate. It adds meaning by mentioning 'skill body, reference, or script file contents', mapping to the three possible loads. However, it does not explain the exact format or relationship of script/reference parameters beyond what dependencies imply.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states two distinct behaviors: listing skills when name is omitted and loading skill body/reference/script contents when name is provided. The verb 'list' and 'load' are specific and the resource is well-defined, distinguishing it from the bernstein_* siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'when name is omitted' versus when a name is provided, giving clear context for both usage modes. It also notes 'executes nothing', implying it is for inspection, not execution, but it does not name alternative tools for execution, so it misses explicit when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev3.18.1
    • Changedbernstein_post_artifact8 fields changed
      • changedInput schema / allOf
        Previous value: -[
        -  {
        -    "if": {
        -      "properties": {
        -        "artifact_type": {
        -          "const": "report"
        -        }
        -      },
        -      "required": [
        -        "artifact_type"
        -      ]
        -    },
        -    "then": {
        -      "required": [
        -        "body"
        -      ]
        -    }
        -  },
        -  {
        -    "if": {
        -      "properties": {
        -        "artifact_type": {
        -          "const": "table"
        -        }
        -      },
        -      "required": [
        -        "artifact_type"
        -      ]
        -    },
        -    "then": {
        -      "required": [
        -        "columns",
        -        "rows"
        -      ]
        -    }
        -  },
        -  {
        -    "if": {
        -      "properties": {
        -        "artifact_type": {
        -          "const": "link"
        -        }
        -      },
        -      "required": [
        -        "artifact_type"
        -      ]
        -    },
        -    "then": {
        -      "required": [
        -        "url",
        -        "link_kind"
        -      ]
        -    }
        -  }
        -]New value: +[
        +  {
        +    "if": {
        +      "properties": {
        +        "artifact_type": {
        +          "const": "report"
        +        }
        +      },
        +      "required": [
        +        "artifact_type"
        +      ]
        +    },
        +    "then": {
        +      "required": [
        +        "body"
        +      ]
        +    }
        +  },
        +  {
        +    "if": {
        +      "properties": {
        +        "artifact_type": {
        +          "const": "table"
        +        }
        +      },
        +      "required": [
        +        "artifact_type"
        +      ]
        +    },
        +    "then": {
        +      "required": [
        +        "columns",
        +        "rows"
        +      ]
        +    }
        +  },
        +  {
        +    "if": {
        +      "properties": {
        +        "artifact_type": {
        +          "const": "link"
        +        }
        +      },
        +      "required": [
        +        "artifact_type"
        +      ]
        +    },
        +    "then": {
        +      "required": [
        +        "url",
        +        "link_kind"
        +      ]
        +    }
        +  },
        +  {
        +    "if": {
        +      "properties": {
        +        "artifact_type": {
        +          "const": "finding"
        +        }
        +      },
        +      "required": [
        +        "artifact_type"
        +      ]
        +    },
        +    "then": {
        +      "required": [
        +        "sarif_result"
        +      ]
        +    }
        +  }
        +]
      • changedInput schema / properties / artifact_type / enum
        Previous value: -[
        -  "report",
        -  "table",
        -  "link"
        -]New value: +[
        +  "report",
        +  "table",
        +  "link",
        +  "finding"
        +]
      • addedInput schema / properties / invocation_argv_hash
        Added value: +{
        +  "maxLength": 256,
        +  "type": "string"
        +}
      • addedInput schema / properties / pinned_ruleset_or_feed_digest
        Added value: +{
        +  "maxLength": 256,
        +  "type": "string"
        +}
      • addedInput schema / properties / sarif_result
        Added value: +{
        +  "type": "object"
        +}
      • addedInput schema / properties / target
        Added value: +{
        +  "maxLength": 4096,
        +  "type": "string"
        +}
      • addedInput schema / properties / tool
        Added value: +{
        +  "maxLength": 256,
        +  "minLength": 1,
        +  "type": "string"
        +}
      • addedInput schema / properties / tool_version
        Added value: +{
        +  "maxLength": 128,
        +  "type": "string"
        +}
  2. 12 tool updatesv3.15.0
    • Changedbernstein_approve1 field changed
      • addedInput schema / description
        Added value: +"Sign off a finished result only when the task is in pending_approval; every other status is refused. This completes the task on the Bernstein server. Host effects: makes network requests."
    • Changedbernstein_cancel1 field changed
      • changedInput schema / description
        Previous value: -"Cancel one task and its subtask tree; the orchestrator keeps running. Cancellable statuses: open, claimed, in_progress, blocked, waiting_for_subtasks, planned. An already-terminal task is reported, not treated as an error."New value: +"Cancel one task and its subtask tree; the orchestrator keeps running. Cancellable statuses: open, claimed, in_progress, blocked, waiting_for_subtasks, planned. An already-terminal task is reported, not treated as an error. This changes task state on the Bernstein server. Host effects: makes network requests."
    • Changedbernstein_claim1 field changed
      • addedInput schema / description
        Added value: +"Claim the next eligible task for an agent. This changes task assignment state on the Bernstein server. Host effects: makes network requests."
    • Changedbernstein_complete1 field changed
      • addedInput schema / description
        Added value: +"Report a task result only from a worker-held state: open, claimed, or in_progress; every other status is refused. This completes the task on the Bernstein server. Host effects: makes network requests."
    • Changedbernstein_post_artifact2 fields changed
      • addedInput schema / description
        Added value: +"Post a versioned artifact to a task on the Bernstein server. Host effects: makes network requests."
      • addedInput schema / timeoutSeconds
        Added value: +30
    • Changedbernstein_post_message1 field changed
      • addedInput schema / description
        Added value: +"Post a progress message to a task mailbox on the Bernstein server. Host effects: makes network requests."
    • Changedbernstein_run1 field changed
      • changedInput schema / description
        Previous value: -"Start an orchestration run. A run does real work and takes minutes to hours; the call returns once the run is queued, not when it finishes. Do not re-issue it while waiting, that starts a second run. Follow the run with bernstein_run_status, passing either the returned task_id or the returned run_id, after waiting the returned poll_after_ms. Pass parent_task_id to create the run as a subtask of an existing task."New value: +"Start an orchestration run. A run does real work and takes minutes to hours; the call returns once the run is queued, not when it finishes. Do not re-issue it while waiting, that starts a second run. Follow the run with bernstein_run_status, passing either the returned task_id or the returned run_id, after waiting the returned poll_after_ms. Pass parent_task_id to create the run as a subtask of an existing task. The queued orchestration writes project state and starts agent work. Host effects: writes files; spawns agent processes; makes network requests."
    • Changedbernstein_run_status1 field changed
      • changedInput schema / description
        Previous value: -"Poll a verifiable handle for a run started with bernstein_run. Accepts either identifier that call returned: the task_id or the run_id."New value: +"Poll a verifiable handle for a run started with bernstein_run. Accepts either identifier that call returned: the task_id or the run_id. Reads the local run journal and audit evidence without changing them. Host effects: reads files."
    • Changedbernstein_shutdown_orchestrator1 field changed
      • addedInput schema / description
        Added value: +"Shut down the ENTIRE Bernstein orchestrator for this project, including every run and worker; use bernstein_cancel to stop one task while the orchestrator keeps running. Writes the local SHUTDOWN signal file. Host effects: writes files."
    • Changedbernstein_status1 field changed
      • changedInput schema / description
        Previous value: -"Liveness, task counts, and cost in one read. Pass status to include the matching tasks; pass detail=true for full per-role and per-task rows."New value: +"Liveness, task counts, and cost in one read. Pass status to include the matching tasks; pass detail=true for full per-role and per-task rows. Retrieves data from the Bernstein server without changing it. Host effects: makes network requests."
    • Changedbernstein_task_capsule1 field changed
      • addedInput schema / description
        Added value: +"Read a task capsule together with its local journal and audit evidence. With verify=true, verification may create the install audit key if it is absent. Host effects: reads files; writes files."
    • Changedload_skill2 fields changed
      • changedInput schema / description
        Previous value: -"List available skills when name is omitted, or load a named skill body, reference, or script."New value: +"List available skills when name is omitted, or load a named skill body, reference, or script file contents. Returns file contents as text; executes nothing. Host effects: reads files."
      • addedInput schema / timeoutSeconds
        Added value: +30
  3. 19 tool updatesv3.11.0
    • Changedbernstein_approve10 fields changed
      • addedInput schema / $schema
        Added value: +"http://json-schema.org/draft-07/schema#"
      • addedInput schema / additionalProperties
        Added value: +false
      • removedInput schema / properties / note / default
        Removed value: -"Approved via MCP"
      • addedInput schema / properties / note / maxLength
        Added value: +8192
      • removedInput schema / properties / note / title
        Removed value: -"Note"
      • addedInput schema / properties / task_id / maxLength
        Added value: +256
      • addedInput schema / properties / task_id / minLength
        Added value: +1
      • addedInput schema / properties / task_id / pattern
        Added value: +"^[A-Za-z0-9_.:-]+$"
      • removedInput schema / properties / task_id / title
        Removed value: -"Task Id"
      • changedInput schema / title
        Previous value: -"bernstein_approveArguments"New value: +"bernstein_approve"
    • Addedbernstein_cancel
    • Changedbernstein_claim39 fields changed
      • addedInput schema / $schema
        Added value: +"http://json-schema.org/draft-07/schema#"
      • addedInput schema / additionalProperties
        Added value: +false
      • removedInput schema / properties / capability / anyOf
        Removed value: -[
        -  {
        -    "type": "string"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]
      • removedInput schema / properties / capability / default
        Removed value: -null
      • addedInput schema / properties / capability / maxLength
        Added value: +256
      • removedInput schema / properties / capability / title
        Removed value: -"Capability"
      • addedInput schema / properties / capability / type
        Added value: +[
        +  "string",
        +  "null"
        +]
      • removedInput schema / properties / claimer_card_fingerprint / anyOf
        Removed value: -[
        -  {
        -    "type": "string"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]
      • removedInput schema / properties / claimer_card_fingerprint / default
        Removed value: -null
      • addedInput schema / properties / claimer_card_fingerprint / maxLength
        Added value: +256
      • removedInput schema / properties / claimer_card_fingerprint / title
        Removed value: -"Claimer Card Fingerprint"
      • addedInput schema / properties / claimer_card_fingerprint / type
        Added value: +[
        +  "string",
        +  "null"
        +]
      • addedInput schema / properties / claimer_id / maxLength
        Added value: +256
      • addedInput schema / properties / claimer_id / minLength
        Added value: +1
      • addedInput schema / properties / claimer_id / pattern
        Added value: +"^[A-Za-z0-9_.:-]+$"
      • removedInput schema / properties / claimer_id / title
        Removed value: -"Claimer Id"
      • removedInput schema / properties / completed_ids / anyOf
        Removed value: -[
        -  {
        -    "items": {
        -      "type": "string"
        -    },
        -    "type": "array"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]
      • removedInput schema / properties / completed_ids / default
        Removed value: -null
      • addedInput schema / properties / completed_ids / items
        Added value: +{
        +  "maxLength": 256,
        +  "minLength": 1,
        +  "pattern": "^[A-Za-z0-9_.:-]+$",
        +  "type": "string"
        +}
      • addedInput schema / properties / completed_ids / maxItems
        Added value: +4096
      • removedInput schema / properties / completed_ids / title
        Removed value: -"Completed Ids"
      • addedInput schema / properties / completed_ids / type
        Added value: +"array"
      • removedInput schema / properties / max_attempts / anyOf
        Removed value: -[
        -  {
        -    "type": "integer"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]
      • removedInput schema / properties / max_attempts / default
        Removed value: -null
      • addedInput schema / properties / max_attempts / maximum
        Added value: +100000
      • addedInput schema / properties / max_attempts / minimum
        Added value: +0
      • removedInput schema / properties / max_attempts / title
        Removed value: -"Max Attempts"
      • addedInput schema / properties / max_attempts / type
        Added value: +[
        +  "integer",
        +  "null"
        +]
      • removedInput schema / properties / project / anyOf
        Removed value: -[
        -  {
        -    "type": "string"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]
      • removedInput schema / properties / project / default
        Removed value: -null
      • addedInput schema / properties / project / maxLength
        Added value: +256
      • removedInput schema / properties / project / title
        Removed value: -"Project"
      • addedInput schema / properties / project / type
        Added value: +[
        +  "string",
        +  "null"
        +]
      • removedInput schema / properties / role / anyOf
        Removed value: -[
        -  {
        -    "type": "string"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]
      • removedInput schema / properties / role / default
        Removed value: -null
      • addedInput schema / properties / role / maxLength
        Added value: +64
      • removedInput schema / properties / role / title
        Removed value: -"Role"
      • addedInput schema / properties / role / type
        Added value: +[
        +  "string",
        +  "null"
        +]
      • changedInput schema / title
        Previous value: -"bernstein_claimArguments"New value: +"bernstein_claim"
    • Addedbernstein_complete
    • Removedbernstein_cost
    • Removedbernstein_create_subtask
    • Removedbernstein_health
    • Changedbernstein_post_artifact41 fields changed
      • addedInput schema / $schema
        Added value: +"http://json-schema.org/draft-07/schema#"
      • addedInput schema / additionalProperties
        Added value: +false
      • addedInput schema / allOf
        Added value: +[
        +  {
        +    "if": {
        +      "properties": {
        +        "artifact_type": {
        +          "const": "report"
        +        }
        +      },
        +      "required": [
        +        "artifact_type"
        +      ]
        +    },
        +    "then": {
        +      "required": [
        +        "body"
        +      ]
        +    }
        +  },
        +  {
        +    "if": {
        +      "properties": {
        +        "artifact_type": {
        +          "const": "table"
        +        }
        +      },
        +      "required": [
        +        "artifact_type"
        +      ]
        +    },
        +    "then": {
        +      "required": [
        +        "columns",
        +        "rows"
        +      ]
        +    }
        +  },
        +  {
        +    "if": {
        +      "properties": {
        +        "artifact_type": {
        +          "const": "link"
        +        }
        +      },
        +      "required": [
        +        "artifact_type"
        +      ]
        +    },
        +    "then": {
        +      "required": [
        +        "url",
        +        "link_kind"
        +      ]
        +    }
        +  }
        +]
      • addedInput schema / properties / artifact_type / enum
        Added value: +[
        +  "report",
        +  "table",
        +  "link"
        +]
      • removedInput schema / properties / artifact_type / title
        Removed value: -"Artifact Type"
      • removedInput schema / properties / body / default
        Removed value: -""
      • addedInput schema / properties / body / maxLength
        Added value: +60000
      • addedInput schema / properties / body / minLength
        Added value: +1
      • removedInput schema / properties / body / title
        Removed value: -"Body"
      • removedInput schema / properties / columns / anyOf
        Removed value: -[
        -  {
        -    "items": {
        -      "type": "string"
        -    },
        -    "type": "array"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]
      • removedInput schema / properties / columns / default
        Removed value: -null
      • addedInput schema / properties / columns / items
        Added value: +{
        +  "maxLength": 256,
        +  "minLength": 1,
        +  "type": "string"
        +}
      • addedInput schema / properties / columns / maxItems
        Added value: +64
      • addedInput schema / properties / columns / minItems
        Added value: +1
      • removedInput schema / properties / columns / title
        Removed value: -"Columns"
      • addedInput schema / properties / columns / type
        Added value: +"array"
      • addedInput schema / properties / key / maxLength
        Added value: +128
      • addedInput schema / properties / key / minLength
        Added value: +1
      • addedInput schema / properties / key / pattern
        Added value: +"^[A-Za-z0-9][A-Za-z0-9_.-]{0,127}$"
      • removedInput schema / properties / key / title
        Removed value: -"Key"
      • removedInput schema / properties / link_kind / default
        Removed value: -""
      • addedInput schema / properties / link_kind / enum
        Added value: +[
        +  "preview",
        +  "dashboard",
        +  "document"
        +]
      • removedInput schema / properties / link_kind / title
        Removed value: -"Link Kind"
      • addedInput schema / properties / poster / maxLength
        Added value: +256
      • addedInput schema / properties / poster / minLength
        Added value: +1
      • removedInput schema / properties / poster / title
        Removed value: -"Poster"
      • removedInput schema / properties / rows / anyOf
        Removed value: -[
        -  {
        -    "items": {
        -      "items": {
        -        "type": "string"
        -      },
        -      "type": "array"
        -    },
        -    "type": "array"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]
      • removedInput schema / properties / rows / default
        Removed value: -null
      • addedInput schema / properties / rows / items
        Added value: +{
        +  "items": {
        +    "maxLength": 4096,
        +    "type": "string"
        +  },
        +  "maxItems": 64,
        +  "type": "array"
        +}
      • addedInput schema / properties / rows / maxItems
        Added value: +4096
      • removedInput schema / properties / rows / title
        Removed value: -"Rows"
      • addedInput schema / properties / rows / type
        Added value: +"array"
      • addedInput schema / properties / task_id / maxLength
        Added value: +256
      • addedInput schema / properties / task_id / minLength
        Added value: +1
      • addedInput schema / properties / task_id / pattern
        Added value: +"^[A-Za-z0-9_.:-]+$"
      • removedInput schema / properties / task_id / title
        Removed value: -"Task Id"
      • removedInput schema / properties / url / default
        Removed value: -""
      • addedInput schema / properties / url / maxLength
        Added value: +4096
      • addedInput schema / properties / url / minLength
        Added value: +1
      • removedInput schema / properties / url / title
        Removed value: -"Url"
      • changedInput schema / title
        Previous value: -"bernstein_post_artifactArguments"New value: +"bernstein_post_artifact"
    • Addedbernstein_post_message
    • Changedbernstein_run33 fields changed
      • addedInput schema / $schema
        Added value: +"http://json-schema.org/draft-07/schema#"
      • addedInput schema / additionalProperties
        Added value: +false
      • addedInput schema / description
        Added value: +"Start an orchestration run. A run does real work and takes minutes to hours; the call returns once the run is queued, not when it finishes. Do not re-issue it while waiting, that starts a second run. Follow the run with bernstein_run_status, passing either the returned task_id or the returned run_id, after waiting the returned poll_after_ms. Pass parent_task_id to create the run as a subtask of an existing task."
      • removedInput schema / properties / complexity / default
        Removed value: -"medium"
      • addedInput schema / properties / complexity / enum
        Added value: +[
        +  "low",
        +  "medium",
        +  "high"
        +]
      • removedInput schema / properties / complexity / title
        Removed value: -"Complexity"
      • removedInput schema / properties / estimated_minutes / default
        Removed value: -30
      • addedInput schema / properties / estimated_minutes / maximum
        Added value: +100000
      • addedInput schema / properties / estimated_minutes / minimum
        Added value: +0
      • removedInput schema / properties / estimated_minutes / title
        Removed value: -"Estimated Minutes"
      • addedInput schema / properties / goal / maxLength
        Added value: +8192
      • addedInput schema / properties / goal / minLength
        Added value: +1
      • removedInput schema / properties / goal / title
        Removed value: -"Goal"
      • addedInput schema / properties / parent_task_id
        Added value: +{
        +  "maxLength": 256,
        +  "minLength": 1,
        +  "pattern": "^[A-Za-z0-9_.:-]+$",
        +  "type": [
        +    "string",
        +    "null"
        +  ]
        +}
      • removedInput schema / properties / priority / default
        Removed value: -2
      • addedInput schema / properties / priority / maximum
        Added value: +3
      • addedInput schema / properties / priority / minimum
        Added value: +1
      • removedInput schema / properties / priority / title
        Removed value: -"Priority"
      • removedInput schema / properties / role / default
        Removed value: -"backend"
      • addedInput schema / properties / role / maxLength
        Added value: +64
      • addedInput schema / properties / role / minLength
        Added value: +1
      • removedInput schema / properties / role / title
        Removed value: -"Role"
      • removedInput schema / properties / scope / default
        Removed value: -"medium"
      • addedInput schema / properties / scope / enum
        Added value: +[
        +  "small",
        +  "medium",
        +  "large"
        +]
      • removedInput schema / properties / scope / title
        Removed value: -"Scope"
      • changedInput schema / title
        Previous value: -"bernstein_runArguments"New value: +"bernstein_run"
      • addedOutput schema / additionalProperties
        Added value: +false
      • addedOutput schema / properties / _meter
        Added value: +{
        +  "additionalProperties": false,
        +  "properties": {
        +    "call_id": {
        +      "type": "string"
        +    },
        +    "cost_usd": {
        +      "type": "number"
        +    },
        +    "error": {
        +      "type": "string"
        +    },
        +    "latency_ms": {
        +      "type": "number"
        +    },
        +    "ok": {
        +      "type": "boolean"
        +    },
        +    "tool": {
        +      "type": "string"
        +    },
        +    "ts": {
        +      "type": "string"
        +    }
        +  },
        +  "required": [
        +    "tool",
        +    "call_id",
        +    "latency_ms",
        +    "cost_usd",
        +    "ok",
        +    "ts"
        +  ],
        +  "type": "object"
        +}
      • addedOutput schema / properties / result / anyOf
        Added value: +[
        +  {
        +    "additionalProperties": false,
        +    "properties": {
        +      "parent_task_id": {
        +        "type": "string"
        +      },
        +      "poll_after_ms": {
        +        "type": "integer"
        +      },
        +      "run_id": {
        +        "type": "string"
        +      },
        +      "status": {
        +        "type": "string"
        +      },
        +      "task_id": {
        +        "type": "string"
        +      },
        +      "title": {
        +        "type": "string"
        +      }
        +    },
        +    "required": [
        +      "task_id",
        +      "title",
        +      "status",
        +      "run_id",
        +      "poll_after_ms"
        +    ],
        +    "type": "object"
        +  },
        +  {
        +    "additionalProperties": true,
        +    "properties": {
        +      "error": {
        +        "type": "string"
        +      },
        +      "hint": {
        +        "type": "string"
        +      }
        +    },
        +    "required": [
        +      "error"
        +    ],
        +    "type": "object"
        +  }
        +]
      • removedOutput schema / properties / result / title
        Removed value: -"Result"
      • removedOutput schema / properties / result / type
        Removed value: -"string"
      • changedOutput schema / required
        Previous value: -[
        -  "result"
        -]New value: +[
        +  "result",
        +  "_meter"
        +]
      • removedOutput schema / title
        Removed value: -"bernstein_runOutput"
    • Addedbernstein_run_status
    • Addedbernstein_shutdown_orchestrator
    • Changedbernstein_status13 fields changed
      • addedInput schema / $schema
        Added value: +"http://json-schema.org/draft-07/schema#"
      • addedInput schema / additionalProperties
        Added value: +false
      • addedInput schema / description
        Added value: +"Liveness, task counts, and cost in one read. Pass status to include the matching tasks; pass detail=true for full per-role and per-task rows."
      • addedInput schema / properties / detail
        Added value: +{
        +  "type": [
        +    "boolean",
        +    "null"
        +  ]
        +}
      • addedInput schema / properties / status
        Added value: +{
        +  "enum": [
        +    null,
        +    "open",
        +    "claimed",
        +    "in_progress",
        +    "done",
        +    "failed",
        +    "blocked",
        +    "cancelled"
        +  ],
        +  "type": [
        +    "string",
        +    "null"
        +  ]
        +}
      • changedInput schema / title
        Previous value: -"bernstein_statusArguments"New value: +"bernstein_status"
      • addedOutput schema / additionalProperties
        Added value: +false
      • addedOutput schema / properties / _meter
        Added value: +{
        +  "additionalProperties": false,
        +  "properties": {
        +    "call_id": {
        +      "type": "string"
        +    },
        +    "cost_usd": {
        +      "type": "number"
        +    },
        +    "error": {
        +      "type": "string"
        +    },
        +    "latency_ms": {
        +      "type": "number"
        +    },
        +    "ok": {
        +      "type": "boolean"
        +    },
        +    "tool": {
        +      "type": "string"
        +    },
        +    "ts": {
        +      "type": "string"
        +    }
        +  },
        +  "required": [
        +    "tool",
        +    "call_id",
        +    "latency_ms",
        +    "cost_usd",
        +    "ok",
        +    "ts"
        +  ],
        +  "type": "object"
        +}
      • addedOutput schema / properties / result / anyOf
        Added value: +[
        +  {
        +    "additionalProperties": false,
        +    "properties": {
        +      "cost": {
        +        "additionalProperties": false,
        +        "properties": {
        +          "per_role": {
        +            "items": {
        +              "additionalProperties": false,
        +              "properties": {
        +                "cost_usd": {
        +                  "type": "number"
        +                },
        +                "role": {
        +                  "type": "string"
        +                }
        +              },
        +              "required": [
        +                "role",
        +                "cost_usd"
        +              ],
        +              "type": "object"
        +            },
        +            "type": "array"
        +          },
        +          "total_cost_usd": {
        +            "type": "number"
        +          }
        +        },
        +        "required": [
        +          "total_cost_usd",
        +          "per_role"
        +        ],
        +        "type": "object"
        +      },
        +      "counts": {
        +        "additionalProperties": false,
        +        "properties": {
        +          "claimed": {
        +            "type": "integer"
        +          },
        +          "done": {
        +            "type": "integer"
        +          },
        +          "failed": {
        +            "type": "integer"
        +          },
        +          "open": {
        +            "type": "integer"
        +          },
        +          "total": {
        +            "type": "integer"
        +          }
        +        },
        +        "required": [
        +          "total",
        +          "open",
        +          "claimed",
        +          "done",
        +          "failed"
        +        ],
        +        "type": "object"
        +      },
        +      "live": {
        +        "type": "boolean"
        +      },
        +      "per_role": {
        +        "items": {
        +          "type": "object"
        +        },
        +        "type": "array"
        +      },
        +      "status_filter": {
        +        "type": "string"
        +      },
        +      "tasks": {
        +        "items": {
        +          "type": "object"
        +        },
        +        "type": "array"
        +      }
        +    },
        +    "required": [
        +      "live",
        +      "counts",
        +      "cost"
        +    ],
        +    "type": "object"
        +  },
        +  {
        +    "additionalProperties": false,
        +    "properties": {
        +      "error": {
        +        "type": "string"
        +      },
        +      "hint": {
        +        "type": "string"
        +      },
        +      "live": {
        +        "type": "boolean"
        +      }
        +    },
        +    "required": [
        +      "live",
        +      "error",
        +      "hint"
        +    ],
        +    "type": "object"
        +  },
        +  {
        +    "additionalProperties": true,
        +    "properties": {
        +      "error": {
        +        "type": "string"
        +      },
        +      "hint": {
        +        "type": "string"
        +      }
        +    },
        +    "required": [
        +      "error"
        +    ],
        +    "type": "object"
        +  }
        +]
      • removedOutput schema / properties / result / title
        Removed value: -"Result"
      • removedOutput schema / properties / result / type
        Removed value: -"string"
      • changedOutput schema / required
        Previous value: -[
        -  "result"
        -]New value: +[
        +  "result",
        +  "_meter"
        +]
      • removedOutput schema / title
        Removed value: -"bernstein_statusOutput"
    • Removedbernstein_stop
    • Addedbernstein_task_capsule
    • Removedbernstein_task_handle
    • Removedbernstein_tasks
    • Removedbernstein_update
    • Changedload_skill23 fields changed
      • addedInput schema / $schema
        Added value: +"http://json-schema.org/draft-07/schema#"
      • addedInput schema / additionalProperties
        Added value: +false
      • addedInput schema / dependencies
        Added value: +{
        +  "reference": [
        +    "name"
        +  ],
        +  "script": [
        +    "name"
        +  ]
        +}
      • addedInput schema / description
        Added value: +"List available skills when name is omitted, or load a named skill body, reference, or script."
      • addedInput schema / properties / name / description
        Added value: +"Skill to load. Omit to return the compact skill index."
      • addedInput schema / properties / name / maxLength
        Added value: +128
      • addedInput schema / properties / name / minLength
        Added value: +1
      • addedInput schema / properties / name / pattern
        Added value: +"^[A-Za-z0-9_.-]+$"
      • removedInput schema / properties / name / title
        Removed value: -"Name"
      • removedInput schema / properties / reference / anyOf
        Removed value: -[
        -  {
        -    "type": "string"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]
      • removedInput schema / properties / reference / default
        Removed value: -null
      • addedInput schema / properties / reference / maxLength
        Added value: +256
      • addedInput schema / properties / reference / pattern
        Added value: +"^[A-Za-z0-9_./-]+$"
      • removedInput schema / properties / reference / title
        Removed value: -"Reference"
      • addedInput schema / properties / reference / type
        Added value: +[
        +  "string",
        +  "null"
        +]
      • removedInput schema / properties / script / anyOf
        Removed value: -[
        -  {
        -    "type": "string"
        -  },
        -  {
        -    "type": "null"
        -  }
        -]
      • removedInput schema / properties / script / default
        Removed value: -null
      • addedInput schema / properties / script / maxLength
        Added value: +256
      • addedInput schema / properties / script / pattern
        Added value: +"^[A-Za-z0-9_./-]+$"
      • removedInput schema / properties / script / title
        Removed value: -"Script"
      • addedInput schema / properties / script / type
        Added value: +[
        +  "string",
        +  "null"
        +]
      • removedInput schema / required
        Removed value: -[
        -  "name"
        -]
      • changedInput schema / title
        Previous value: -"load_skillArguments"New value: +"load_skill"
  4. 3 tool updatesv3.7.0
    • Addedbernstein_claim
    • Addedbernstein_post_artifact
    • Addedbernstein_update
  5. 1 tool updatev3.4.4
    • Addedbernstein_task_handle

TDQS

A3.5/5.0

Scored across 12 tools

Disambiguation4/5

Most tools target a distinct lifecycle action (claim, run, cancel, approve, shutdown), and the descriptions give explicit status constraints that separate complete from approve and status from run_status. The only real ambiguity is between bernstein_status and bernstein_run_status, and between complete/approve, which the status wording helps resolve.

Naming Consistency4/5

The overwhelming majority of tools follow the bernstein_<verb>_<noun> pattern with a consistent snake_case prefix. It is slightly marred by load_skill, which lacks the prefix, and bernstein_task_capsule, which uses a noun rather than an action verb.

Tool Count5/5

Twelve tools is well within the ideal range for an orchestration server and each one earns its place in the run/task/artifact lifecycle. The count feels complete without bloat or redundant utilities.

Completeness3/5

The surface covers run creation, claiming, progress messaging, artifacts, cancellation, shutdown, monitoring, and completion, which is substantial. However, there is no explicit failure/reject path: a task stuck in pending_approval cannot be rejected, and an agent that hits an error has no dedicated way to report failure instead of completing or canceling.

Maintenance

ActivityActive
ResponsivenessWithin a week

Related MCP Connectors

Related MCP Servers

  • -
    license
    B
    quality
    Not graded
    maintenance
    Enables AI-driven orchestration of GitHub development workflows including automated issue analysis, code generation, code review, and PR creation through multiple specialized agents. Integrates with GitHub Actions to automate the complete development process from issue to pull request.
    7
    -
  • F
    license
    Not graded
    quality
    Not graded
    maintenance
    A multi-agent runtime that coordinates six specialized agents through a typed artifact pipeline with 41 RPC methods. It features dynamic autonomy levels and context sufficiency scoring that adjust agent behavior based on the operator's state and task requirements.
    -
  • A
    license
    A
    quality
    D
    maintenance
    Multi-agent orchestration server that enables parallel task delegation, sequential pipelines, cron scheduling, and cross-model peer review via CLI providers like Codex, Antigravity, OpenCode, and Claude Code.
    42
    16 npm
    5
    MIT