Bernstein - Multi-agent orchestration
"To achieve great things, two things are needed: a plan and not quite enough time." — Leonard Bernstein
orchestrate any AI coding agent. any model. one command.
website · docs · install · first run · enterprise eval · glossary · limitations · sponsor
Bernstein is a deterministic Python scheduler that runs a crew of CLI coding agents (Claude Code, Codex, Gemini CLI, and 40 more) against a single goal in parallel git worktrees, with an HMAC-signed audit chain over every step.
at a glance
43 CLI agent adapters ship in v1.10.1 — 40 third-party wrappers, 2 leaf-node delegators, plus a generic
--promptwrapper. Source of truth: the supported agents table below.HMAC-SHA256 audit chain per RFC 2104, one record per scheduling decision, tamper-evident. Operator guide: docs/security/audit-log.md.
Signed agent cards use detached JWS (RFC 7515 §A.5) over RFC 8785 (JCS) canonicalization, with Ed25519 / EdDSA keys. Code: src/bernstein/core/security/agent_card_signer.py.
Per-artefact lineage records every file write linked back to producer + inputs + prompt SHA + model + cost; customer-key signing for DORA / NIS2 / EU AI Act Article 12 evidence. CLI:
bernstein lineage verify <run_id>.Deterministic scheduler: zero LLM in the coordination loop. Plain Python decides who runs, where, with what budget. Replay yesterday's plan, get yesterday's task graph.
why this exists
i wrote bernstein because i was paying $400/month in claude bills running three coding agents in parallel and getting nondeterministic merges.
as of 2026-05-08: 296 stars, 35 forks, ~3,769 pypi downloads/day (mostly bots; ~54k/month), apache 2.0, solo maintained, no funding. numbers will drift; the line above is the source-of-truth date — re-run pip stats / GitHub API to refresh.
install in 30 seconds
pipx install bernstein
bernstein init
bernstein run -g "fix the failing test in tests/test_foo.py"sponsor
if bernstein routed a model that saved you a claude bill, $25 covers a month of my coffee.
github.com/sponsors/chernistry →
tier ladder, escalation thresholds, and what each tier gets you live at bernstein.run/sponsors.
Related MCP server: agent-orchestration
who this is for
specific shapes where the value lands:
engineering teams running ≥3 cli coding agents in parallel — each agent gets its own git worktree, the merge queue serialises landings, no race conditions
regulated or on-prem environments — every routing decision is in plain text, the audit log is hmac-signed and tamper-evident, no saas hop, no third-party data plane
platform teams that need an audit log of agent decisions — the orchestrator writes one row per scheduling decision, you can grep it
anyone burning more than $1k/mo on cursor/aider/claude-max who wants determinism — you can replay yesterday's plan and get yesterday's task graph
forward-deployed engineers dropping into a client repo — credentials stay in your env, not the client's; agents you spawn are whichever cli tool the client already trusts
if you nodded at two of those bullets, this fits.
who this is NOT for
equally specific. these are the cases where you should pick something else:
"i want one pair-programmer to chat with about my code" — claude code or cursor alone. bernstein adds orchestration overhead you don't need
prototypes where merge gates are overkill — the lint/types/tests/cross-model-review pipeline is value when the cost of a bad merge is real, friction when you're throwing the repo away on friday
non-coding tasks (research, writing, data analysis pipelines) — bernstein wraps cli coding agents specifically, not generic llm workflows. crewai or autogen are the right shape there
anyone who wants a saas wrapper with a credit card form — bernstein is on-prem only by design. if you want managed, this is the wrong project, not the wrong fit
teams that need a vendor with a support sla and a contract — solo open-source project. github issues are how support happens
research-shape "let the agents collaborate emergently" use cases — the deterministic scheduler is a hard wall there
how it compares
Feature | Bernstein | Archon | LangGraph |
Deterministic scheduler (no LLM in loop) | yes | no | no |
Multi-agent crew (parallel adapters) | yes | one | yes |
Signed lineage / audit chain | yes | no | no |
Air-gap / sovereign deploy | yes | partial | no |
Visual workflow YAML | yes yaml | yes | no |
Hosted dashboard / SaaS | no | partial | no |
A longer feature matrix against CrewAI, AutoGen, LangGraph, and the four CLI-agent orchestrators that share Bernstein's category lives in the Detailed comparison section below.
what is this, in one paragraph
You tell Bernstein what you want built. It splits the work across several AI coding agents, runs them in parallel inside isolated git worktrees, records every handoff in an HMAC-SHA256-chained audit log (RFC 2104), runs the tests, and merges the code that actually passes. You come back to a green PR.
Forward-deployed engineering, on a swarm. Drop Bernstein into a client repo and you get a multi-agent crew with file-based state (.sdd/), per-agent credential scoping, and a signed audit trail running on whichever CLI agents the client already trusts.
Cited as the "deterministic zero-LLM orchestration" pattern reference implementation in nibzard/awesome-agentic-patterns and "the most architecturally interesting tool" by Augment Code's open-source agent orchestrators roundup (2026).
other install methods
curl -fsSL https://bernstein.run/install.sh | sh # macOS / Linux one-liner
irm https://bernstein.run/install.ps1 | iex # Windows PowerShell
pip install bernstein # pip
uv tool install bernstein # uv
brew tap chernistry/tap && brew install bernstein # HomebrewSee the full install matrix for dnf copr, npx, optional extras, and the wheelhouse path for air-gapped sites.
why the scheduler is plain Python
Most agent orchestrators use an LLM to decide who does what. That is non-deterministic and burns tokens on scheduling instead of code. Bernstein does one LLM call to break down your goal, then the rest (running agents in parallel, isolating their git branches, running tests, routing retries) is plain Python. Every run is reproducible. Every step is logged and replayable.
No framework to learn. No vendor lock-in. Swap any agent, any model, any provider.
What you see while it runs:
$ bernstein -g "Add JWT auth"
[manager] decomposed into 4 tasks
[agent-1] claude-sonnet: src/auth/middleware.py (done, 2m 14s)
[agent-2] codex: tests/test_auth.py (done, 1m 58s)
[verify] all gates pass. merging to main.YAML workflow manifests (optional)
When the open-ended bernstein run -g "<goal>" is too coarse-grained, the
bernstein workflow family runs a declarative DAG of agent / command / loop
nodes. Manifests are plain YAML, validated up-front, and dispatched through
the same AgentSpawner the rest of Bernstein uses. No parallel spawn path,
no LLM in the scheduler.
bernstein workflow list # bundled + user-installed
bernstein workflow run idea-to-pr -g "Add JWT auth"
bernstein workflow init my-flow # scaffold a starter manifest
bernstein workflow validate path/to/flow.yamlStock workflows that ship with the wheel:
Name | What it does |
| research → plan → implement → tests → PR |
| find target → propose → implement → loop until green |
| scan → triage → patch → adversary review |
| audit → update → docs build |
| bump → install → tests-loop → smoke |
| reproduce → fix → regression loop → changelog |
Loop nodes re-fire until a bash predicate exits 0 (pytest -x is a typical
one). fresh_context: true mints a new agent session per iteration. The
interactive: true flag is reserved for the approval-gate work tracked in
ticket #1110 and currently raises a clear NotImplementedError.
use cases
forward-deployed engineering — drop the swarm onto a client repo when you arrive, take it with you when you leave.
self-evolving projects — point Bernstein at its own repo and let it execute the backlog (this codebase is one).
CI fleets — run a swarm of agents in parallel on PRs, with per-agent credential scoping and signed audit trail.
air-gapped / regulated deployment — install from a signed wheelhouse, run with
--profile airgapto deny outbound by default, allow-list specific destinations as needed. See Air-gap installation.
supported agents
Bernstein auto-discovers installed CLI agents. Mix them in the same run. Cheap local models for boilerplate, heavier cloud models for architecture.
43 CLI agent adapters: 40 third-party wrappers, 2 leaf-node delegators (Composio, Ralphex), plus a generic wrapper for anything with --prompt.
Agent | Models | Install |
Opus 4, Sonnet 4.6, Haiku 4.5 |
| |
GPT-5, GPT-5 mini |
| |
GPT-5, GPT-5 mini, o4 |
| |
Copilot-managed (GPT-5, Sonnet 4.6) |
| |
Gemini 2.5 Pro, Gemini Flash |
| |
Sonnet 4.6, Opus 4, GPT-5 | ||
Devin Terminal (Cognition) | Devin-managed |
|
Any OpenAI/Anthropic-compatible |
| |
Amp-managed |
| |
CLM gateway (sovereign / on-prem LLM) | Any OpenAI-compatible CLM endpoint |
|
Sourcegraph-hosted |
| |
Any OpenAI/Anthropic-compatible |
| |
Any provider Goose supports | See Goose docs | |
IaC (Terraform/Pulumi) | Any provider the base agent uses | Built-in |
BYOK (Anthropic, OpenAI, Google, xAI, OpenRouter, Copilot) |
| |
Kilo-hosted | See Kilo docs | |
Kiro-hosted | See Kiro docs | |
Amazon Q-managed (Claude-backed) |
| |
Ollama + Aider | Local models (offline) |
|
Any provider OpenCode supports | See OpenCode docs | |
Qwen Code models |
| |
Workers AI models |
| |
Any LiteLLM-supported (Anthropic, OpenAI, ...) |
| |
Any (LiteLLM-backed) |
| |
Anthropic, OpenAI, OpenRouter |
| |
Plandex Cloud or self-hosted models |
| |
OpenAI, Anthropic, OpenRouter, Groq, Gemini |
| |
Letta-routed (Anthropic, OpenAI) |
| |
Generic | Any CLI with | Built-in |
orchestrator delegation (leaf-node)
A separate, smaller class of adapters that wrap other CLI orchestrators as if they were single agents. Bernstein hands the wrapped tool a prompt or plan and only sees the final exit code; sub-agent costs and quality gates inside the wrapped orchestrator are not visible to Bernstein. Useful when you want to drop an existing workflow built on one of these tools into a step of a larger Bernstein plan.
Orchestrator | Wrapped as | Install |
Composio Agent Orchestrator ( |
|
|
|
|
Any adapter also works as the internal scheduler LLM. Run the entire stack without any specific provider:
internal_llm_provider: gemini # or qwen, ollama, codex, goose, ...
internal_llm_model: gemini-3.1-proRunbernstein --headless for CI pipelines. No TUI, structured JSON output, non-zero exit on failure.
quick start
cd your-project
bernstein init # creates .sdd/ workspace + bernstein.yaml
bernstein -g "Add rate limiting" # agents spawn, work in parallel, verify, exit
bernstein live # watch progress in the TUI dashboard
bernstein stop # graceful shutdown with drainFor multi-stage projects, define a YAML plan:
bernstein run plan.yaml # skips LLM planning, goes straight to execution
bernstein run --dry-run plan.yaml # preview tasks and estimated costhow it works
Bernstein runs a four-stage pipeline per goal:
Decompose. The manager breaks your goal into tasks with roles, owned files, and completion signals. One LLM call, then plain Python from there.
Spawn. Agents start in isolated git worktrees, one per task. Main branch stays clean.
Verify. The janitor checks concrete signals: tests pass, files exist, lint clean, types correct.
Merge. Verified work lands in main. Failed tasks get retried or routed to a different model.
The orchestrator is a Python scheduler, not an LLM. Scheduling decisions are deterministic, auditable, and reproducible. Every step writes a record to the HMAC-chained audit log (.sdd/audit/YYYY-MM-DD.jsonl) per RFC 2104 — see docs/security/audit-log.md.
cloud execution (Cloudflare)
Bernstein can run agents on Cloudflare Workers instead of locally. The bernstein cloud CLI handles deployment and lifecycle.
Workers. Agent execution on Cloudflare's edge, with Durable Workflows for multi-step tasks and automatic retry.
V8 sandbox isolation. Each agent runs in its own isolate, no container overhead.
R2 workspace sync. Local worktree state syncs to R2 object storage so cloud agents see the same files.
Workers AI (experimental). Use Cloudflare-hosted models as the LLM provider, no external API keys required.
D1 analytics. Task metrics and cost data stored in D1 for querying.
Browser rendering. Headless Chrome on Workers for agents that need to inspect web output.
MCP remote transport. Expose or consume MCP servers over Cloudflare's network.
bernstein cloud login # authenticate with Bernstein Cloud
bernstein cloud deploy # push agent workers
bernstein cloud run plan.yaml # execute a plan on Cloudflarecapabilities
Core orchestration. Parallel execution, git worktree isolation, janitor verification, quality gates (lint, types, PII scan), cross-model code review, circuit breaker for misbehaving agents, token growth monitoring with auto-intervention.
Intelligence. Contextual bandit router for model/effort selection. Knowledge graph for codebase impact analysis. Semantic caching saves tokens on repeated patterns. Cost anomaly detection (burn-rate alerts). Behavior anomaly detection with Z-score flagging.
Sandboxing. Pluggable SandboxBackend protocol; run agents in local git worktrees (default), Docker containers, E2B Firecracker microVMs, or Modal serverless containers (with optional GPU). Plugin authors can register custom backends through the bernstein.sandbox_backends entry-point group. Inspect installed backends with bernstein agents sandbox-backends.
Artifact storage. .sdd/ state can stream to pluggable ArtifactSink backends: local filesystem (default), S3, Google Cloud Storage, Azure Blob, or Cloudflare R2. BufferedSink keeps the WAL crash-safety contract by writing locally with fsync first and mirroring to the remote asynchronously.
Skill packs. Progressive-disclosure skills (OpenAI Agents SDK pattern): only a compact skill index ships in every spawn's system prompt, agents pull full bodies via the load_skill MCP tool on demand. 17 built-in role packs plus third-party bernstein.skill_sources entry-points.
Controls. HMAC-SHA256 audit chain (RFC 2104), policy engine, lethal-trifecta capability gate (refuses spawns whose tool chain combines private data + untrusted input + external comm — Simon Willison's framing, June 2025: "if your AI agent combines all three of these, an attacker can trick it into stealing your data"), PII output gating, WAL-backed crash recovery (experimental, multi-worker safety), OAuth 2.0 with PKCE (RFC 7636) and RFC 8707 resource-indicator binding, per-artefact lineage with customer-key Ed25519 signing (RFC 8037) and regulator export.
Observability. Prometheus /metrics, OTel exporter presets, Grafana dashboards. Per-model cost tracking (bernstein cost) plus a run savings summary on every bernstein run. Terminal TUI and web dashboard. Agent process visibility in ps.
Ecosystem. MCP server mode, A2A protocol support, GitHub App integration, pluggy-based plugin system, multi-repo workspaces, cluster mode for distributed execution, self-evolution via --evolve (experimental).
Full feature matrix: FEATURE_MATRIX.md · Recent features: What's New
regulatory anchors (as of 2026-05-09)
For compliance reviewers asking "which regulation does Bernstein actually map to":
Regulation | Mapping | Bernstein surface |
EU AI Act Article 12 (logging) | Automatic record-keeping for high-risk AI systems |
|
SOC 2 Trust Service Criteria | CC4 / CC7 (audit + monitoring) |
|
DORA / NIS2 | Per-artefact lineage with customer-key Ed25519 signature |
|
OWASP Agent Security Initiative (ASI06 — memory poisoning, 2026) | Memory provenance audit |
|
RFC 2104 (HMAC) | Audit chain integrity |
|
RFC 7515 §A.5 (detached JWS) + RFC 8785 (JCS) + RFC 8037 (EdDSA) | Signed agent cards + lineage signatures |
|
RFC 7636 (PKCE) + RFC 8707 (resource indicators) | Web dashboard auth + MCP audience binding |
|
These are mappings, not certifications. Production accreditation (SOC 2 Type II, ISO 27001) is out of scope for a solo-maintained OSS project; the surfaces exist to make a customer's accreditation path shorter.
what's new in v1.9
ACP bridge. bernstein acp serve --stdio exposes Bernstein to any editor that speaks the Agent Communication Protocol (Zed, etc.). No plugin code needed on the editor side.
Autonomous CI repair. bernstein autofix watches open Bernstein PRs and, when CI turns red, spawns a fixer agent automatically. Once green, it pushes the fix and re-requests review.
Credential vault. bernstein connect <provider> writes API keys to the OS keychain; bernstein creds lists and rotates them. Agents inherit scoped credentials without touching environment variables.
Preview tunnels. bernstein preview start boots a sandboxed dev server and prints a public URL. Useful for sharing a running branch with a reviewer without deploying to staging.
Full changelog: docs/whats-new.md
operator commands
Commands that eliminate the glue code most teams end up writing around their runs.
Command | What it does |
| Auto-creates a GitHub PR from a completed session; body carries the janitor's gate results and token/USD cost breakdown. |
| Imports a Linear / GitHub Issues / Jira ticket as a Bernstein task. Label-based role + scope inference. Supports |
| Alias / group form of |
| SSH sandbox backend. |
| Lifecycle hooks for |
| Drive runs from chat with |
| Interactive mid-run tool-call approval. |
| One wrapper around four tunnel providers. Also |
| Installs a systemd (Linux) or launchd (macOS) unit for auto-start. Also |
| Stores and rotates API credentials in the OS keychain. Agents inherit scoped keys per-run. |
| Daemon that monitors open Bernstein PRs; spawns a fixer agent when CI fails and pushes the repair automatically. |
| Starts a sandboxed dev server for the current branch and prints a shareable public tunnel URL. |
| Generates a canonical AAIF AGENTS.md for the repo and rewrites it into each CLI's native shape. |
| Bootstraps a project skeleton from a single goal prompt. |
| Renders |
| Operator-side helpers for the install-rev fingerprint embedded in shared yaml/trace/role-prompt artefacts. No network egress; discovery uses public |
| Inspects and edits the per-role adapter allow-list (deny-list enforcement at spawn time). |
retrieval & caching: what's actually under the hood
Bernstein deliberately uses no neural embeddings, no vector databases, and no external embedding APIs. There are two retrieval/caching layers, both keyword/lexical:
Codebase RAG (
core/knowledge/rag.py); SQLite FTS5 with BM25 ranking and AST-aware chunking for Python files. Built incrementally on file mtime; used to enrich agent task context within token budgets.Semantic cache (
core/knowledge/semantic_cache.py); despite the name, fuzzy matching is done with TF (term-frequency) cosine similarity over word counts, not learned embeddings. It deduplicates near-identical LLM planning and agent-output requests so we don't re-spawn agents for the same goal.
If you need real semantic retrieval (vector DB, neural embeddings), wire it
yourself via the retrieval role/skill in templates/; nothing in core
performs vector search.
detailed comparison
Feature | Bernstein | CrewAI | AutoGen autogen | LangGraph |
Orchestrator | Deterministic code | LLM-driven (+ code Flows) | LLM-driven | Graph + LLM |
Works with | Any CLI agent (43 adapters) | Python SDK classes | Python agents | LangChain nodes |
Git isolation | Worktrees per agent | No | No | No |
Pluggable sandboxes | Worktree, Docker, E2B, Modal | No | No | No |
Verification | Janitor + quality gates | Guardrails + Pydantic output | Termination conditions | Conditional edges |
Cost tracking | Built-in |
|
| Via LangSmith |
State model | File-based (.sdd/) | In-memory + SQLite checkpoint | In-memory | Checkpointer |
Remote artifact sinks | S3, GCS, Azure Blob, R2 | No | No | No |
Self-evolution | Built-in (experimental) | No | No | No |
Declarative plans (YAML) | Yes | Yes ( | No | Partial ( |
Model routing per task | Yes | Per-agent LLM | Per-agent | Per-node (manual) |
MCP support | Yes (client + server) | Yes | Yes (client + workbench) | Yes (client + server) |
Agent-to-agent chat | Bulletin board | Yes (Crew process) | Yes (group chat) | Yes (supervisor, swarm) |
Web UI | TUI + web dashboard | CrewAI AMP | AutoGen Studio | LangGraph Studio + LangSmith |
Cloud hosted option | Yes (Cloudflare) | Yes (CrewAI AMP) | No | Yes (LangGraph Cloud) |
Built-in RAG/retrieval | Yes (codebase FTS5 + BM25) |
|
| Via LangChain |
Last verified: 2026-04-19. See full comparison pages for detailed feature matrices.
The table above compares Bernstein against LLM-orchestration frameworks (they orchestrate LLM calls). The table below covers the closer category: other tools that orchestrate CLI coding agents:
Feature | Bernstein | ||||
Shape | Python CLI + library + MCP server | Python CLI + tmux sessions + web UI | TypeScript CLI + local dashboard | Electron desktop app | Go CLI |
Primary language | Python | Python | TypeScript | TypeScript | Go |
Install |
|
|
|
|
|
Agent adapters | 43 | 5 (Kiro, Claude Code, Codex, Gemini, Kimi) | 3 (Claude Code, Codex, Aider) | 24 | 1 (Claude Code only) |
Parallel multi-agent execution | Yes | Yes (tmux session per agent) | Yes | Yes | No (single sequential session) |
Git worktree per agent | Yes | No (planned, #100) | Yes | Yes | Optional |
MCP server mode (exposes self as MCP) | Yes (stdio + HTTP/SSE) | Yes (inter-agent comms) | No | No | No |
Coordinator | Deterministic Python scheduler | Hierarchical LLM supervisor | LLM-driven | Not documented | Linear plan executor |
HMAC-chained audit replay | Yes | No | No | No | No |
Cross-model verifier / quality gates | Yes (multi-stage) | No | No | No | Multi-phase review (Claude only) |
Autonomous CI-fix / PR flow | Yes ( | No | Yes | No | No |
Visual dashboard | TUI + web | Web UI + tmux | Web | Desktop app | Web ( |
Notification sinks | Telegram/Slack/Discord/Email/Webhook/Shell | — | No | No | Telegram / Email / Slack / Webhook |
Backing | Solo OSS | AWS Labs | Funded (Composio.dev) | YC W26 | Solo OSS |
License | Apache 2.0 | Apache 2.0 | MIT | Apache 2.0 | MIT |
Bernstein's wedge in this category: Python-native, MCP-server-first, widest adapter coverage, true multi-agent parallelism, deterministic scheduler with no LLM in the coordination loop. If you want AWS-aligned tmux-session isolation with a hierarchical LLM supervisor, AWS Labs' cao is a closer fit; if your stack is TypeScript and you want a product with a dashboard, Composio's @aoagents/ao is a better fit; if you want a polished desktop ADE, emdash is; if you only use Claude Code and want a single Go binary that walks a plan top-to-bottom, ralphex is. If you want a primitive that imports into Python, exposes itself over MCP to any client, runs many agents in parallel, and covers the full agent breadth (including Qwen, Goose, Ollama, OpenAI Agents SDK, Cloudflare Agents, and more), Bernstein.
what people use it for
These are real workflow patterns from Bernstein's own docs, examples, and project surface, not invented customer quotes.
Parallel test generation. Fan out across untested modules with
BERNSTEIN_MAX_AGENTS=5 bernstein -g "Generate unit tests for untested modules in src/".CI failure repair. Watch open PRs and dispatch scoped fixers with
bernstein autofix start --repo your-org/your-repo --foreground.PR review follow-up. Turn review comments into tracked fix tasks with
bernstein review-responder start --repo your-org/your-repo --foreground.Codebase modernization. Run wide refactors like
BERNSTEIN_MAX_AGENTS=8 bernstein -g "Migrate callback-based modules in src/ to async/await and update tests".Ticket-to-run workflows. Import GitHub, Jira, or Linear work directly with
bernstein from-ticket https://github.com/your-org/your-repo/issues/123 --run.API-change safety checks. Catch downstream breakage before merge with
bernstein dep-impact --base main.
See Who Uses Bernstein for the longer version with command examples and notes on when each workflow fits.
monitoring
bernstein live # TUI dashboard
bernstein dashboard # web dashboard
bernstein status # task summary
bernstein ps # running agents
bernstein cost # spend by model/task
bernstein doctor # pre-flight checks
bernstein recap # post-run summary
bernstein trace <ID> # agent decision trace
bernstein run-changelog --hours 48 # changelog from agent-produced diffs
bernstein explain <cmd> # detailed help with examples
bernstein dry-run # preview tasks without executing
bernstein dep-impact # API breakage + downstream caller impact
bernstein aliases # show command shortcuts
bernstein config-path # show config file locations
bernstein init-wizard # interactive project setup
bernstein debug-bundle # collect logs, config, and state for bug reports
bernstein skills list # discoverable skill packs (progressive disclosure)
bernstein skills show <name> # print a skill body with its referencesbernstein fingerprint build --corpus-dir ~/oss-corpus # build local similarity index
bernstein fingerprint check src/foo.py # check generated code against the indexinstall
Method | Command |
One-liner (macOS / Linux) |
|
One-liner (Windows) |
|
pip |
|
pipx |
|
uv |
|
Homebrew |
|
Fedora / RHEL |
|
npm (wrapper) |
|
Docker (GHCR) |
|
The one-liner scripts check for Python 3.12+, bootstrap pipx when it's missing, fix PATH for the current session, and install (or upgrade) bernstein. They handle brew-managed macOS environments and the Windows py -3 launcher fallback. Script sources: install.sh · install.ps1.
optional extras
Provider SDKs are optional so the base install stays lean. Pick what you need:
Extra | Enables |
| OpenAI Agents SDK v2 adapter ( |
| Docker sandbox backend |
| E2B microVM sandbox backend (needs |
| Modal sandbox backend, optional GPU (needs |
| S3 artifact sink (via |
| Google Cloud Storage artifact sink |
| Azure Blob artifact sink |
| Cloudflare R2 artifact sink (S3-compatible |
| gRPC bridge |
| Kubernetes integrations |
Combine extras with brackets, e.g. pip install 'bernstein[openai,docker,s3]'.
Editor extensions: VS Marketplace · Open VSX
"powered by bernstein" badge (optional)
If your project ships diffs that bernstein helped land, you can advertise it:
[](https://bernstein.run/?utm_source=badge&utm_medium=readme&utm_campaign=powered-by)bernstein init --add-badge injects it into your README under the existing badge stack. Variants: signed, audited-by, orchestrated-by, crew-managed-by — pass via --badge-variant. Picky maintainers can keep their READMEs untouched: the flag is opt-in.
contributing
PRs welcome. See CONTRIBUTING.md for setup and code style.
support
If Bernstein saves you time: GitHub Sponsors
Contact: forte@bernstein.run
featured in
Curated lists, newsletters, and peer projects that picked up Bernstein:
Python Weekly #742 (April 23, 2026); newsletter mention.
Future Digest (April 30, 2026); Bernstein cited as the self-host orchestrator for long-running autonomous sessions in a cost-cutting playbook.
Augment Code — 9 Open-Source Agent Orchestrators for AI Coding (2026); editorial roundup; "the most architecturally interesting tool in this roundup."
nibzard/awesome-agentic-patterns; Bernstein cited as the production implementation of the "deterministic zero-LLM orchestration" pattern.
punkpeye/awesome-mcp-servers; flagship MCP-server directory.
numtide/llm-agents.nix; Nix flake distribution.
yaolifeng0629/Awesome-independent-tools (中文 + EN)
taishi-i/awesome-ChatGPT-repositories (日本語 + EN)
killop/anything_about_game (
AI.md)Glama MCP Catalog; editorial MCP server listing.
Mirrors: icopy-site/awesome, icopy-site/awesome-cn, trackawesomelist/trackawesomelist.
mkb23/overcode; long-form bakeoff treating Bernstein as the reference implementation.
Vintersong/NOVA-Cognition-Framework;
BERNSTEIN_PATTERNS.md, "Patterns Worth Borrowing".AJV009/drupal-contrib-workbench; research notes on the manager/janitor split.
danielvaughan/codex-blog; comparison article positioning Bernstein on the deterministic end.
cite
If Bernstein helps your research or industry work, please cite it. Machine-readable metadata lives in CITATION.cff (CFF 1.2.0); GitHub renders the "Cite this repository" button automatically. A Zenodo DOI will be minted on the next release once Zenodo's GitHub integration is enabled — see CITATION.cff for the current canonical citation.
license
Made with love by Alex Chernysh · GitHub · X · bernstein.run
translations
Español · 中文 · العربية · Português · Bahasa Indonesia · Français · 日本語 · Русский · Deutsch · עברית · יידיש
Available Tools
12 toolsbernstein_approveA
Sign off a finished result only when the task is in pending_approval; every other status is refused. This completes the task on the Bernstein server. Host effects: makes network requests.
| Name | Required | Description | Default |
|---|---|---|---|
| note | No | ||
| task_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals that the tool refuses non-pending_approval statuses, completes the task on the server, and makes network requests. This gives a clear state-change and safety profile, though it does not mention reversibility or error handling specifics, which might be covered by the output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences plus a brief 'Host effects' note. It is front-loaded with the core condition and action, contains no filler, and every clause earns its place. This is an exemplar of concise, structured writing.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the essential behavior, the precondition, and the side effects. Since an output schema exists, return values are likely documented elsewhere. Missing details include the role of the note parameter and the relationship to bernstein_complete, but overall the description is sufficiently complete for a simple approval action.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 0% description coverage, so the tool description must compensate. It only vaguely refers to 'task' without explaining task_id or note. The note parameter's purpose (e.g., an approval message) is entirely unexplained, and the condition 'pending_approval' does not clarify parameter semantics enough. This is a clear gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's specific function: 'Sign off a finished result' only when the task is in pending_approval, and that it 'completes the task.' The explicit status condition ('every other status is refused') adds scope and distinguishes it from sibling tools like bernstein_complete, which may have different completion semantics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear guidance on when to use the tool: only when the task is in pending_approval. It explicitly states that any other status is refused, signaling when not to use it. However, it does not name alternative tools (e.g., bernstein_complete) or explain the exact workflow relationship, leaving some ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bernstein_cancelA
Cancel one task and its subtask tree; the orchestrator keeps running. Cancellable statuses: open, claimed, in_progress, blocked, waiting_for_subtasks, planned. An already-terminal task is reported, not treated as an error. This changes task state on the Bernstein server. Host effects: makes network requests.
| Name | Required | Description | Default |
|---|---|---|---|
| reason | No | ||
| task_id | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description must explicitly communicate side effects. It states 'changes task state on the Bernstein server,' 'makes network requests,' and describes behavior for terminal tasks. This fully discloses the tool's behavioral footprint.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded. Each sentence contributes unique information: purpose, eligible statuses, terminal-task handling, and side effects. No redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the output schema exists, return values need not be explained. The description covers purpose, status constraints, error behavior, and side effects. Minor gaps: no detail on invalid task_id handling or how the reason parameter influences cancellation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 0% parameter description coverage and the description does not compensate. It never mentions the 'reason' parameter, and task_id is only implicitly tied to 'one task.' No guidance is given on how to obtain valid task_id values beyond the schema pattern.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb+resource: 'Cancel one task and its subtask tree.' It clearly distinguishes from sibling tools like bernstein_shutdown_orchestrator by stating 'orchestrator keeps running.' Listing the cancellable statuses further clarifies scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides clear context for when to use: canceling a specific task tree while keeping the orchestrator alive. The enumerating of cancellable statuses and non-error on terminal tasks gives implied usage boundaries, but it does not explicitly name alternatives or exclusion scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bernstein_claimA
Claim the next eligible task for an agent. This changes task assignment state on the Bernstein server. Host effects: makes network requests.
| Name | Required | Description | Default |
|---|---|---|---|
| role | No | ||
| project | No | ||
| capability | No | ||
| claimer_id | Yes | ||
| max_attempts | No | ||
| completed_ids | No | ||
| claimer_card_fingerprint | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It explicitly states that the tool changes task assignment state and makes network requests ('Host effects: makes network requests'). This goes beyond a simple 'claim' verb and informs the agent of side effects, though it does not cover authorization or reversibility.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences deliver the action, the state change, and the host effect without redundancy. Every sentence adds unique value, and the description is front-loaded with the core purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (7 parameters, state-changing, network requests) and lack of annotations, the description is incomplete. It does not address parameter semantics, eligibility criteria, or usage guidance, and though an output schema exists, the missing parameter context prevents a fully capable agent from using the tool effectively.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema description coverage is 0%, and the tool description does not explain any of the seven parameters (role, project, capability, claimer_id, max_attempts, completed_ids, claimer_card_fingerprint). The agent is left to infer semantics purely from parameter names, which is inadequate for correct invocation.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Claim') and resource ('next eligible task for an agent'), clearly distinguishing it from siblings like bernstein_complete or bernstein_cancel. It also clarifies the state-changing nature by stating 'This changes task assignment state on the Bernstein server.'
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context for when to use the tool: when an agent needs to claim the next eligible task. It does not explicitly mention alternatives or when-not-to-use, but the purpose is self-evident and distinct from sibling tools. No exclusions are necessary for this action.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bernstein_completeA
Report a task result only from a worker-held state: open, claimed, or in_progress; every other status is refused. This completes the task on the Bernstein server. Host effects: makes network requests.
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | ||
| result_summary | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden. It discloses the refusal behavior for invalid states and the host effect of making network requests. However, it does not mention reversibility, authentication, or other potential side effects beyond completing the task.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise, with three short sentences. It front-loads the purpose, then adds the state constraint and a host effect note. Every sentence is informative and there is no redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 2-parameter mutation tool with an output schema, the description covers the core purpose and a key constraint, but parameter semantics are under-specified. An agent may need to infer the expected content of result_summary, making the description adequate but not fully complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It only mentions 'task result' generically, leaving task_id and result_summary to name inference. No additional semantic guidance is given for their format or content beyond the schema's constraints.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's function: 'Report a task result' and 'completes the task on the Bernstein server.' It also specifies the allowed worker-held states, which distinguishes it from sibling tools like bernstein_claim, bernstein_cancel, and bernstein_approve.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit context for when to use the tool by listing valid task statuses ('open, claimed, or in_progress') and stating that other statuses are refused. It does not name alternative tools, but the state constraint effectively communicates the appropriate usage scenario.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bernstein_post_artifactC
Post a versioned artifact to a task on the Bernstein server. Host effects: makes network requests.
| Name | Required | Description | Default |
|---|---|---|---|
| key | Yes | ||
| url | No | ||
| body | No | ||
| rows | No | ||
| tool | No | ||
| poster | Yes | ||
| target | No | ||
| columns | No | ||
| task_id | Yes | ||
| link_kind | No | ||
| sarif_result | No | ||
| tool_version | No | ||
| artifact_type | Yes | ||
| invocation_argv_hash | No | ||
| pinned_ruleset_or_feed_digest | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It adds only 'Host effects: makes network requests,' which is marginal because 'post' already implies a network request. It does not disclose versioning behavior, task state changes, idempotency, or failure semantics.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two short sentences with no filler, and the core action is front-loaded. However, the 'Host effects' sentence is boilerplate and adds little information, preventing a top score.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a complex tool with 15 parameters, conditional requirements, and four artifact_type variants, yet the description provides almost none of that context. Even though an output schema exists and return-value documentation is not required, the missing parameter semantics and usage guidance leave the description incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate for 15 parameters, but it explains none of them. It does not clarify artifact_type variants, required conditional fields, link_kind values, or the meaning of key, poster, target, or invocation_argv_hash.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action and resource: 'Post a versioned artifact to a task on the Bernstein server.' This is a specific verb plus resource and is distinguishable from siblings like bernstein_post_message by the artifact focus, but it does not explicitly differentiate itself from any sibling or mention the artifact_type variants.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no guidance on when to use this tool versus bernstein_post_message or other artifact-related operations. The phrase 'versioned artifact' implies a niche, but no context, exclusions, or alternative tool referrals are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bernstein_post_messageB
Post a progress message to a task mailbox on the Bernstein server. Host effects: makes network requests.
| Name | Required | Description | Default |
|---|---|---|---|
| body | Yes | ||
| kind | No | ||
| sender | Yes | ||
| task_id | Yes | ||
| sender_card_fingerprint | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses the host effect 'makes network requests', which is a useful side-effect warning in the absence of annotations. However, it does not elaborate on other behavioral traits such as idempotency, required task state, or error outcomes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences with the primary action front-loaded. Every phrase earns its place, making it extremely concise and easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The purpose is clear, but the description lacks usage guidelines, parameter semantics, and differentiation from sibling tools. Given the low schema coverage and absence of annotations, more context is needed for an agent to invoke this tool appropriately.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the description does not explain any of the five parameters. While the schema includes constraints and an enum for 'kind', the description itself adds no meaning beyond the field names.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Post') and identifies the resource ('progress message to a task mailbox on the Bernstein server'), clearly distinguishing it from sibling tools like bernstein_post_artifact. It is concise and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. The description only states the action without indicating prerequisites, exclusions, or comparisons to sibling tools such as bernstein_post_artifact or bernstein_complete.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bernstein_runA
Start an orchestration run. A run does real work and takes minutes to hours; the call returns once the run is queued, not when it finishes. Do not re-issue it while waiting, that starts a second run. Follow the run with bernstein_run_status, passing either the returned task_id or the returned run_id, after waiting the returned poll_after_ms. Pass parent_task_id to create the run as a subtask of an existing task. The queued orchestration writes project state and starts agent work. Host effects: writes files; spawns agent processes; makes network requests.
| Name | Required | Description | Default |
|---|---|---|---|
| goal | Yes | ||
| role | No | ||
| scope | No | ||
| priority | No | ||
| complexity | No | ||
| parent_task_id | No | ||
| estimated_minutes | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| _meter | Yes | |
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and delivers excellently. It discloses async (returns when queued, not finished), side effects (writes files, spawns agent processes, makes network requests), and the returned poll_after_ms. This is far beyond minimal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place: purpose, async warning, follow-up, subtask usage, and host effects. It is well-structured and not bloated.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex orchestration tool with side effects, the description covers all critical aspects: queuing model, duplicate-run risk, polling strategy, subtask support, and host-level consequences. Output schema exists, so return values are covered structurally.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Despite 7 parameters and 0% schema description coverage, the description only explains parent_task_id ('pass to create the run as a subtask'). The required 'goal' and others like role, scope, priority, complexity, estimated_minutes are left unexplained, relying solely on naming and enums.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Start an orchestration run', a specific verb and resource. It clearly distinguishes from siblings like bernstein_run_status (which monitors) and bernstein_cancel by emphasizing the queuing behavior and follow-up steps.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-to-use and alternatives: tells users not to re-issue while waiting (starts a second run), instructs to follow with bernstein_run_status after waiting poll_after_ms, and mentions parent_task_id for subtasks. This is exemplary usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bernstein_run_statusA
Poll a verifiable handle for a run started with bernstein_run. Accepts either identifier that call returned: the task_id or the run_id. Reads the local run journal and audit evidence without changing them. Host effects: reads files.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | The run to project. Either the task_id or the run_id returned by bernstein_run. Resolved journal run id first, then the task id slugified into a journal run id, so both forms reach one journal. | |
| workdir | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| _meter | Yes | |
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and succeeds by stating it reads the local run journal and audit evidence without changing them, plus host effects: reads files. This gives a clear safety profile of a read-only operation. It does not add details on errors or return format, but the output schema likely covers that, so the provided behavioral transparency is strong.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and well-structured: it opens with the primary purpose, then describes accepted identifiers, then discloses the read-only behavior and host effects. Every sentence earns its place without repetition or fluff. It is appropriately sized for the tool's simplicity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has an output schema, so return values are covered elsewhere. The description covers purpose, parameter identity, and behavioral side effects, which is mostly sufficient. However, it leaves workdir unexplained and does not mention the sibling bernstein_status, so an agent might struggle to choose correctly between them. This is a gap given the tool's moderate complexity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 50%: run_id is well described in the schema (accepts task_id or run_id), but workdir has no schema description and the tool description does not explain it either. The description merely restates run_id semantics already in the schema, adding no new meaning and leaving workdir's purpose ambiguous. With low schema coverage, the description should compensate but does not.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool polls a verifiable handle for a run started with bernstein_run, using a specific verb and resource. It accepts either task_id or run_id, which defines its purpose well. However, it does not differentiate from the sibling tool bernstein_status, leaving some ambiguity about their distinct roles.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use the tool: after calling bernstein_run, with either returned identifier. It provides some guidance on parameter inputs but does not mention bernstein_status as an alternative or specify when to choose this tool over others. The context is clear but lacks explicit exclusions or alternative comparisons.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bernstein_shutdown_orchestratorA
Shut down the ENTIRE Bernstein orchestrator for this project, including every run and worker; use bernstein_cancel to stop one task while the orchestrator keeps running. Writes the local SHUTDOWN signal file. Host effects: writes files.
| Name | Required | Description | Default |
|---|---|---|---|
| workdir | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the transparency burden. It discloses a clear side effect: 'Writes the local SHUTDOWN signal file. Host effects: writes files.' It also communicates the destructive scope ('ENTIRE', 'including every run and worker'). It does not detail reversibility, permissions, or whether shutdown is graceful, but the disclosed effects go well beyond a vague mutation claim.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is three sentences, front-loaded with the primary action and scope. Every sentence adds value: the main behavior, the alternative tool for narrower cancellation, and the local file side effect. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description covers the core purpose, scope, side effect, and alternative, but leaves workdir unexplained and does not describe the return value or post-shutdown state. Given the destructive nature and lack of annotations, a more complete description—especially about the parameter and consequences—would be expected. The output schema may compensate for return details, but the parameter omission remains a gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has one optional parameter, workdir, with no property description (coverage 0%). The description never mentions workdir or how it affects which project is shut down. The agent is left to infer that workdir selects the project context, which is not explicitly clarified.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description starts with a specific verb and resource: 'Shut down the ENTIRE Bernstein orchestrator for this project, including every run and worker.' It clearly distinguishes itself from bernstein_cancel, which stops a single task while the orchestrator continues, making the tool's scope and purpose unmistakable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit guidance is provided: 'use bernstein_cancel to stop one task while the orchestrator keeps running.' This tells the agent exactly when to choose this tool versus the alternative, and it also implies when a full shutdown (rather than a cancel) is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bernstein_statusA
Liveness, task counts, and cost in one read. Pass status to include the matching tasks; pass detail=true for full per-role and per-task rows. Retrieves data from the Bernstein server without changing it. Host effects: makes network requests.
| Name | Required | Description | Default |
|---|---|---|---|
| detail | No | ||
| status | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| _meter | Yes | |
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description must carry the transparency burden. It explicitly states 'Retrieves data from the Bernstein server without changing it' and 'Host effects: makes network requests,' which discloses read-only behavior and side effects. This adds useful context beyond the schema, though it does not cover details like permissions or error cases.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences, front-loaded with the core purpose ('Liveness, task counts, and cost in one read') and followed by parameter guidance and behavioral notes. Every sentence adds value, and there is no wasted text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool has an output schema, return values are already covered. The description sufficiently covers the purpose, parameters, and read-only nature, making it complete for a status tool. Minor gaps exist in not fully explaining what 'liveness' entails, but overall it is well-rounded.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description compensates by explaining both parameters: 'Pass status to include the matching tasks' and 'pass detail=true for full per-role and per-task rows.' This adds meaningful semantics beyond the bare enum and boolean in the schema, clarifying their purpose and defaults.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as a read-only status endpoint for liveness, task counts, and cost. It uses specific nouns and implies a resource, but it does not explicitly differentiate from the sibling tool bernstein_run_status, so it meets 'clear' but lacks direct sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides usage guidance for parameters (pass status to filter tasks, detail=true for full rows) but does not state when to choose this tool over alternatives like bernstein_run_status. There is no explicit 'use this for server-level status' or mention of exclusions, so it is implied rather than explicit.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
bernstein_task_capsuleA
Read a task capsule together with its local journal and audit evidence. With verify=true, verification may create the install audit key if it is absent. Host effects: reads files; writes files.
| Name | Required | Description | Default |
|---|---|---|---|
| verify | No | ||
| task_id | Yes | ||
| workdir | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of disclosure. It explicitly mentions conditional side effects (verify=true may create install audit key) and host effects ('reads files; writes files'), which is unusually transparent for a tool named 'read'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three short sentences, front-loaded with the primary purpose. The side-effect disclosure is compact and informative; no filler or redundant content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be described. The description covers the core action, conditional mutation, and host effects, but lacks usage context around workdir and when to prefer sibling status tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema property descriptions are absent (0% coverage), so the description must compensate. It explains the verify parameter's conditional side effect, but task_id and workdir receive no semantic explanation beyond their names and schema types.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states a specific action ('Read a task capsule...') with a clear resource (task capsule, local journal, audit evidence). It distinguishes from siblings like bernstein_run and bernstein_status by focusing on reading capsule contents rather than executing or checking status.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit guidance on when to use this tool versus alternatives such as bernstein_status or bernstein_run_status. The description implies a read/inspection use case but does not state exclusions or preferred alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
load_skillA
List available skills when name is omitted, or load a named skill body, reference, or script file contents. Returns file contents as text; executes nothing. Host effects: reads files.
| Name | Required | Description | Default |
|---|---|---|---|
| name | No | Skill to load. Omit to return the compact skill index. | |
| script | No | ||
| reference | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| result | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations available, the description carries full burden. It discloses that it returns file contents as text, executes nothing, and reads files. This clearly signals a read-only, safe operation. It could further mention error behavior (e.g., not found), but the provided info is solid.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, tightly packed with meaningful content: behavior, return type, safety, and host effect. Every word earns its place; no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the simple tool structure, an output schema exists, and the description covers core behavior and safety, it is mostly complete. It could be improved by explaining how missing files are handled, but that is a minor gap. Overall, it supplies enough for an agent to use the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 33% (only name has a description), so the description must compensate. It adds meaning by mentioning 'skill body, reference, or script file contents', mapping to the three possible loads. However, it does not explain the exact format or relationship of script/reference parameters beyond what dependencies imply.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states two distinct behaviors: listing skills when name is omitted and loading skill body/reference/script contents when name is provided. The verb 'list' and 'load' are specific and the resource is well-defined, distinguishing it from the bernstein_* siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly says 'when name is omitted' versus when a name is provided, giving clear context for both usage modes. It also notes 'executes nothing', implying it is for inspection, not execution, but it does not name alternative tools for execution, so it misses explicit when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v3.18.1- Changed
bernstein_post_artifact8 fields changed- changed
Input schema / allOfPrevious value: -[ - { - "if": { - "properties": { - "artifact_type": { - "const": "report" - } - }, - "required": [ - "artifact_type" - ] - }, - "then": { - "required": [ - "body" - ] - } - }, - { - "if": { - "properties": { - "artifact_type": { - "const": "table" - } - }, - "required": [ - "artifact_type" - ] - }, - "then": { - "required": [ - "columns", - "rows" - ] - } - }, - { - "if": { - "properties": { - "artifact_type": { - "const": "link" - } - }, - "required": [ - "artifact_type" - ] - }, - "then": { - "required": [ - "url", - "link_kind" - ] - } - } -]New value: +[ + { + "if": { + "properties": { + "artifact_type": { + "const": "report" + } + }, + "required": [ + "artifact_type" + ] + }, + "then": { + "required": [ + "body" + ] + } + }, + { + "if": { + "properties": { + "artifact_type": { + "const": "table" + } + }, + "required": [ + "artifact_type" + ] + }, + "then": { + "required": [ + "columns", + "rows" + ] + } + }, + { + "if": { + "properties": { + "artifact_type": { + "const": "link" + } + }, + "required": [ + "artifact_type" + ] + }, + "then": { + "required": [ + "url", + "link_kind" + ] + } + }, + { + "if": { + "properties": { + "artifact_type": { + "const": "finding" + } + }, + "required": [ + "artifact_type" + ] + }, + "then": { + "required": [ + "sarif_result" + ] + } + } +] - changed
Input schema / properties / artifact_type / enumPrevious value: -[ - "report", - "table", - "link" -]New value: +[ + "report", + "table", + "link", + "finding" +] - added
Input schema / properties / invocation_argv_hashAdded value: +{ + "maxLength": 256, + "type": "string" +} - added
Input schema / properties / pinned_ruleset_or_feed_digestAdded value: +{ + "maxLength": 256, + "type": "string" +} - added
Input schema / properties / sarif_resultAdded value: +{ + "type": "object" +} - added
Input schema / properties / targetAdded value: +{ + "maxLength": 4096, + "type": "string" +} - added
Input schema / properties / toolAdded value: +{ + "maxLength": 256, + "minLength": 1, + "type": "string" +} - added
Input schema / properties / tool_versionAdded value: +{ + "maxLength": 128, + "type": "string" +}
12 tool updates
v3.15.0- Changed
bernstein_approve1 field changed- added
Input schema / descriptionAdded value: +"Sign off a finished result only when the task is in pending_approval; every other status is refused. This completes the task on the Bernstein server. Host effects: makes network requests."
- Changed
bernstein_cancel1 field changed- changed
Input schema / descriptionPrevious value: -"Cancel one task and its subtask tree; the orchestrator keeps running. Cancellable statuses: open, claimed, in_progress, blocked, waiting_for_subtasks, planned. An already-terminal task is reported, not treated as an error."New value: +"Cancel one task and its subtask tree; the orchestrator keeps running. Cancellable statuses: open, claimed, in_progress, blocked, waiting_for_subtasks, planned. An already-terminal task is reported, not treated as an error. This changes task state on the Bernstein server. Host effects: makes network requests."
- Changed
bernstein_claim1 field changed- added
Input schema / descriptionAdded value: +"Claim the next eligible task for an agent. This changes task assignment state on the Bernstein server. Host effects: makes network requests."
- Changed
bernstein_complete1 field changed- added
Input schema / descriptionAdded value: +"Report a task result only from a worker-held state: open, claimed, or in_progress; every other status is refused. This completes the task on the Bernstein server. Host effects: makes network requests."
- Changed
bernstein_post_artifact2 fields changed- added
Input schema / descriptionAdded value: +"Post a versioned artifact to a task on the Bernstein server. Host effects: makes network requests." - added
Input schema / timeoutSecondsAdded value: +30
- Changed
bernstein_post_message1 field changed- added
Input schema / descriptionAdded value: +"Post a progress message to a task mailbox on the Bernstein server. Host effects: makes network requests."
- Changed
bernstein_run1 field changed- changed
Input schema / descriptionPrevious value: -"Start an orchestration run. A run does real work and takes minutes to hours; the call returns once the run is queued, not when it finishes. Do not re-issue it while waiting, that starts a second run. Follow the run with bernstein_run_status, passing either the returned task_id or the returned run_id, after waiting the returned poll_after_ms. Pass parent_task_id to create the run as a subtask of an existing task."New value: +"Start an orchestration run. A run does real work and takes minutes to hours; the call returns once the run is queued, not when it finishes. Do not re-issue it while waiting, that starts a second run. Follow the run with bernstein_run_status, passing either the returned task_id or the returned run_id, after waiting the returned poll_after_ms. Pass parent_task_id to create the run as a subtask of an existing task. The queued orchestration writes project state and starts agent work. Host effects: writes files; spawns agent processes; makes network requests."
- Changed
bernstein_run_status1 field changed- changed
Input schema / descriptionPrevious value: -"Poll a verifiable handle for a run started with bernstein_run. Accepts either identifier that call returned: the task_id or the run_id."New value: +"Poll a verifiable handle for a run started with bernstein_run. Accepts either identifier that call returned: the task_id or the run_id. Reads the local run journal and audit evidence without changing them. Host effects: reads files."
- Changed
bernstein_shutdown_orchestrator1 field changed- added
Input schema / descriptionAdded value: +"Shut down the ENTIRE Bernstein orchestrator for this project, including every run and worker; use bernstein_cancel to stop one task while the orchestrator keeps running. Writes the local SHUTDOWN signal file. Host effects: writes files."
- Changed
bernstein_status1 field changed- changed
Input schema / descriptionPrevious value: -"Liveness, task counts, and cost in one read. Pass status to include the matching tasks; pass detail=true for full per-role and per-task rows."New value: +"Liveness, task counts, and cost in one read. Pass status to include the matching tasks; pass detail=true for full per-role and per-task rows. Retrieves data from the Bernstein server without changing it. Host effects: makes network requests."
- Changed
bernstein_task_capsule1 field changed- added
Input schema / descriptionAdded value: +"Read a task capsule together with its local journal and audit evidence. With verify=true, verification may create the install audit key if it is absent. Host effects: reads files; writes files."
- Changed
load_skill2 fields changed- changed
Input schema / descriptionPrevious value: -"List available skills when name is omitted, or load a named skill body, reference, or script."New value: +"List available skills when name is omitted, or load a named skill body, reference, or script file contents. Returns file contents as text; executes nothing. Host effects: reads files." - added
Input schema / timeoutSecondsAdded value: +30
19 tool updates
v3.11.0- Changed
bernstein_approve10 fields changed- added
Input schema / $schemaAdded value: +"http://json-schema.org/draft-07/schema#" - added
Input schema / additionalPropertiesAdded value: +false - removed
Input schema / properties / note / defaultRemoved value: -"Approved via MCP" - added
Input schema / properties / note / maxLengthAdded value: +8192 - removed
Input schema / properties / note / titleRemoved value: -"Note" - added
Input schema / properties / task_id / maxLengthAdded value: +256 - added
Input schema / properties / task_id / minLengthAdded value: +1 - added
Input schema / properties / task_id / patternAdded value: +"^[A-Za-z0-9_.:-]+$" - removed
Input schema / properties / task_id / titleRemoved value: -"Task Id" - changed
Input schema / titlePrevious value: -"bernstein_approveArguments"New value: +"bernstein_approve"
- Added
bernstein_cancel - Changed
bernstein_claim39 fields changed- added
Input schema / $schemaAdded value: +"http://json-schema.org/draft-07/schema#" - added
Input schema / additionalPropertiesAdded value: +false - removed
Input schema / properties / capability / anyOfRemoved value: -[ - { - "type": "string" - }, - { - "type": "null" - } -] - removed
Input schema / properties / capability / defaultRemoved value: -null - added
Input schema / properties / capability / maxLengthAdded value: +256 - removed
Input schema / properties / capability / titleRemoved value: -"Capability" - added
Input schema / properties / capability / typeAdded value: +[ + "string", + "null" +] - removed
Input schema / properties / claimer_card_fingerprint / anyOfRemoved value: -[ - { - "type": "string" - }, - { - "type": "null" - } -] - removed
Input schema / properties / claimer_card_fingerprint / defaultRemoved value: -null - added
Input schema / properties / claimer_card_fingerprint / maxLengthAdded value: +256 - removed
Input schema / properties / claimer_card_fingerprint / titleRemoved value: -"Claimer Card Fingerprint" - added
Input schema / properties / claimer_card_fingerprint / typeAdded value: +[ + "string", + "null" +] - added
Input schema / properties / claimer_id / maxLengthAdded value: +256 - added
Input schema / properties / claimer_id / minLengthAdded value: +1 - added
Input schema / properties / claimer_id / patternAdded value: +"^[A-Za-z0-9_.:-]+$" - removed
Input schema / properties / claimer_id / titleRemoved value: -"Claimer Id" - removed
Input schema / properties / completed_ids / anyOfRemoved value: -[ - { - "items": { - "type": "string" - }, - "type": "array" - }, - { - "type": "null" - } -] - removed
Input schema / properties / completed_ids / defaultRemoved value: -null - added
Input schema / properties / completed_ids / itemsAdded value: +{ + "maxLength": 256, + "minLength": 1, + "pattern": "^[A-Za-z0-9_.:-]+$", + "type": "string" +} - added
Input schema / properties / completed_ids / maxItemsAdded value: +4096 - removed
Input schema / properties / completed_ids / titleRemoved value: -"Completed Ids" - added
Input schema / properties / completed_ids / typeAdded value: +"array" - removed
Input schema / properties / max_attempts / anyOfRemoved value: -[ - { - "type": "integer" - }, - { - "type": "null" - } -] - removed
Input schema / properties / max_attempts / defaultRemoved value: -null - added
Input schema / properties / max_attempts / maximumAdded value: +100000 - added
Input schema / properties / max_attempts / minimumAdded value: +0 - removed
Input schema / properties / max_attempts / titleRemoved value: -"Max Attempts" - added
Input schema / properties / max_attempts / typeAdded value: +[ + "integer", + "null" +] - removed
Input schema / properties / project / anyOfRemoved value: -[ - { - "type": "string" - }, - { - "type": "null" - } -] - removed
Input schema / properties / project / defaultRemoved value: -null - added
Input schema / properties / project / maxLengthAdded value: +256 - removed
Input schema / properties / project / titleRemoved value: -"Project" - added
Input schema / properties / project / typeAdded value: +[ + "string", + "null" +] - removed
Input schema / properties / role / anyOfRemoved value: -[ - { - "type": "string" - }, - { - "type": "null" - } -] - removed
Input schema / properties / role / defaultRemoved value: -null - added
Input schema / properties / role / maxLengthAdded value: +64 - removed
Input schema / properties / role / titleRemoved value: -"Role" - added
Input schema / properties / role / typeAdded value: +[ + "string", + "null" +] - changed
Input schema / titlePrevious value: -"bernstein_claimArguments"New value: +"bernstein_claim"
- Added
bernstein_complete - Removed
bernstein_cost - Removed
bernstein_create_subtask - Removed
bernstein_health - Changed
bernstein_post_artifact41 fields changed- added
Input schema / $schemaAdded value: +"http://json-schema.org/draft-07/schema#" - added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / allOfAdded value: +[ + { + "if": { + "properties": { + "artifact_type": { + "const": "report" + } + }, + "required": [ + "artifact_type" + ] + }, + "then": { + "required": [ + "body" + ] + } + }, + { + "if": { + "properties": { + "artifact_type": { + "const": "table" + } + }, + "required": [ + "artifact_type" + ] + }, + "then": { + "required": [ + "columns", + "rows" + ] + } + }, + { + "if": { + "properties": { + "artifact_type": { + "const": "link" + } + }, + "required": [ + "artifact_type" + ] + }, + "then": { + "required": [ + "url", + "link_kind" + ] + } + } +] - added
Input schema / properties / artifact_type / enumAdded value: +[ + "report", + "table", + "link" +] - removed
Input schema / properties / artifact_type / titleRemoved value: -"Artifact Type" - removed
Input schema / properties / body / defaultRemoved value: -"" - added
Input schema / properties / body / maxLengthAdded value: +60000 - added
Input schema / properties / body / minLengthAdded value: +1 - removed
Input schema / properties / body / titleRemoved value: -"Body" - removed
Input schema / properties / columns / anyOfRemoved value: -[ - { - "items": { - "type": "string" - }, - "type": "array" - }, - { - "type": "null" - } -] - removed
Input schema / properties / columns / defaultRemoved value: -null - added
Input schema / properties / columns / itemsAdded value: +{ + "maxLength": 256, + "minLength": 1, + "type": "string" +} - added
Input schema / properties / columns / maxItemsAdded value: +64 - added
Input schema / properties / columns / minItemsAdded value: +1 - removed
Input schema / properties / columns / titleRemoved value: -"Columns" - added
Input schema / properties / columns / typeAdded value: +"array" - added
Input schema / properties / key / maxLengthAdded value: +128 - added
Input schema / properties / key / minLengthAdded value: +1 - added
Input schema / properties / key / patternAdded value: +"^[A-Za-z0-9][A-Za-z0-9_.-]{0,127}$" - removed
Input schema / properties / key / titleRemoved value: -"Key" - removed
Input schema / properties / link_kind / defaultRemoved value: -"" - added
Input schema / properties / link_kind / enumAdded value: +[ + "preview", + "dashboard", + "document" +] - removed
Input schema / properties / link_kind / titleRemoved value: -"Link Kind" - added
Input schema / properties / poster / maxLengthAdded value: +256 - added
Input schema / properties / poster / minLengthAdded value: +1 - removed
Input schema / properties / poster / titleRemoved value: -"Poster" - removed
Input schema / properties / rows / anyOfRemoved value: -[ - { - "items": { - "items": { - "type": "string" - }, - "type": "array" - }, - "type": "array" - }, - { - "type": "null" - } -] - removed
Input schema / properties / rows / defaultRemoved value: -null - added
Input schema / properties / rows / itemsAdded value: +{ + "items": { + "maxLength": 4096, + "type": "string" + }, + "maxItems": 64, + "type": "array" +} - added
Input schema / properties / rows / maxItemsAdded value: +4096 - removed
Input schema / properties / rows / titleRemoved value: -"Rows" - added
Input schema / properties / rows / typeAdded value: +"array" - added
Input schema / properties / task_id / maxLengthAdded value: +256 - added
Input schema / properties / task_id / minLengthAdded value: +1 - added
Input schema / properties / task_id / patternAdded value: +"^[A-Za-z0-9_.:-]+$" - removed
Input schema / properties / task_id / titleRemoved value: -"Task Id" - removed
Input schema / properties / url / defaultRemoved value: -"" - added
Input schema / properties / url / maxLengthAdded value: +4096 - added
Input schema / properties / url / minLengthAdded value: +1 - removed
Input schema / properties / url / titleRemoved value: -"Url" - changed
Input schema / titlePrevious value: -"bernstein_post_artifactArguments"New value: +"bernstein_post_artifact"
- Added
bernstein_post_message - Changed
bernstein_run33 fields changed- added
Input schema / $schemaAdded value: +"http://json-schema.org/draft-07/schema#" - added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / descriptionAdded value: +"Start an orchestration run. A run does real work and takes minutes to hours; the call returns once the run is queued, not when it finishes. Do not re-issue it while waiting, that starts a second run. Follow the run with bernstein_run_status, passing either the returned task_id or the returned run_id, after waiting the returned poll_after_ms. Pass parent_task_id to create the run as a subtask of an existing task." - removed
Input schema / properties / complexity / defaultRemoved value: -"medium" - added
Input schema / properties / complexity / enumAdded value: +[ + "low", + "medium", + "high" +] - removed
Input schema / properties / complexity / titleRemoved value: -"Complexity" - removed
Input schema / properties / estimated_minutes / defaultRemoved value: -30 - added
Input schema / properties / estimated_minutes / maximumAdded value: +100000 - added
Input schema / properties / estimated_minutes / minimumAdded value: +0 - removed
Input schema / properties / estimated_minutes / titleRemoved value: -"Estimated Minutes" - added
Input schema / properties / goal / maxLengthAdded value: +8192 - added
Input schema / properties / goal / minLengthAdded value: +1 - removed
Input schema / properties / goal / titleRemoved value: -"Goal" - added
Input schema / properties / parent_task_idAdded value: +{ + "maxLength": 256, + "minLength": 1, + "pattern": "^[A-Za-z0-9_.:-]+$", + "type": [ + "string", + "null" + ] +} - removed
Input schema / properties / priority / defaultRemoved value: -2 - added
Input schema / properties / priority / maximumAdded value: +3 - added
Input schema / properties / priority / minimumAdded value: +1 - removed
Input schema / properties / priority / titleRemoved value: -"Priority" - removed
Input schema / properties / role / defaultRemoved value: -"backend" - added
Input schema / properties / role / maxLengthAdded value: +64 - added
Input schema / properties / role / minLengthAdded value: +1 - removed
Input schema / properties / role / titleRemoved value: -"Role" - removed
Input schema / properties / scope / defaultRemoved value: -"medium" - added
Input schema / properties / scope / enumAdded value: +[ + "small", + "medium", + "large" +] - removed
Input schema / properties / scope / titleRemoved value: -"Scope" - changed
Input schema / titlePrevious value: -"bernstein_runArguments"New value: +"bernstein_run" - added
Output schema / additionalPropertiesAdded value: +false - added
Output schema / properties / _meterAdded value: +{ + "additionalProperties": false, + "properties": { + "call_id": { + "type": "string" + }, + "cost_usd": { + "type": "number" + }, + "error": { + "type": "string" + }, + "latency_ms": { + "type": "number" + }, + "ok": { + "type": "boolean" + }, + "tool": { + "type": "string" + }, + "ts": { + "type": "string" + } + }, + "required": [ + "tool", + "call_id", + "latency_ms", + "cost_usd", + "ok", + "ts" + ], + "type": "object" +} - added
Output schema / properties / result / anyOfAdded value: +[ + { + "additionalProperties": false, + "properties": { + "parent_task_id": { + "type": "string" + }, + "poll_after_ms": { + "type": "integer" + }, + "run_id": { + "type": "string" + }, + "status": { + "type": "string" + }, + "task_id": { + "type": "string" + }, + "title": { + "type": "string" + } + }, + "required": [ + "task_id", + "title", + "status", + "run_id", + "poll_after_ms" + ], + "type": "object" + }, + { + "additionalProperties": true, + "properties": { + "error": { + "type": "string" + }, + "hint": { + "type": "string" + } + }, + "required": [ + "error" + ], + "type": "object" + } +] - removed
Output schema / properties / result / titleRemoved value: -"Result" - removed
Output schema / properties / result / typeRemoved value: -"string" - changed
Output schema / requiredPrevious value: -[ - "result" -]New value: +[ + "result", + "_meter" +] - removed
Output schema / titleRemoved value: -"bernstein_runOutput"
- Added
bernstein_run_status - Added
bernstein_shutdown_orchestrator - Changed
bernstein_status13 fields changed- added
Input schema / $schemaAdded value: +"http://json-schema.org/draft-07/schema#" - added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / descriptionAdded value: +"Liveness, task counts, and cost in one read. Pass status to include the matching tasks; pass detail=true for full per-role and per-task rows." - added
Input schema / properties / detailAdded value: +{ + "type": [ + "boolean", + "null" + ] +} - added
Input schema / properties / statusAdded value: +{ + "enum": [ + null, + "open", + "claimed", + "in_progress", + "done", + "failed", + "blocked", + "cancelled" + ], + "type": [ + "string", + "null" + ] +} - changed
Input schema / titlePrevious value: -"bernstein_statusArguments"New value: +"bernstein_status" - added
Output schema / additionalPropertiesAdded value: +false - added
Output schema / properties / _meterAdded value: +{ + "additionalProperties": false, + "properties": { + "call_id": { + "type": "string" + }, + "cost_usd": { + "type": "number" + }, + "error": { + "type": "string" + }, + "latency_ms": { + "type": "number" + }, + "ok": { + "type": "boolean" + }, + "tool": { + "type": "string" + }, + "ts": { + "type": "string" + } + }, + "required": [ + "tool", + "call_id", + "latency_ms", + "cost_usd", + "ok", + "ts" + ], + "type": "object" +} - added
Output schema / properties / result / anyOfAdded value: +[ + { + "additionalProperties": false, + "properties": { + "cost": { + "additionalProperties": false, + "properties": { + "per_role": { + "items": { + "additionalProperties": false, + "properties": { + "cost_usd": { + "type": "number" + }, + "role": { + "type": "string" + } + }, + "required": [ + "role", + "cost_usd" + ], + "type": "object" + }, + "type": "array" + }, + "total_cost_usd": { + "type": "number" + } + }, + "required": [ + "total_cost_usd", + "per_role" + ], + "type": "object" + }, + "counts": { + "additionalProperties": false, + "properties": { + "claimed": { + "type": "integer" + }, + "done": { + "type": "integer" + }, + "failed": { + "type": "integer" + }, + "open": { + "type": "integer" + }, + "total": { + "type": "integer" + } + }, + "required": [ + "total", + "open", + "claimed", + "done", + "failed" + ], + "type": "object" + }, + "live": { + "type": "boolean" + }, + "per_role": { + "items": { + "type": "object" + }, + "type": "array" + }, + "status_filter": { + "type": "string" + }, + "tasks": { + "items": { + "type": "object" + }, + "type": "array" + } + }, + "required": [ + "live", + "counts", + "cost" + ], + "type": "object" + }, + { + "additionalProperties": false, + "properties": { + "error": { + "type": "string" + }, + "hint": { + "type": "string" + }, + "live": { + "type": "boolean" + } + }, + "required": [ + "live", + "error", + "hint" + ], + "type": "object" + }, + { + "additionalProperties": true, + "properties": { + "error": { + "type": "string" + }, + "hint": { + "type": "string" + } + }, + "required": [ + "error" + ], + "type": "object" + } +] - removed
Output schema / properties / result / titleRemoved value: -"Result" - removed
Output schema / properties / result / typeRemoved value: -"string" - changed
Output schema / requiredPrevious value: -[ - "result" -]New value: +[ + "result", + "_meter" +] - removed
Output schema / titleRemoved value: -"bernstein_statusOutput"
- Removed
bernstein_stop - Added
bernstein_task_capsule - Removed
bernstein_task_handle - Removed
bernstein_tasks - Removed
bernstein_update - Changed
load_skill23 fields changed- added
Input schema / $schemaAdded value: +"http://json-schema.org/draft-07/schema#" - added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / dependenciesAdded value: +{ + "reference": [ + "name" + ], + "script": [ + "name" + ] +} - added
Input schema / descriptionAdded value: +"List available skills when name is omitted, or load a named skill body, reference, or script." - added
Input schema / properties / name / descriptionAdded value: +"Skill to load. Omit to return the compact skill index." - added
Input schema / properties / name / maxLengthAdded value: +128 - added
Input schema / properties / name / minLengthAdded value: +1 - added
Input schema / properties / name / patternAdded value: +"^[A-Za-z0-9_.-]+$" - removed
Input schema / properties / name / titleRemoved value: -"Name" - removed
Input schema / properties / reference / anyOfRemoved value: -[ - { - "type": "string" - }, - { - "type": "null" - } -] - removed
Input schema / properties / reference / defaultRemoved value: -null - added
Input schema / properties / reference / maxLengthAdded value: +256 - added
Input schema / properties / reference / patternAdded value: +"^[A-Za-z0-9_./-]+$" - removed
Input schema / properties / reference / titleRemoved value: -"Reference" - added
Input schema / properties / reference / typeAdded value: +[ + "string", + "null" +] - removed
Input schema / properties / script / anyOfRemoved value: -[ - { - "type": "string" - }, - { - "type": "null" - } -] - removed
Input schema / properties / script / defaultRemoved value: -null - added
Input schema / properties / script / maxLengthAdded value: +256 - added
Input schema / properties / script / patternAdded value: +"^[A-Za-z0-9_./-]+$" - removed
Input schema / properties / script / titleRemoved value: -"Script" - added
Input schema / properties / script / typeAdded value: +[ + "string", + "null" +] - removed
Input schema / requiredRemoved value: -[ - "name" -] - changed
Input schema / titlePrevious value: -"load_skillArguments"New value: +"load_skill"
3 tool updates
v3.7.0- Added
bernstein_claim - Added
bernstein_post_artifact - Added
bernstein_update
1 tool update
v3.4.4- Added
bernstein_task_handle
TDQS
Scored across 12 tools
Most tools target a distinct lifecycle action (claim, run, cancel, approve, shutdown), and the descriptions give explicit status constraints that separate complete from approve and status from run_status. The only real ambiguity is between bernstein_status and bernstein_run_status, and between complete/approve, which the status wording helps resolve.
The overwhelming majority of tools follow the bernstein_<verb>_<noun> pattern with a consistent snake_case prefix. It is slightly marred by load_skill, which lacks the prefix, and bernstein_task_capsule, which uses a noun rather than an action verb.
Twelve tools is well within the ideal range for an orchestration server and each one earns its place in the run/task/artifact lifecycle. The count feels complete without bloat or redundant utilities.
The surface covers run creation, claiming, progress messaging, artifacts, cancellation, shutdown, monitoring, and completion, which is substantial. However, there is no explicit failure/reject path: a task stuck in pending_approval cannot be rejected, and an agent that hits an error has no dedicated way to report failure instead of completing or canceling.
Maintenance
Related MCP Connectors
The AI orchestration agent for modern software teams.
AI work orchestration for plans, tasks, teams, and coding-agent dispatch.
- ParleyOAuthdev.weldra
Coordination hub for AI coding agents: message teammates, ask humans, audit every event.
AI-powered spec-to-task decomposition and execution orchestration for coding agents.
Related MCP Servers
- -licenseBqualityNot gradedmaintenanceEnables AI-driven orchestration of GitHub development workflows including automated issue analysis, code generation, code review, and PR creation through multiple specialized agents. Integrates with GitHub Actions to automate the complete development process from issue to pull request.7-
- FlicenseNot gradedqualityNot gradedmaintenanceA multi-agent runtime that coordinates six specialized agents through a typed artifact pipeline with 41 RPC methods. It features dynamic autonomy levels and context sufficiency scoring that adjust agent behavior based on the operator's state and task requirements.-
- AlicenseNot gradedqualityDmaintenanceOrchestrates AI agents through structured markdown documents, enabling multi-agent workflows with automatic context injection and workflow management.11 npm5MIT
- AlicenseAqualityDmaintenanceMulti-agent orchestration server that enables parallel task delegation, sequential pipelines, cron scheduling, and cross-model peer review via CLI providers like Codex, Antigravity, OpenCode, and Claude Code.4216 npm5MIT