Skip to main content
Glama

forge

Front door for the erewhon code agents — a forge CLI and a forge-mcp MCP server over a fleet of coding agents: review ensembles, parallel model comparison, research harnesses, dependency bumping, testing loops, and an architect/worker coding pipeline.

forge <verb> [args] — run forge <verb> --help for each agent's own options.

  • Where tasks go: verbs that emit (audit/testing/refactor/deps, the pipeline) file into the configured task store — a Nous Forge notebook (needs the [nous] extra), GitHub issues with TASK_STORE_BACKEND=github, or git-bug bugs stored in the repo itself with TASK_STORE_BACKEND=git-bug (tasks travel with the clone — no tracker API, no server).

  • Models: the review/analysis ensembles run against any OpenAI-compatible endpoint (point *_OPENAI_BASE_URL-style settings at your router); edit and task/build can also drive Claude/OpenCode per their model settings.

  • MCP: research, review, book are exposed as MCP tools via forge-mcp; the rest are CLI-only.

Research & writing

Example

Does

forge research "why did the Bronze Age collapse?"

Iterative research: plan → research → verify → synthesize (--dry-run, --max-sprints N)

forge book config.yaml

Book-length research via generator–evaluator sprint cycles

Related MCP server: CornMCP

Review & analysis (read-only)

Example

Does

forge review --pr 123 --repo owner/repo

PR review ensemble; --post-comment posts back; --pass digest|supply-chain; or --diff-file x.diff

forge audit crates/foo/src --project Gaol --emit-tasks

Adversarial multi-model audit (discover→dedup→verify); --focus "data loss", --min-severity high, --dry-run-emit

forge testing crates/foo/src --project Gaol --emit-tasks

Find untested behavior → file test-gap tasks (--auto instead generates+gates+pushes tests)

forge refactor crates/foo/src --project Gaol --emit-tasks

Find smells, verify safe+worthwhile, file refactor tasks

forge code-review

Nightly code review of recent commits (markdown by default; optional Nous daily-note sink)

Code generation & autonomy

Example

Does

forge edit --prompt "convert callbacks to async" --repo ~/src/foo --models "claude-opus-4-8,opencode:glm-5.1"

Same prompt, N models (2–26), compare the diffs

forge task --project Gaol

Autonomous worker: pick the top ready task, run it in the sandbox, commit (--dry-run executes then reverts)

forge queue --auto

Backlog report: non-done tasks grouped by project, worker dispatch gate resolved per row; --auto narrows to Auto-OK/Auto-Preferred tasks, --project to one project

forge build plan --epic <slug> --project Meta --repo <path>

Coding pipeline A0+A1 framing; --approve unlocks decompose + emit

forge build run --epic <slug> --project Meta --repo <path>

Orchestrator wave loop (dispatch → verify → replan); --concurrency N

forge build gate --epic <slug> --repo <path>

Final full-quorum epic sign-off (a human merges after)

forge deps --project Meta --dry-run

Dependency bumper: scan + gate; --auto-merge advances main on clean low-risk bumps; Python (uv) / Go / pnpm / Cargo repos auto-detected (--ecosystem to force); --redundancy-report prints a read-only markdown report of overlapping-purpose dependency clusters (uv-only)

forge upstream --dry-run

Upstream sync for additive forks: fetch + merge on a sync branch in a disposable worktree; green-suite + collision-seat gates; --auto-merge advances the default branch when all green

forge sweep --dry-run

Fleet sweep over a Soft Serve instance: enumerate repos via SSH, keep workdir clones fresh, run deps/upstream per repo with git-bug advisories filed in-repo; fail-isolated per repo

Eval

Example

Does

forge evals run

Score models against frozen gold sets and print the report

forge evals baseline / compare

Save a scorecard as baseline / compare a fresh one against it

Install

uv sync                     # library + CLI, no Nous
uv sync --extra nous        # + the Nous task-store backend

# or as a global tool:
uv tool install --editable ~/path/to/forge          # `forge` + `forge-mcp`
uv tool install --editable '~/path/to/forge[nous]'  # with the Nous backend

Machine-local settings (router URL, API key, project paths) live in a gitignored .env at the repo root — every agent's settings read it when run from the repo.

Privacy tiers. Every forge call to the router carries X-Router-Privacy, and the router enforces it on the wire (local: fleet hardware only, no overflow off-box; zdr: local or a cloud endpoint held to zero data retention per request — OpenRouter, never OpenCode Zen; any: no requirement of forge's own). The tier is a required argument on every executor, so no verb can fall back to the permissive default by omission. Who sends what: research/book derive it from the lane (--locallocal, otherwise zdr; --privacy overrides); review sends any on the frontier roster and local on the local roster (PR_REVIEW_ENSEMBLE_FRONTIER_PRIVACY / _LOCAL_PRIVACY); map is local unconditionally; everything else defaults to zdr via a *_ROUTER_PRIVACY setting. A seat the tier excludes comes back from the router as a 403 the harness reports as "refused by privacy policy" and fails over from, never retries. See forge/shared/privacy.py.

Setting up on a work machine (Bedrock strong tiers + a local model)? See docs/RUNBOOK-work-install.md and the router recipe at docs/work/litellm-work.yaml.


Conventions. Path-based verbs (audit/testing/refactor) read files/dirs into one context (no file tools) — scope them to a crate/module, not a whole repo. First pass any emit with --dry-run-emit to preview volume. Emitted tasks land Spec Needed + Manual (proposals), deduped by a stable external_ref, so re-running the same sweep never duplicates.

License

Apache-2.0

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    F
    maintenance
    Enables orchestrating multiple AI CLI agents (Claude Code, Codex, Gemini CLI, Copilot CLI) through a unified MCP interface for task delegation, cross-agent comparison, and specialized tools like code review and debugging.
    14
    13 npm
    14
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables AI coding agents to perform surgical code analysis, semantic memory, and quality enforcement via 18 specialized MCP tools, with a real-time analytics dashboard.
    66
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Provides coding agents with structured planning, persistent project memory, automated verification, and safety permission controls through MCP tools, enabling better planning, context retention, self-checking, and guarded execution.
    6 npm
    MIT