Skip to main content
Glama

Capitoline

In the Grand Temple, before the end, Desmond does not find a single voice. He finds three: Jupiter, Juno, Minerva. The Capitoline Triad, what remains of Those Who Came Before. They do not agree. Minerva asks him to let the world burn and start again. Juno asks him to save it, at a price. Jupiter has already spoken, and is silent. The answer does not come from consensus. It comes from the one who listens to all of them, and then chooses.

Capitoline is a self-hosted personal AI gateway. It exposes an OpenAI-compatible HTTP API and an MCP server, and behind them it runs the official Claude Code, Codex and Gemini CLIs, authenticated with the subscriptions of whoever hosts it. Three voices, one endpoint.

You can give the floor to a single member: claude-fable, codex-gpt-6-astra, antigravity-gemini-pro, or any other model the three CLIs serve, since every one of them is exposed by name. Or the capitoline model convenes the Triad: every member answers, every member judges the others without knowing who wrote what, and a judge synthesizes. As in the Temple, the value is not in the agreement but in hearing the dissent before deciding.

The council is Andrej Karpathy's idea, and the credit is his: llm-council is a local web app that sends a question to several models through OpenRouter, has them review and rank each other's answers anonymously, and lets a chairman model write the final response. Capitoline keeps those three stages and changes what is around them. The models are reached through their official CLIs, on your own subscriptions, with no pay-per-use API, and the council is a model name any OpenAI client or MCP client can ask for. Each seat is a model family with a fallback chain, not a single model. The judge is seated apart and blind by default: llm-council's chairman is a member, and here that is an option (judge_allow_member). The synthesis builds on the top-ranked answer and asserts nothing the answers do not support. And every change to the strategy was measured before it shipped (docs/measurements/).

Run it

Requirements: Linux for a deployment (Debian or Ubuntu; macOS works for development), Node 24.x (pinned in .nvmrc; engines is >=24 <25), and the CLIs claude, codex, agy installed and logged in, on subscriptions of your own, for the user that runs them. The shipped configuration seats all three in its councils and sizes its concurrency for a host of about 8 GB (the startup log says whether yours fits, design §4.1). A CLI that is missing or signed out is reported unhealthy and its models unavailable; the councils then seat the members that are left.

On a development machine, a short overlay runs the CLIs as yourself instead of through sudo, and keeps the sandboxes and the database in the working copy:

mkdir -p tmp && cat > tmp/dev-overlay.yaml <<'EOF'
runner: { user: null, sandbox_root: ./tmp/sandboxes }
usage: { db_path: ./tmp/usage.sqlite }
EOF
npm ci && npm run build
CAPITOLINE_OVERLAY=tmp/dev-overlay.yaml node dist/main.js    # base config/capitoline.yaml, listens on 127.0.0.1:8080

To work on it, git config core.hooksPath .githooks enables a check that refuses a commit carrying a session link or a key (.githooks/public-check, which also reads a private pattern list of your own if you keep one).

A production host — a separate runner user that alone holds the CLI logins, a systemd service, backups, updates — is described step by step in docs/deploy.md. Cloudflare Tunnel and Access are optional there: the gateway issues and checks its own API keys, so it can serve a network of your own with nothing in front.

Connecting another application of your own: docs/connecting-an-application.md, which covers the two credentials — a key issued by the gateway (Authorization: Bearer cap_…, managed through /v1/admin/keys or npm run keys on the host) and a Cloudflare service token (scripts/cf-service-token.sh creates one and binds it to a name) — where the secret goes and what it may be used for. Every call is recorded under the caller's name: GET /v1/usage is the per-caller and per-model breakdown.

Related MCP server: consult-mcp

Use it

curl http://127.0.0.1:8080/v1/models
curl http://127.0.0.1:8080/v1/chat/completions -H 'content-type: application/json' \
  -d '{"model":"codex-gpt-6-luna","reasoning_effort":"low","messages":[{"role":"user","content":"Reply with the single word: ok"}]}'
curl http://127.0.0.1:8080/v1/images/generations -H 'content-type: application/json' \
  -d '{"prompt":"a red fox in the snow, 16:9"}' | jq -r '.data[0].b64_json' | base64 -d > fox.jpg
curl -N http://127.0.0.1:8080/v1/chat/completions -H 'content-type: application/json' \
  -d '{"model":"capitoline","stream":true,"messages":[{"role":"user","content":"Is a retry after a refusal worth one more call?"}]}'
curl -N http://127.0.0.1:8080/v1/chat/completions -H 'content-type: application/json' \
  -d '{"model":"capitoline","reasoning_effort":"low","stream":true,"messages":[{"role":"user","content":"Same question, five calls: no peer ranking."}]}'

GET /v1/usage answers two questions: .callers is who spent the last day, and .models is which real model served each gateway name over the last week. The second exists because the configuration names CLI aliases rather than dated ids — opus meant Opus 5 until 2026-09-22 and Opus 5.5 after it, with nothing here changed — so two rows under one name is an alias that moved, and without them every measurement would be undated underneath.

Any OpenAI-compatible client works by setting its base URL to /v1 (Open WebUI, the official SDKs, LiteLLM). Supported: model, messages (text and base64 image parts), stream, reasoning_effort (low, medium, high, xhigh, max, ultra — each provider prices the levels its CLI accepts and a request asking for one it does not runs at the nearest, never at the CLI's own default). Rejected with 400: tools, n>1, logprobs, response_format. Ignored with the X-Capitoline-Ignored header: temperature, top_p, max_tokens and other sampling knobs. Responses carry an extra capitoline field, which names the provider that served the call and, when the CLI reported one, cliModelId: the dated id of the model that actually answered. model stays the name you asked for, as an OpenAI client expects. Only Claude reports an id, because only its names are aliases — a Codex slug and an Antigravity id are the model itself.

Conversations work the standard Chat Completions way: the API is stateless, and a client that wants a second turn sends the whole history again in messages — which is what Open WebUI and the SDKs do. The gateway keeps no conversation state: it renders the history as labelled turns into one prompt, and every call runs in a fresh CLI process with its sessions disabled. Each turn therefore costs the whole history in tokens. The Responses API (/v1/responses, with previous_response_id) is not served; why, and how it would be built if a client ever needs it, is in docs/backlog.md.

Images: POST /v1/images/generations serves the models declared with kind: image (the kind is reported by /v1/models); omit model and the first available image model answers. Two exist: codex-image, Codex's built-in generation on the ChatGPT subscription, which comes first and so is the default; and antigravity-image, whose quota is small (12 per 5 hours, 58 per week) and which answers when a client names it or when Codex's is unavailable. Neither uses an API key. One image per call, returned inline as b64_json, with the real mime, width, height and bytes in the capitoline field. Rejected with 400: n other than 1, a response_format other than b64_json, an output_format other than jpeg (the gateway returns the format the CLI produced), a chat request against an image model and an image request against a text model. Ignored with the X-Capitoline-Ignored header: size, quality, style and the other style knobs — the CLI's image tool takes only a prompt, so there is nothing to map a size onto. A generation takes 11-45 s and the provider's quota is small: docs/spike-2026-09.md §8 has the two windows.

Council: model: capitoline convenes the Triad instead of giving the floor to one member. One question is nine calls in three stages — four independent answers, four anonymous peer rankings of those answers, one synthesis written by a judge seated apart from the panel — spread over three subscriptions, so it costs about what nine direct requests cost and takes minutes rather than seconds. The synthesis is the message content, so a client that knows nothing of the council reads an ordinary completion; the rest is in the capitoline.council field: every member with its real model name, its label and its answer, which seats fell back and which were lost, each member's ranking and the panel's aggregate, the judge and whether it was blind, the shape that ran and the strategy version, and a deliberation id — the same id that ties that deliberation's rows in the usage table together. Below two answers no council takes place: with one, that answer is returned as its member wrote it and the field says plainly that nobody ranked it. Ask for it with stream: true through a tunnel or any other proxy (docs/deploy.md §9): the silent stages then send a progress chunk per stage and per member ({"stage":"rankings","done":2,"total":4} in the chunk's own capitoline field, which an OpenAI client ignores) and the synthesis streams as ordinary content. The seats, the judge's chain, the quorum, the ranking stage and the per-stage timeout are configuration (council: in config/capitoline.yaml); the three prompts are not — they are the strategy itself, they live in src/council/prompts.ts, and changing them changes the model's name.

Two councils are configured, and a client asks for either the same way: in model, or as the council argument of the MCP tool ask_council. The price is one call per seat, one more per seat when the ranking stage runs, and one for the judge. The shape is also a request option, as the effort is for every model: capitoline with reasoning_effort: low (or effort: low on ask_council) skips the ranking stage and costs five calls; high, or no effort, is the full council, and the other levels resolve to the nearer of the two. A council configured without the ranking stage is pinned to it and reports reasoning_effort as ignored.

Model

What it convenes

Calls

capitoline

the reference panel: four families — Anthropic, OpenAI, Google, open weights — answer, rank each other blind, and a judge seated apart synthesizes. The shape to ask when the panel's own verdict on its answers is worth its price

9

capitoline-fast

the same four families and the same judge, without the ranking stage (ranking: false): four independent perspectives and a synthesis for half the price, and no panel verdict on them. It is capitoline at reasoning_effort: low, pinned under a name of its own for clients that cannot send the field. The response carries an empty rankings and aggregate and says which shape ran, so a fast deliberation is never read as one whose rankings all failed

5

The model list keeps itself current. Once a day, and at startup, the gateway asks the Codex and Antigravity CLIs which models they serve (codex debug models, agy models): a model they add is served under the door's prefix, one they drop disappears from /v1/models and the councils step past it, and /health shows what changed. Claude's names are aliases that already follow the latest model. Nothing is edited on disk, and a change can optionally be announced as one plain-text POST — to an ntfy topic, say, or any endpoint that takes one (docs/deploy.md §7.2).

A single model is named <door>-<model>: the door is the CLI the request goes through (claude, codex, antigravity), the model is what that door calls it. antigravity-claude-opus is Claude Opus through Antigravity and claude-opus is the same lineage through Claude Code, on a different subscription with a different quota, which is the distinction the prefix exists to draw.

To find out which model is enough for a task, a capability ladder seats three models of one family — three sizes, or one model at three reasoning levels — that rank each other blind. It is a measuring instrument rather than a council to use, so none ships: docs/measure-a-model.md has three ready to add to a host's overlay, and the measurement that showed the instrument works.

After the capitoline- prefix a shape word says how the council deliberates (-fast), a family name says who sits (-gemini, -claude, -openai for the ladders of that guide), and a number is a new version of the same shape (-2); capitoline alone stays the reference panel. A third council is configuration and a restart — other seats, another judge, a different quorum or deadline — and only a change to the sequence of stages needs the engine (spec §12.9).

MCP: POST /mcp (streamable HTTP) with tools list_models, ask_model, ask_council and generate_image (the image comes back as an MCP image content block). Registration from Claude Code is in docs/deploy.md §10. Raise the tool timeout on the client side first — export MCP_TOOL_TIMEOUT=1200000 in the shell that starts Claude Code: a CLI answer can take minutes, an image 11-45 s, and a deliberation has no deadline of its own beyond the 300 s each member of each stage gets, so it can run past 900 s — all well beyond the default. Capitoline sends a progress notification at every council stage and every 5 s while it draws, to clients that ask for one (a request with a progress token), but a notification postpones the deadline only in a client that sets resetTimeoutOnProgress (off by default in the MCP TypeScript SDK), so the raised timeout is what actually carries the call (spec §6.2).

Token counts — the usage block of every chat response, the usage of the last chunk of a stream, the usage of the MCP ask_model result, and the per-caller sums in GET /v1/usage — follow each provider's own convention: OpenAI counts cached and reasoning tokens inside the prompt and completion totals, Anthropic reports cached reads beside the input, Antigravity leaves them out of its own total and keeps its thinking tokens inside the output (measured, docs/spike-2026-09.md §10). Each adapter is right for the CLI it reads, so one call's count is comparable with the same provider's own history and not comparable with another provider's — and the sums in GET /v1/usage, which add up whatever providers a caller used, are an order of magnitude rather than a measure (docs/deploy.md §9). For one figure across all three, count calls, not tokens.

After updating a CLI, run scripts/smoke.sh (see docs/update-clis.md); it needs curl and jq on the machine it runs from.

Providers considered and declined

Grok (xAI). Grok Build, xAI's official CLI, can serve as a provider: a subscription signs in on a headless host, and its non-interactive output carries the text, token usage and the real model id (docs/spike-2026-09.md §12). It was then measured as a fifth council member, on 2026-09-29, with Grok Build 1.0.41 and grok-4.7, against a rule written before the runs (docs/measurements/2026-09-29-grok-seat/). On the two open design questions of the six, the other members ranked its answer first, blind, and the syntheses built on it were more complete. Twice in five answers, though, it stated a precise detail that is not true — a kernel source comment that does not exist, a vSphere requirement described wrongly for the way it was used — and the council carried both into the final answer: the peer ranking rewards the best answer as a whole, and the judge is told to build on it. It was also the slowest member by minutes per stage, and lost one answer and three of five rankings to the stage timeout or a malformed reply. So there is no Grok provider. The question is reopened with a new run of the same measurement, not by argument, when a later model or CLI might change the result; the protocol and the questions are there to rerun.

Principles

  • The CLIs are used as processes, never their tokens. This is an architectural constraint, not just a policy.

  • The CLIs are agents: they run in an empty sandbox, with tools disabled, as a separate user that is the only one holding the credentials.

  • Standard protocol, declared subset: whatever cannot be honored is rejected explicitly, never silently degraded.

  • Everything that depends on the CLIs lives in configuration, because the CLIs change every month.

A note on terms of service

Capitoline uses consumer subscriptions through the providers' official clients, for personal use. What each provider allows in headless, automated mode must be checked provider by provider and can change. Never pool multiple accounts of the same provider: it violates explicit clauses.

docs/terms-of-service.md answers one question — can you use Capitoline with your accounts, and for what — from the owner's reading of the four providers' terms (Anthropic, OpenAI, Google and xAI), with the clauses that matter quoted verbatim and dated. In short: yes for your own use on your own plan; a business seat is still one person's; a service for colleagues needs credentials the company holds. It is a personal interpretation by the author, who is not a lawyer, and not legal advice: anyone installing Capitoline reads the terms that bind their own accounts. docs/deployment-policy.md is how this installation applies that reading.

How this was built

Capitoline was written with Claude Code: most of the code and the documents are Claude's, and every commit carries a Co-Authored-By line saying so. The design, the decisions — what to build, the constraints above, the council's strategy, which providers to trust with a seat — and the review of every change are the author's.

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    C
    maintenance
    An MCP server that enables users to query, compare, and synthesize responses from multiple local and cloud LLMs simultaneously using existing subscriptions. It provides tools for parallel model evaluation, consensus polling with an LLM-as-judge, and response synthesis across different model providers.
    8
    34 npm
    16
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    MCP server that wraps AI CLI tools — Claude Code, Antigravity CLI, and Codex CLI — so any MCP client can call them as tools.
    266 npm
    16
    MIT