Skip to main content
Glama
README.md
<div align="center">

# Gauntlet

Make Claude's code run the gauntlet.

Codex, Gemini and GPT-OSS review what Claude Code writes. Every finding has to quote the
line it's about, and Gauntlet checks that quote against your files before Claude reads it.
It runs on the ChatGPT and Google plans you already pay for. No API keys.

[![CI](https://github.com/Raffymimii/gauntlet/actions/workflows/ci.yml/badge.svg)](https://github.com/Raffymimii/gauntlet/actions/workflows/ci.yml)
![Node 20+](https://img.shields.io/badge/node-20%2B-339933)
![MCP](https://img.shields.io/badge/MCP-server-6E56CF)
![License: MIT](https://img.shields.io/badge/license-MIT-blue)

</div>

---

## Why I built this

I do most of my coding with Claude Code, and I kept hitting the same two walls.

The first: asking Claude to review its own code mostly gets you Claude agreeing with
itself. The bug it didn't see while writing, it doesn't see while reviewing.

The second came when I started asking other models instead. They found real bugs, and
they also made some up. A confident paragraph about a race condition on line 84, and line
84 is a comment. I'd lose half an hour finding that out.

Gauntlet is what I ended up with. Claude still writes the code and is still the only thing
allowed to touch my files. When a change matters, it sends a small packet (the objective,
the files, the diff) to a model from a different family and gets a structured review back.
Before Claude sees that review, Gauntlet looks up every quoted line in the real file and
sorts the findings into three groups:
- verified;
- nothing to check against a file;
- "could not be confirmed", meaning the quoted code isn't there.

Here's a real run, Gauntlet reviewing its own JSON parser with two families:

```text
# Council (review): 2/2 voices answered

## Verdicts
- codex (gpt-5.6-terra): approve, confidence high
- gemini (gemini-3.8-flash-high): approve_with_changes, confidence high

Disagreement: approve vs approve_with_changes. Resolve it with evidence, not by averaging.

## gemini
[gemini / gemini-3.8-flash-high, advisory and read-only: tier standard, 108.4s,
 2/2 claims verified against the files]

## Findings verified against the files (2)
1. [medium] src/json.mjs:32: extractJson aborts on the first unclosed opening brace,
   failing to extract subsequent valid JSON objects.
   Fix: Replace `break;` with `continue;` ...
2. [low] src/json.mjs:28: extractJson returns arrays when the input is a standalone
   JSON array ...
```

Codex approved the file. Gemini found two real bugs, both quoted from the source, and both
got fixed in the next commit. That's the whole idea in one screen.

## What's in it

- An MCP server with ten tools: reviews, a diagnosis tool, plan critique, a security
  audit, edge cases, and `council`, which asks Codex, Gemini and GPT-OSS the same question
  at once and tells you where they agree. Findings that two families reached on their own
  are the ones I trust most.
- A fallback chain, because quotas run out. If the flagship model is rate-limited, the
  call goes to the smaller model of the same family, then to another family, then to
  Claude through Antigravity. A model that just failed sits out for a while, so the next
  call doesn't wait on it again.
- Four read-only Claude subagents (`haiku-navigator`, `sonnet-reviewer`, `opus-architect`,
  and `sonnet-test-analyst`, which is opt-in) for when you'd rather stay in the Claude
  family.
- Optional hooks:
  - a pre-turn hook that runs a small pipeline of internal agents (comprehension,
    anti-hallucination, text review, jury, legal, learning);
  - a gate that refuses `git push` while edited code hasn't had a review.

The pipeline config is versioned. Publish a new version and chats that are already open
pick it up on their next message, without a restart. It's the part I'm proudest of,
and the one I'd point a curious engineer at first. See
[architecture](docs/architecture.md).

## Numbers

These are from my machine (Windows 11, personal ChatGPT and Google plans, October 2026),
so take them as one data point. `gauntlet stats` will give you yours.

| What | Result |
|---|---|
| `quick_check` on a one-file question (Gemini, light tier) | 25 s; about 31k tokens on Google's side, none on Claude's for the review itself |
| Same question again, file unchanged | answered from cache in 0.0 s |
| Two-family council on that file | 108 s; two real bugs that one family missed |
| A 44,500-character document through the text-review agent | 32 s; it caught the typo in the last sentence |
| Regression suite for the six internal agents | 5/5 on live models |
| Same suite with Codex broken, then Codex and Gemini both broken | 5/5, answered by Gemini, then by Claude |

A word on tokens. Gauntlet doesn't make reviews free. It moves them off your Claude plan and
keeps the packets small. [docs/token-savings.md](docs/token-savings.md) explains where the
savings come from and how to measure them on your own work. I'd rather you measure than
take my word for it.

## Installing

### 1. What you need

- Node.js 20 or newer (`node --version`).
- [Claude Code](https://docs.claude.com/en/docs/claude-code), with the `claude` command
  working in your terminal.
- At least one of the two provider CLIs below. You don't need both; Gauntlet uses
  whatever is installed.

### 2. The provider CLIs

Codex, using your ChatGPT plan:

```bash
npm install -g @openai/codex
codex login        # sign in with your ChatGPT account
```

Antigravity, using your Google account. This is the one that gives you Gemini, GPT-OSS
and the Claude fallback lane. Install Google's Antigravity CLI (`agy`) by following
Google's instructions for your OS, then run it once so it can sign you in:

```bash
agy
```

### 3. Gauntlet itself

```bash
git clone https://github.com/Raffymimii/gauntlet.git
cd gauntlet
npm install
node bin/gauntlet.mjs init --dry-run    # prints every change, touches nothing
node bin/gauntlet.mjs init
```

What `init` does by default:
- registers the MCP server with `claude mcp add --scope user gauntlet ...`;
- copies the read-only subagents into `~/.claude/agents`.

Without a flag it doesn't touch your `settings.json` or your `CLAUDE.md`. These are the
opt-ins:

```bash
node bin/gauntlet.mjs init --claude-md        # adds one @-include line to ~/.claude/CLAUDE.md
node bin/gauntlet.mjs init --hooks turn       # pre-turn hook: live config + internal agents
node bin/gauntlet.mjs init --hooks review-gate
node bin/gauntlet.mjs init --test-analyst     # the subagent that runs your tests
```

I'd recommend `--claude-md`. It's the file that tells Claude when a quick check is enough
and when to call the council. Without it, Claude only uses the tools when you ask.

If you want a plain `gauntlet` command instead of `node bin/gauntlet.mjs`, run `npm link`
inside the folder.

### 4. Check it

```bash
node bin/gauntlet.mjs status
```

You should see your CLIs as installed and signed in, the model for each tier, and no
paused models. Then **open a new Claude Code chat** (MCP servers are only loaded when a chat
starts) and try:

> get a quick_check of src/whatever.ts: is the retry loop bounded?

> run a council on this diff before I push it

### 5. Removing it

```bash
node bin/gauntlet.mjs uninstall
```

This removes the MCP server, the subagents it installed, its hook entries and its
`CLAUDE.md` line, and nothing else. Your data in `~/.gauntlet` stays until you delete it.

## When something doesn't work

| Symptom | Likely cause |
|---|---|
| Claude doesn't see the `gauntlet` tools | You're in a chat opened before `init`. Open a new one. `claude mcp list` should show `gauntlet`. |
| `status` says a CLI is not signed in | Run `codex login`, or run `agy` once interactively. |
| "model not available" errors | Providers rename models. Put the current names in `~/.gauntlet/config.json` (see [configuration](docs/configuration.md)). |
| A review times out | The packet is too big, or the tier too high. Name fewer files, or ask for `tier: "light"`. |
| A model is "paused" | It failed on quota or login recently. It comes back by itself. `status` shows how long. |
| `agy` isn't found by the MCP server | Set `providers.antigravity.command` to the full path of the executable. |

## Safety

The reviewers are read-only:
- Codex runs in its read-only sandbox, and Antigravity in plan mode with its sandbox on;
- both have a hard timeout and a scrubbed environment;
- the prompt goes in on stdin, never on the command line;
- a reviewer can't call another reviewer.

What leaves your machine is the packet, and only the files you name. Gauntlet refuses:
- `.env` files and keys;
- anything under `.ssh`, `.aws` or `.git`;
- symlinks that lead outside the project;
- binaries.

It also redacts token-shaped strings. `gauntlet packet` shows you exactly what would be
sent, without sending it.

What you send to OpenAI or Google is covered by your own agreement with them, and checking
that your use fits your plans' terms is on you. Gauntlet uses the official CLIs under
your account and backs off when a provider says you're out of quota.

The full picture is in [docs/security.md](docs/security.md), including what it doesn't
protect against.

## FAQ

**Does it cost anything per token?** No. It uses your existing logins through the
official CLIs.

**Can a reviewer change my code?** No. Reviewers only answer; Claude decides what to apply.

**Why not just ask Claude to review its own work?** Sometimes that's fine. But another
family catches different things. And a reviewer whose quotes get checked against the file
can't send you after a bug that doesn't exist.

**Will this work with my model names?** The defaults are what I use. When yours differ,
`~/.gauntlet/config.json` overrides any of them.

## Contributing

Issues and PRs welcome. If you touch a prompt in `prompts/`, run the agent regressions
(`node bin/gauntlet.mjs regress`) and paste the result in the PR. If you measure the token
savings on real work, I'd love to see the numbers, along with how you measured them.

## License

MIT

TDQS

A3.6/5.0

Scored across 10 tools

Disambiguation4/5

Most tools have clearly distinct purposes, such as status, quick check, security audit, edge cases, and plan critique. However, codex_review, gemini_review, and council overlap in the review space, and an agent might be unsure when to use a single-model review versus the multi-model council, despite the descriptions offering guidance.

Naming Consistency3/5

All names use snake_case, which is consistent, but the underlying pattern is mixed: some are model_action (codex_review, gemini_analyze), some are noun_noun (security_audit, edge_cases, plan_critique), and others are adjective_noun (quick_check) or a single noun (council). This is readable but not a predictable verb_noun convention throughout.

Tool Count5/5

The server provides 10 tools, which fits comfortably within the ideal range of 3–15 for a focused code-review and analysis assistant. Each tool appears to earn its place by covering a specific analysis angle or model combination.

Completeness4/5

The surface covers a broad range of code quality concerns: status, quick checks, reviews, diagnosis, architecture analysis, security, edge cases, plan critique, and multi-model council. Minor gaps exist, such as no explicit tool for performance analysis or dependency auditing, but the core lifecycle for reviewing and understanding code is well represented.

Maintenance

ActivityMaintained
ResponsivenessNo issues