Skip to main content
Glama

Gauntlet

Make Claude's code run the gauntlet.

Codex, Gemini and GPT-OSS review what Claude Code writes. Every finding has to quote the line it's about, and Gauntlet checks that quote against your files before Claude reads it. It runs on the ChatGPT and Google plans you already pay for. No API keys.

CI Node 20+ MCP License: MIT


Why I built this

I do most of my coding with Claude Code, and I kept hitting the same two walls.

The first: asking Claude to review its own code mostly gets you Claude agreeing with itself. The bug it didn't see while writing, it doesn't see while reviewing.

The second came when I started asking other models instead. They found real bugs, and they also made some up. A confident paragraph about a race condition on line 84, and line 84 is a comment. I'd lose half an hour finding that out.

Gauntlet is what I ended up with. Claude still writes the code and is still the only thing allowed to touch my files. When a change matters, it sends a small packet (the objective, the files, the diff) to a model from a different family and gets a structured review back. Before Claude sees that review, Gauntlet looks up every quoted line in the real file and sorts the findings into three groups:

  • verified;

  • nothing to check against a file;

  • "could not be confirmed", meaning the quoted code isn't there.

Here's a real run, Gauntlet reviewing its own JSON parser with two families:

# Council (review): 2/2 voices answered

## Verdicts
- codex (gpt-5.6-terra): approve, confidence high
- gemini (gemini-3.8-flash-high): approve_with_changes, confidence high

Disagreement: approve vs approve_with_changes. Resolve it with evidence, not by averaging.

## gemini
[gemini / gemini-3.8-flash-high, advisory and read-only: tier standard, 108.4s,
 2/2 claims verified against the files]

## Findings verified against the files (2)
1. [medium] src/json.mjs:32: extractJson aborts on the first unclosed opening brace,
   failing to extract subsequent valid JSON objects.
   Fix: Replace `break;` with `continue;` ...
2. [low] src/json.mjs:28: extractJson returns arrays when the input is a standalone
   JSON array ...

Codex approved the file. Gemini found two real bugs, both quoted from the source, and both got fixed in the next commit. That's the whole idea in one screen.

Related MCP server: claude-consult-mcp

What's in it

  • An MCP server with ten tools: reviews, a diagnosis tool, plan critique, a security audit, edge cases, and council, which asks Codex, Gemini and GPT-OSS the same question at once and tells you where they agree. Findings that two families reached on their own are the ones I trust most.

  • A fallback chain, because quotas run out. If the flagship model is rate-limited, the call goes to the smaller model of the same family, then to another family, then to Claude through Antigravity. A model that just failed sits out for a while, so the next call doesn't wait on it again.

  • Four read-only Claude subagents (haiku-navigator, sonnet-reviewer, opus-architect, and sonnet-test-analyst, which is opt-in) for when you'd rather stay in the Claude family.

  • Optional hooks:

    • a pre-turn hook that runs a small pipeline of internal agents (comprehension, anti-hallucination, text review, jury, legal, learning);

    • a gate that refuses git push while edited code hasn't had a review.

The pipeline config is versioned. Publish a new version and chats that are already open pick it up on their next message, without a restart. It's the part I'm proudest of, and the one I'd point a curious engineer at first. See architecture.

Numbers

These are from my machine (Windows 11, personal ChatGPT and Google plans, October 2026), so take them as one data point. gauntlet stats will give you yours.

What

Result

quick_check on a one-file question (Gemini, light tier)

25 s; about 31k tokens on Google's side, none on Claude's for the review itself

Same question again, file unchanged

answered from cache in 0.0 s

Two-family council on that file

108 s; two real bugs that one family missed

A 44,500-character document through the text-review agent

32 s; it caught the typo in the last sentence

Regression suite for the six internal agents

5/5 on live models

Same suite with Codex broken, then Codex and Gemini both broken

5/5, answered by Gemini, then by Claude

A word on tokens. Gauntlet doesn't make reviews free. It moves them off your Claude plan and keeps the packets small. docs/token-savings.md explains where the savings come from and how to measure them on your own work. I'd rather you measure than take my word for it.

Installing

1. What you need

  • Node.js 20 or newer (node --version).

  • Claude Code, with the claude command working in your terminal.

  • At least one of the two provider CLIs below. You don't need both; Gauntlet uses whatever is installed.

2. The provider CLIs

Codex, using your ChatGPT plan:

npm install -g @openai/codex
codex login        # sign in with your ChatGPT account

Antigravity, using your Google account. This is the one that gives you Gemini, GPT-OSS and the Claude fallback lane. Install Google's Antigravity CLI (agy) by following Google's instructions for your OS, then run it once so it can sign you in:

agy

3. Gauntlet itself

git clone https://github.com/Raffymimii/gauntlet.git
cd gauntlet
npm install
node bin/gauntlet.mjs init --dry-run    # prints every change, touches nothing
node bin/gauntlet.mjs init

What init does by default:

  • registers the MCP server with claude mcp add --scope user gauntlet ...;

  • copies the read-only subagents into ~/.claude/agents.

Without a flag it doesn't touch your settings.json or your CLAUDE.md. These are the opt-ins:

node bin/gauntlet.mjs init --claude-md        # adds one @-include line to ~/.claude/CLAUDE.md
node bin/gauntlet.mjs init --hooks turn       # pre-turn hook: live config + internal agents
node bin/gauntlet.mjs init --hooks review-gate
node bin/gauntlet.mjs init --test-analyst     # the subagent that runs your tests

I'd recommend --claude-md. It's the file that tells Claude when a quick check is enough and when to call the council. Without it, Claude only uses the tools when you ask.

If you want a plain gauntlet command instead of node bin/gauntlet.mjs, run npm link inside the folder.

4. Check it

node bin/gauntlet.mjs status

You should see your CLIs as installed and signed in, the model for each tier, and no paused models. Then open a new Claude Code chat (MCP servers are only loaded when a chat starts) and try:

get a quick_check of src/whatever.ts: is the retry loop bounded?

run a council on this diff before I push it

5. Removing it

node bin/gauntlet.mjs uninstall

This removes the MCP server, the subagents it installed, its hook entries and its CLAUDE.md line, and nothing else. Your data in ~/.gauntlet stays until you delete it.

When something doesn't work

Symptom

Likely cause

Claude doesn't see the gauntlet tools

You're in a chat opened before init. Open a new one. claude mcp list should show gauntlet.

status says a CLI is not signed in

Run codex login, or run agy once interactively.

"model not available" errors

Providers rename models. Put the current names in ~/.gauntlet/config.json (see configuration).

A review times out

The packet is too big, or the tier too high. Name fewer files, or ask for tier: "light".

A model is "paused"

It failed on quota or login recently. It comes back by itself. status shows how long.

agy isn't found by the MCP server

Set providers.antigravity.command to the full path of the executable.

Safety

The reviewers are read-only:

  • Codex runs in its read-only sandbox, and Antigravity in plan mode with its sandbox on;

  • both have a hard timeout and a scrubbed environment;

  • the prompt goes in on stdin, never on the command line;

  • a reviewer can't call another reviewer.

What leaves your machine is the packet, and only the files you name. Gauntlet refuses:

  • .env files and keys;

  • anything under .ssh, .aws or .git;

  • symlinks that lead outside the project;

  • binaries.

It also redacts token-shaped strings. gauntlet packet shows you exactly what would be sent, without sending it.

What you send to OpenAI or Google is covered by your own agreement with them, and checking that your use fits your plans' terms is on you. Gauntlet uses the official CLIs under your account and backs off when a provider says you're out of quota.

The full picture is in docs/security.md, including what it doesn't protect against.

FAQ

Does it cost anything per token? No. It uses your existing logins through the official CLIs.

Can a reviewer change my code? No. Reviewers only answer; Claude decides what to apply.

Why not just ask Claude to review its own work? Sometimes that's fine. But another family catches different things. And a reviewer whose quotes get checked against the file can't send you after a bug that doesn't exist.

Will this work with my model names? The defaults are what I use. When yours differ, ~/.gauntlet/config.json overrides any of them.

Contributing

Issues and PRs welcome. If you touch a prompt in prompts/, run the agent regressions (node bin/gauntlet.mjs regress) and paste the result in the PR. If you measure the token savings on real work, I'd love to see the numbers, along with how you measured them.

License

MIT

Available Tools

10 tools
codex_diagnosecodex_diagnoseA

Root cause of a bug or failing test. Call it as soon as your first hypothesis fails; give it the symptom, the error output and the few files involved. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
diffNoUnified diff, when the question is about a change.
tierNolight: fast, cheap model for small ordinary questions. standard: normal reviews. deep: flagship models, for security, money, concurrency, production or a hard bug. auto: decided from the size and subject of the packet. Pick the cheapest tier that can do the job.
modelNoPin a model for the first attempt. Usually leave unset and pick a tier.
pathsNoProject-relative files to include. Keep the list short.
checksNoSpecific checks, e.g. "does the retry loop stop on a 429".
effortNoCodex reasoning effort override. Usually leave unset.
noCacheNoSkip the result cache.
workdirYesAbsolute path of the project. The specialist reads nothing above it and writes nothing at all.
objectiveYesWhat you want decided or found, in one or two sentences. This is the whole task the specialist sees.
timeoutMsNoTime budget per attempt.
constraintsNoRules or context to respect: framework version, invariants, what is out of scope.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. 'Read-only' discloses the safety profile, and the guidance on inputs adds context, but it omits cost/tier behavior, caching (noCache), timeouts, or what the response looks like — all left entirely to the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the purpose, then the trigger, then the required inputs and safety note. No wasted words.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 11-parameter tool with full schema coverage and no output schema, the description adequately covers purpose, trigger, and inputs. It leaves return-format and cost behavior to the schema, which is acceptable, though a note on the read-only/no-write guarantee relative to siblings would round it out.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 11 parameters are already documented in structured data. The description's mention of 'symptom, the error output and the few files involved' loosely maps to objective/paths but adds no syntax or detail beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific goal (find the root cause of a bug or failing test), which an agent can act on. However, it does not distinguish itself from plausible siblings like gemini_analyze, quick_check, or codex_review, so an agent must infer which analysis tool applies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Call it as soon as your first hypothesis fails' gives a concrete trigger condition for when to invoke it, and it specifies what to supply (symptom, error output, files). It stops short of naming alternatives or conditions to prefer a sibling instead.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

codex_reviewcodex_reviewA

Adversarial review of a diff or a few files by Codex. Use after any non-trivial change. Returns a verdict and ranked findings with file:line, each checked against the files. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
diffNoUnified diff, when the question is about a change.
tierNolight: fast, cheap model for small ordinary questions. standard: normal reviews. deep: flagship models, for security, money, concurrency, production or a hard bug. auto: decided from the size and subject of the packet. Pick the cheapest tier that can do the job.
modelNoPin a model for the first attempt. Usually leave unset and pick a tier.
pathsNoProject-relative files to include. Keep the list short.
checksNoSpecific checks, e.g. "does the retry loop stop on a 429".
effortNoCodex reasoning effort override. Usually leave unset.
noCacheNoSkip the result cache.
workdirYesAbsolute path of the project. The specialist reads nothing above it and writes nothing at all.
objectiveYesWhat you want decided or found, in one or two sentences. This is the whole task the specialist sees.
timeoutMsNoTime budget per attempt.
constraintsNoRules or context to respect: framework version, invariants, what is out of scope.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden; it does disclose the safety-relevant "Read-only" and the return shape (verdict plus ranked file:line findings validated against the files). However, it omits other meaningful traits: that results are cached by default (evidenced only by the noCache param), the cost/latency implications of tier selection, and any timeout behavior. Adequate but materially incomplete for an 11-param tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four short sentences, front-loaded with purpose, then trigger, then return contract, then safety. No filler and each clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, yet the description compensates by naming the return contract (verdict plus ranked findings with file:line). It is nearly complete, with the residual gap being differentiation from the review siblings rather than any missing call mechanics.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all eleven parameters including the tier semantics. The description's "a diff or a few files" loosely maps to diff/paths but adds no syntax, format or constraint detail beyond the schema; baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (adversarial review of a diff or files) and names the engine (Codex), which is the exact axis that separates it from the sibling gemini_review. The scope is narrow and immediately understandable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Use after any non-trivial change" gives a trigger condition, but the sibling set contains gemini_review, quick_check, security_audit and edge_cases, and the description never says when to pick this over those. It implies usage rather than routing the agent.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

councilcouncilA

The same question to Codex, Gemini and GPT-OSS in parallel, merged: each verdict, the findings two or more families agree on (the strongest signal), and where they disagree. Use for risky changes, security-sensitive code or a design you are unsure of. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
diffNoUnified diff, when the question is about a change.
modeNoDefault review.
planNoFor mode=critique: the plan.
tierNolight: fast, cheap model for small ordinary questions. standard: normal reviews. deep: flagship models, for security, money, concurrency, production or a hard bug. auto: decided from the size and subject of the packet. Pick the cheapest tier that can do the job.
lanesNoWhich families sit on the council. Default: all three.
modelNoPin a model for the first attempt. Usually leave unset and pick a tier.
pathsNoProject-relative files to include. Keep the list short.
checksNoSpecific checks, e.g. "does the retry loop stop on a 429".
effortNoCodex reasoning effort override. Usually leave unset.
noCacheNoSkip the result cache.
workdirYesAbsolute path of the project. The specialist reads nothing above it and writes nothing at all.
objectiveYesWhat you want decided or found, in one or two sentences. This is the whole task the specialist sees.
timeoutMsNoTime budget per attempt.
constraintsNoRules or context to respect: framework version, invariants, what is out of scope.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses 'Read-only' and describes the merged output shape (each verdict, agreements, disagreements), but says nothing about cost, latency, or auth implications of fanning out to three model families. Adequate but with real gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Compact and mostly front-loaded: the mechanism and merge structure come first, then the use contexts, then the read-only note. No wasted sentences, though the opening could lead with the outcome rather than the process.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 14-parameter tool with no output schema, the description covers the core purpose and describes what comes back (verdicts, agreement findings, disagreements). The tier/cost and path guidance lives in the schema. The main omission is any note on the expense or latency of a multi-family call.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 14 parameters in detail. The description adds no parameter-level syntax or format beyond mentioning the three families, which the lanes enum already conveys. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific mechanism: the same question dispatched to Codex, Gemini and GPT-OSS in parallel and merged into a consensus verdict. This distinguishes it from single-model siblings like codex_review and gemini_review without naming them explicitly. Clear but not maximally differentiated.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives clear use contexts: risky changes, security-sensitive code, or a design you are unsure of. However, it offers no when-not guidance and does not route the agent away from overlapping siblings such as security_audit or plan_critique, so the alternative-selection guidance is incomplete.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

edge_casesedge_casesA

Inputs and states the code mishandles, phrased as test cases. Use after writing logic with branches, parsing, dates, money or retries. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
diffNoUnified diff, when the question is about a change.
tierNolight: fast, cheap model for small ordinary questions. standard: normal reviews. deep: flagship models, for security, money, concurrency, production or a hard bug. auto: decided from the size and subject of the packet. Pick the cheapest tier that can do the job.
modelNoPin a model for the first attempt. Usually leave unset and pick a tier.
pathsNoProject-relative files to include. Keep the list short.
checksNoSpecific checks, e.g. "does the retry loop stop on a 429".
effortNoCodex reasoning effort override. Usually leave unset.
noCacheNoSkip the result cache.
workdirYesAbsolute path of the project. The specialist reads nothing above it and writes nothing at all.
objectiveYesWhat you want decided or found, in one or two sentences. This is the whole task the specialist sees.
timeoutMsNoTime budget per attempt.
constraintsNoRules or context to respect: framework version, invariants, what is out of scope.

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden; 'Read-only' is a genuinely useful disclosure of side-effect profile. However it says nothing about cost/latency across the four tiers, whether code is ever executed, or how results are staged, all relevant for an 11-parameter analysis tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero waste: purpose first, usage trigger second, safety note third. Nothing padded, and the most decision-relevant information is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 11-parameter tool with no output schema and no annotations, the description covers identity and one usage trigger but omits result format, cost/latency expectations, and how it relates to the other review tools. Adequate, not complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, including detailed tier and objective semantics, so the baseline is 3. The description adds no parameter-level meaning beyond what the schema already documents.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Strong specific verb+resource: it enumerates inputs and states the code mishandles and delivers them as test cases. That is a distinct niche from the sibling reviewers, though the description never names or contrasts any sibling explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete trigger: use after writing logic involving branches, parsing, dates, money or retries. Clear context for when the tool applies, but no when-not guidance and no named alternative among quick_check/codex_review for overlapping cases.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gauntlet_statusgauntlet_statusA

Which CLIs are installed and signed in, which model each tier uses, which models are paused after failures, and the limits in force. Call it after a failed call. Never shows credentials.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden, and it does disclose two meaningful traits: that failing models get paused (so the report reflects evolving state) and that credentials are never returned. It still does not state whether the tool is read-only or requires any auth, but for a zero-parameter status probe the disclosure is substantive.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, front-loaded with the report contents and closed with the caller condition and the credential guarantee. No filler and every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description must convey what comes back, and its four-part enumeration does that adequately for a no-argument status tool. It could be tighter on when the paused/limit state applies, but an agent has enough to call it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate and the baseline of 4 applies. The enumeration of report contents is a bonus rather than a compensation for missing parameter documentation.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description enumerates precisely what the tool surfaces: installed/signed-in CLIs, per-tier model assignment, paused models, and active limits. That is a concrete, specific purpose rather than a restatement of the name. It stops short of 5 because there is no explicit verb and no differentiation from the nearest sibling, quick_check, which an agent could plausibly confuse it with.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Call it after a failed call' gives a clear triggering condition rather than leaving usage to inference. It does not name an alternative or state when not to call it (e.g., quick_check for a cheap pre-flight), so it falls short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gemini_analyzegemini_analyzeB

How a broad or unfamiliar part of the code fits together. Use before editing code you have not read. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
diffNoUnified diff, when the question is about a change.
tierNolight: fast, cheap model for small ordinary questions. standard: normal reviews. deep: flagship models, for security, money, concurrency, production or a hard bug. auto: decided from the size and subject of the packet. Pick the cheapest tier that can do the job.
modelNoPin a model for the first attempt. Usually leave unset and pick a tier.
pathsNoProject-relative files to include. Keep the list short.
checksNoSpecific checks, e.g. "does the retry loop stop on a 429".
effortNoCodex reasoning effort override. Usually leave unset.
noCacheNoSkip the result cache.
workdirYesAbsolute path of the project. The specialist reads nothing above it and writes nothing at all.
objectiveYesWhat you want decided or found, in one or two sentences. This is the whole task the specialist sees.
timeoutMsNoTime budget per attempt.
constraintsNoRules or context to respect: framework version, invariants, what is out of scope.

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full disclosure burden, and it does supply the single most important trait: 'Read-only.' That is a genuine, useful behavioral commitment. However, it says nothing about the fact that this dispatches work to an external specialist model, the cost/latency implications, or caching behavior (params like tier/model/noCache hint at this but the description is silent).

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three terse fragments with no filler, and the usage trigger and safety note are packed into minimal space. The opening phrase is a noun clause rather than a front-loaded verb, which slightly weakens scannability, but overall it is efficiently sized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

This is a complex 11-parameter tool with no output schema and no annotations, so the description must carry more weight than it does. It omits what the tool returns, that it delegates to an external LLM specialist (with cost implications), and how it relates to the sibling analysis/review tools, leaving meaningful gaps for an agent deciding whether and how to call it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% across 11 parameters, so the schema already explains diff, tier, model, paths, checks, workdir, objective, timeoutMs and constraints thoroughly (e.g. the tier enum guidance). The description adds no additional parameter meaning beyond what the schema provides, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose3/5

Does the description clearly state what the tool does and how it differs from similar tools?

The fragment 'How a broad or unfamiliar part of the code fits together' implies a code-comprehension/analysis function, but it never uses a verb+resource construction to state what the tool actually does. It also offers no differentiation from the many sibling review/diagnose tools (gemini_review, codex_review, codex_diagnose, edge_cases), so an agent must infer the boundary from the name alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The trigger 'Use before editing code you have not read' is a concrete, actionable condition that tells the agent when this tool applies. It stops short of naming alternatives or stating when-not-to-use (e.g. versus gemini_review), so it is clear context without exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

gemini_reviewgemini_reviewB

Second review by Gemini, strongest on interface contracts, error paths, state and user-facing regressions. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
diffNoUnified diff, when the question is about a change.
tierNolight: fast, cheap model for small ordinary questions. standard: normal reviews. deep: flagship models, for security, money, concurrency, production or a hard bug. auto: decided from the size and subject of the packet. Pick the cheapest tier that can do the job.
modelNoPin a model for the first attempt. Usually leave unset and pick a tier.
pathsNoProject-relative files to include. Keep the list short.
checksNoSpecific checks, e.g. "does the retry loop stop on a 429".
effortNoCodex reasoning effort override. Usually leave unset.
noCacheNoSkip the result cache.
workdirYesAbsolute path of the project. The specialist reads nothing above it and writes nothing at all.
objectiveYesWhat you want decided or found, in one or two sentences. This is the whole task the specialist sees.
timeoutMsNoTime budget per attempt.
constraintsNoRules or context to respect: framework version, invariants, what is out of scope.

TDQS

B3.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full behavioral burden. It discloses read-only status and its focus areas, and the schema adds that the workdir is read-only and nothing above it is read, but the description omits auth needs, cost/rate implications, and what the review returns. Adequate but thin for an annotation-free tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short, front-loaded sentences with no filler; the purpose leads and the read-only trait is stated compactly. Terse, though it trades some completeness for brevity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 11-parameter review tool with no output schema and no annotations, the description covers the role and focus but leaves behavioral details (auth, cost, return format) to the schema or unstated. Reasonable but with clear gaps.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter (tier, model, diff, paths, checks, objective, constraints, etc.) is already documented in the schema. The description adds nothing parameter-specific, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States that this is a review by Gemini and names its strong areas (interface contracts, error paths, state, user-facing regressions), which gives a real sense of what it does. However, it never explicitly states the verb+resource ('reviews your code/diff') nor names the sibling it complements (codex_review, gemini_analyze), so differentiation is only implicit.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'Second review' implies it is used as a follow-up pass, and the listed strengths hint at when to prefer it, but there is no explicit when-to-use/when-not or named alternative such as codex_review or gemini_analyze. Usage is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

plan_critiqueplan_critiqueA

Critique of an implementation plan before any code is written: wrong assumptions, missing steps, simpler paths, production risks. Use on any task that touches more than one file. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
diffNoUnified diff, when the question is about a change.
planYesThe plan: steps, files to touch, approach, assumptions.
tierNolight: fast, cheap model for small ordinary questions. standard: normal reviews. deep: flagship models, for security, money, concurrency, production or a hard bug. auto: decided from the size and subject of the packet. Pick the cheapest tier that can do the job.
modelNoPin a model for the first attempt. Usually leave unset and pick a tier.
pathsNoProject-relative files to include. Keep the list short.
checksNoSpecific checks, e.g. "does the retry loop stop on a 429".
effortNoCodex reasoning effort override. Usually leave unset.
noCacheNoSkip the result cache.
workdirYesAbsolute path of the project. The specialist reads nothing above it and writes nothing at all.
objectiveYesWhat you want decided or found, in one or two sentences. This is the whole task the specialist sees.
timeoutMsNoTime budget per attempt.
constraintsNoRules or context to respect: framework version, invariants, what is out of scope.

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the behavioral burden, and it does disclose 'Read-only' and that it operates before code is written. However it omits cost/latency behavior, caching semantics, and what the critique actually returns; the cost tradeoffs are relegated entirely to the tier parameter's schema text rather than the description itself.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with zero waste: the scope of the critique is front-loaded, followed immediately by the usage trigger and a one-word safety disclosure. Every fragment earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 12-parameter tool with full schema coverage and no annotations, the description supplies purpose, scope, usage condition, and the critical read-only guarantee. The notable gap is that, with no output schema, it never hints at what the critique returns (format, structure, actionable findings).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 12 parameters (tier, workdir, effort, etc.). The description adds no parameter-level detail beyond what the schema provides, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (critique) and resource (an implementation plan) and enumerates the kinds of issues it surfaces (wrong assumptions, missing steps, simpler paths, production risks). The phrase 'before any code is written' implicitly distinguishes it from the review siblings, though no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a concrete trigger: 'Use on any task that touches more than one file.' That is a clear condition for invocation. It stops short of naming alternatives (e.g. codex_review for written code) or stating when NOT to use it, so it is not a full 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

quick_checkquick_checkA

Fast second look at something small: a function, a regex, a query, a config value, a one-file change. Cheap and quick; use it often. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
diffNoUnified diff, when the question is about a change.
tierNolight: fast, cheap model for small ordinary questions. standard: normal reviews. deep: flagship models, for security, money, concurrency, production or a hard bug. auto: decided from the size and subject of the packet. Pick the cheapest tier that can do the job.
modelNoPin a model for the first attempt. Usually leave unset and pick a tier.
pathsNoProject-relative files to include. Keep the list short.
checksNoSpecific checks, e.g. "does the retry loop stop on a 429".
effortNoCodex reasoning effort override. Usually leave unset.
noCacheNoSkip the result cache.
workdirYesAbsolute path of the project. The specialist reads nothing above it and writes nothing at all.
objectiveYesWhat you want decided or found, in one or two sentences. This is the whole task the specialist sees.
timeoutMsNoTime budget per attempt.
constraintsNoRules or context to respect: framework version, invariants, what is out of scope.

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It does disclose the key behavioral trait ("Read-only"), which is valuable, but says nothing about what is returned, that results are cached, or the cost/latency implications of the tier system — significant gaps for an 11-param tool with no annotation coverage.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the scope, with no filler. Every clause carries information an agent can act on.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 11 parameters, no annotations, and no output schema, the description is thin relative to the tool's complexity — it never says what the specialist returns or in what form. The 100% schema coverage compensates for parameter gaps but not for the missing behavioral/return context.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter including tier, workdir, and checks is already documented in the schema. The description's artifact list loosely maps onto paths/diff but adds no syntax or selection guidance beyond the schema, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a concrete purpose: a fast, cheap "second look" scoped to small artifacts (function, regex, query, config value, one-file change). This contrasts implicitly with heavier siblings like codex_review or security_audit, but no sibling is named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Cheap and quick; use it often" is a genuine usage directive, and the enumerated small-artifact list bounds when it applies. However, it never says when NOT to use it or which alternative to pick for larger changes, leaving the routing decision to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

security_auditsecurity_auditA

Exploitable issues: injection, authorization gaps, secrets, SSRF, path traversal, races, trust of client input. Use on auth, payments, user input, file paths, shell commands. Read-only.

ParametersJSON Schema
NameRequiredDescriptionDefault
diffNoUnified diff, when the question is about a change.
tierNolight: fast, cheap model for small ordinary questions. standard: normal reviews. deep: flagship models, for security, money, concurrency, production or a hard bug. auto: decided from the size and subject of the packet. Pick the cheapest tier that can do the job.
modelNoPin a model for the first attempt. Usually leave unset and pick a tier.
pathsNoProject-relative files to include. Keep the list short.
checksNoSpecific checks, e.g. "does the retry loop stop on a 429".
effortNoCodex reasoning effort override. Usually leave unset.
noCacheNoSkip the result cache.
workdirYesAbsolute path of the project. The specialist reads nothing above it and writes nothing at all.
objectiveYesWhat you want decided or found, in one or two sentences. This is the whole task the specialist sees.
timeoutMsNoTime budget per attempt.
constraintsNoRules or context to respect: framework version, invariants, what is out of scope.

TDQS

A3.6/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses one important trait, 'Read-only', but says nothing about cost/model spend, caching, per-attempt time budget, or the shape of the result. The schema's workdir description also states it reads nothing above the root and writes nothing, so 'Read-only' is partly redundant with structured data; the description adds a little but leaves the behavioral picture thin for an 11-parameter tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short telegraphic sentences: what it finds, where to use it, and the read-only guarantee, with the issue classes front-loaded. There is no filler. The fragmented list style trades a little readability for density, which keeps it just under a 5.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex 11-parameter tool with 100% schema coverage and no output schema, the description covers purpose and when-to-use but omits anything about how it differs operationally from codex_review/gemini_review or what the agent gets back. Given the schema covers parameters fully and no output schema exists, the remaining gap is sibling differentiation, so an adequate-but-not-complete 3.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 11 parameters thoroughly (tier semantics, workdir sandboxing, objective, checks). The description adds no syntax, format, or defaults beyond that. Baseline 3 is correct when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

It states a specific domain: finding exploitable security issues, and enumerates the classes it covers (injection, authz gaps, secrets, SSRF, path traversal, races, client-input trust). That is enough to distinguish it from the generic review siblings (codex_review, gemini_review) and the structural siblings (edge_cases, plan_critique) without naming them. It is a noun-phrase list rather than a clean verb+resource, but the purpose is unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives concrete triggers: use on auth, payments, user input, file paths, shell commands. That is a clear context-of-use. It stops short of explicit exclusions or naming the sibling to pick instead when the question is not security-relevant, so it is not a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 10 tool updatesv0.1.0
    • First observedcodex_diagnose
    • First observedcodex_review
    • First observedcouncil
    • First observededge_cases
    • First observedgauntlet_status
    • First observedgemini_analyze
    • First observedgemini_review
    • First observedplan_critique
    • First observedquick_check
    • First observedsecurity_audit

TDQS

A3.6/5.0

Scored across 10 tools

Disambiguation4/5

Most tools have clearly distinct purposes, such as status, quick check, security audit, edge cases, and plan critique. However, codex_review, gemini_review, and council overlap in the review space, and an agent might be unsure when to use a single-model review versus the multi-model council, despite the descriptions offering guidance.

Naming Consistency3/5

All names use snake_case, which is consistent, but the underlying pattern is mixed: some are model_action (codex_review, gemini_analyze), some are noun_noun (security_audit, edge_cases, plan_critique), and others are adjective_noun (quick_check) or a single noun (council). This is readable but not a predictable verb_noun convention throughout.

Tool Count5/5

The server provides 10 tools, which fits comfortably within the ideal range of 3–15 for a focused code-review and analysis assistant. Each tool appears to earn its place by covering a specific analysis angle or model combination.

Completeness4/5

The surface covers a broad range of code quality concerns: status, quick checks, reviews, diagnosis, architecture analysis, security, edge cases, plan critique, and multi-model council. Minor gaps exist, such as no explicit tool for performance analysis or dependency auditing, but the core lifecycle for reviewing and understanding code is well represented.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers