Skip to main content
Glama

Review MCP Server

Get expert code reviews from multiple AI models integrated into Claude Code. Catches bugs, security issues, and design problems automatically.

Quick Start

git clone https://github.com/je4550/review-mcp.git
cd review-mcp
npm install
npm run build

Configure Claude Code - Add to ~/.config/claude-code/mcp.json:

{
  "mcpServers": {
    "review-mcp": {
      "command": "node",
      "args": ["/absolute/path/to/review-mcp/dist/index.js"]
    }
  }
}

Set up a reviewer CLI (at least one):

# Option 1: Codex CLI (recommended)
codex --version  # If you already have it

# Option 2: OpenAI CLI
npm install -g openai
# No API key needed if logged in with ChatGPT subscription
# Otherwise: export OPENAI_API_KEY="sk-..."

# Option 3: Gemini CLI
npm install -g @google/gemini-cli
# No API key needed if logged in with Google account
# Otherwise: export GOOGLE_API_KEY="..."

Restart Claude Code and you're ready!

Related MCP server: Gemini Collaboration MCP Server

Usage

Just ask Claude naturally:

"Review this authentication function"
"Get a second opinion on src/auth.ts"
"Check the payment processing code for security issues"
"Review all files in the api/ directory"

Claude will get reviews from Codex/Gemini, analyze them, and present comprehensive feedback.

What You Get

Real Results from Testing

100-line authentication service:

  • Found: 6 critical security issues

  • SQL injection (5 locations), hardcoded secrets, insecure random, missing JWT expiration

  • Time: ~5 seconds

Payment processing module:

  • Found: 5 issues (2 critical, 2 high, 1 medium)

  • Hardcoded API keys, SQL injection, missing transactions, floating point errors

  • Time: ~4 seconds

React component (90 lines):

  • Found: 5 bugs

  • Null pointer crash, XSS vulnerability, state mutation bugs, missing dependencies

  • Time: ~5 seconds

Utility functions:

  • Found: 4 security flaws

  • Weak password hashing, insecure tokens, insufficient sanitization

  • Time: ~4 seconds

Example Review

Your code:

function auth(user, pass) {
  if (user === "admin" && pass === "12345") {
    return true;
  }
  return false;
}

Codex review:

- High: auth hard-codes "admin" and "12345" (auth.js:2). Anyone with
  source access gains full access, credentials can't be rotated without
  redeploying, and password is stored in clear text.

- High: Plain string comparison leaks timing information (auth.js:2).
  An attacker can measure response times to infer correct characters;
  use constant-time comparison.

- Medium: No hashing or KDF applied to password before comparison.
  Even if you moved the secret out of source control, you'd still want
  to hash user-supplied passwords.

Next steps: Replace hardcoded credential with configurable secret store,
hash/verify using a KDF, add constant-time compare helper.

Claude's synthesis:

Both reviewers identified critical security issues. The hardcoded credentials and timing attacks need immediate attention. I also notice there's no rate limiting or audit logging. Let me help you fix these...

Features

Senior-level reviews - Catches security, bugs, performance issues ✅ Multiple perspectives - Get Codex + Gemini + Claude's analysis ✅ Auto-detection - Works with whichever CLIs you have installed ✅ Smart validation - Filters out code rewrites and unhelpful responses ✅ Fast - ~5 seconds per 100 lines of code ✅ Comprehensive - Reviews snippets, files, or entire directories ✅ Prioritized - Issues marked as Critical/High/Medium/Low

Available Tools

Tool

Use Case

check_cli_status

Check which review CLIs are installed

review_code

Review a code snippet directly

review_file

Review a specific file

review_directory

Review all code files in a directory

You don't need to remember these - Claude calls them automatically when you ask for reviews.

Supported Languages

.js .ts .jsx .tsx .py .rb .go .java .c .cpp .cs .php .swift .kt .rs

How It Works

  1. You write code and ask Claude for a review

  2. MCP server detects which CLIs are available (Codex/Gemini)

  3. Sends your code with a simple prompt: "You are a senior software engineer. Code review the changes and implementation. Don't change anything, just review."

  4. Reviewers analyze in parallel (5-minute timeout each)

  5. Validation filters out invalid responses (code rewrites, errors, off-topic)

  6. Claude receives feedback and adds its own expert analysis

  7. You get comprehensive results with multiple AI perspectives

Troubleshooting

"No review CLIs available"

  • Run "Check CLI status" in Claude Code

  • Install at least one: codex, openai, or gemini CLI

Reviews timing out

  • 5-minute timeout should be plenty

  • Check internet connection and API keys

API keys not working

# Note: API keys not needed if you're logged in with:
# - ChatGPT subscription (for OpenAI CLI)
# - Google account (for Gemini CLI)

# If you need to set API keys manually:
# Check if keys are set
echo $OPENAI_API_KEY
echo $GOOGLE_API_KEY

# Add to ~/.bashrc or ~/.zshrc
export OPENAI_API_KEY="sk-..."
export GOOGLE_API_KEY="..."

Performance

Based on real testing:

  • Speed: ~5 seconds per 100 lines

  • Accuracy: Zero false positives in testing

  • Coverage: Finds security, bugs, performance, design issues

  • Cost: ~$0.10 per review at GPT-4 rates

  • Tokens: ~3,000 per 100-line file

Architecture

You write code
     ↓
Claude Code asks for review
     ↓
Review MCP Server
     ├─→ Detects available CLIs
     ├─→ Sends code to Codex/Gemini (parallel)
     ├─→ Validates responses
     └─→ Returns formatted feedback
          ↓
Claude analyzes and synthesizes
     ↓
You get expert recommendations

Development

npm run watch  # Auto-rebuild on changes

License

MIT

Contributing

Issues and PRs welcome at https://github.com/je4550/review-mcp

Available Tools

4 tools
check_cli_statusA

Check which code review CLIs (Codex/OpenAI and Gemini) are installed and available. Use this before requesting reviews to see what's available.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It clearly indicates this is a read-only check operation without side effects, but doesn't specify what format the availability information will be returned in, whether there are authentication requirements, or what happens if no CLIs are found. The description adds basic behavioral context but lacks detail about the output format.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is perfectly concise with two focused sentences. The first sentence states the purpose, the second provides usage guidance. Every word earns its place, and the information is front-loaded with the core functionality stated immediately.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a zero-parameter tool with no annotations and no output schema, the description provides good context about what the tool does and when to use it. However, it doesn't describe the return format or what specific information will be provided about CLI availability, leaving some ambiguity about the tool's output.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters with 100% schema description coverage, so the baseline is 4. The description appropriately doesn't discuss parameters since none exist, focusing instead on the tool's purpose and usage context.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Check which code review CLIs are installed and available') and identifies the resources (Codex/OpenAI and Gemini CLIs). It distinguishes this tool from its siblings (review_code, review_directory, review_file) by focusing on CLI availability checking rather than performing reviews.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit usage guidance: 'Use this before requesting reviews to see what's available.' This clearly indicates when to use this tool (as a prerequisite check) versus when to use its sibling review tools, establishing a clear workflow relationship.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

review_codeA

Request a code review from Codex and Gemini CLIs. Provide code directly as a string. Returns feedback from both reviewers for Claude to consider.

ParametersJSON Schema
NameRequiredDescriptionDefault
codeYesThe code to review
contextNoAdditional context about the code (optional)
reviewersNoWhich reviewers to use (default: both)

TDQS

A3.9/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses that the tool requests reviews from specific CLIs (Codex and Gemini) and returns feedback from both reviewers, but lacks details on permissions, rate limits, error handling, or what 'feedback' entails. It adds some behavioral context but leaves gaps for a mutation-like operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately sized and front-loaded with key information in two concise sentences. Every sentence earns its place by stating the action, input method, and output purpose without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, and 3 parameters with full schema coverage, the description is moderately complete. It covers the basic operation and output intent but lacks details on behavioral traits like error cases or feedback format, which are important for a tool that interacts with external CLIs.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters. The description adds minimal value beyond the schema by implying the 'code' parameter is provided as a string and 'reviewers' defaults to 'both', but does not elaborate on parameter interactions or usage nuances.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose with specific verbs ('Request a code review') and resources ('from Codex and Gemini CLIs'), and distinguishes it from siblings by specifying it reviews code provided as a string (unlike review_directory or review_file which likely handle files/directories).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context for when to use this tool ('Provide code directly as a string'), but does not explicitly state when not to use it or name alternatives like review_directory or review_file. It implies usage for string-based code review without file system access.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

review_directoryB

Request a code review of all files in a directory from Codex and Gemini CLIs. Returns feedback from both reviewers for Claude to consider.

ParametersJSON Schema
NameRequiredDescriptionDefault
directoryYesPath to the directory to review
contextNoAdditional context about the code (optional)
reviewersNoWhich reviewers to use (default: both)

TDQS

B3.4/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden. It discloses that the tool requests reviews from external CLIs and returns feedback, but doesn't mention authentication needs, rate limits, error conditions, or what happens if the directory doesn't exist. For a tool interacting with external services, this leaves significant behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is appropriately concise with two sentences that efficiently convey purpose and outcome. It's front-loaded with the core functionality. The second sentence about Claude integration could be slightly more integrated, but overall it's well-structured with minimal waste.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 3 parameters with full schema coverage but no annotations or output schema, the description provides adequate purpose but lacks behavioral context for a tool that interacts with external CLIs. It doesn't explain the format or structure of the returned feedback, which is important since there's no output schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters thoroughly. The description adds no additional parameter semantics beyond what's in the schema. The baseline of 3 is appropriate when the schema does all the parameter documentation work.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the specific action ('Request a code review'), target resource ('all files in a directory'), and tools involved ('Codex and Gemini CLIs'). It distinguishes from sibling tools like 'review_file' (single file) and 'review_code' (unclear scope) by specifying directory-level review.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage for directory-level code reviews but doesn't explicitly state when to use this tool versus alternatives like 'review_file' or 'review_code'. It mentions 'for Claude to consider' which suggests integration context, but lacks clear when/when-not guidance or prerequisite conditions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

review_fileB

Request a code review of a specific file from Codex and Gemini CLIs. Returns feedback from both reviewers for Claude to consider.

ParametersJSON Schema
NameRequiredDescriptionDefault
filePathYesPath to the file to review
contextNoAdditional context about the code (optional)
reviewersNoWhich reviewers to use (default: both)

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It mentions that the tool returns feedback from reviewers for Claude to consider, but lacks details on permissions, rate limits, error handling, or what the feedback format entails. For a tool that interacts with external CLIs and returns results, this is a significant gap in transparency.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise with two sentences that efficiently state the tool's purpose and outcome. It's front-loaded with the main action and avoids unnecessary details, though it could be slightly more structured by explicitly separating purpose from usage context.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity (interacting with multiple CLIs, returning feedback), lack of annotations, and no output schema, the description is moderately complete. It covers the basic purpose and outcome but misses behavioral details like authentication, error cases, or feedback structure. It's adequate for a minimal understanding but has clear gaps for effective agent use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all parameters (filePath, context, reviewers) with descriptions. The description adds no additional meaning beyond what the schema provides, such as explaining parameter interactions or constraints. Baseline score of 3 is appropriate when the schema handles parameter documentation adequately.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Request a code review of a specific file from Codex and Gemini CLIs.' It specifies the verb ('request a code review'), resource ('specific file'), and reviewers ('Codex and Gemini CLIs'), but doesn't explicitly distinguish it from sibling tools like 'review_code' or 'review_directory', which likely have different scopes.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies usage by mentioning 'Returns feedback from both reviewers for Claude to consider,' suggesting it's intended for Claude to process feedback. However, it doesn't provide explicit guidance on when to use this tool versus alternatives like 'review_code' or 'review_directory,' nor does it specify prerequisites or exclusions, leaving the context somewhat vague.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections. Dates show when Glama detected each change.

  1. 4 tool updates
    • First observedcheck_cli_status
    • First observedreview_code
    • First observedreview_directory
    • First observedreview_file

TDQS

A3.8/5.0
Disambiguation4/5

The tools are mostly distinct with clear boundaries: check_cli_status verifies CLI availability, while review_code, review_directory, and review_file handle code reviews at different granularities. However, review_directory and review_file could be slightly confused as both target files, but their descriptions clarify the scope difference (directory vs. single file).

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern with snake_case: check_cli_status, review_code, review_directory, and review_file. This predictability makes it easy for an agent to understand and select tools without confusion.

Tool Count5/5

With 4 tools, the server is well-scoped for its purpose of code review CLI integration. Each tool earns its place by covering distinct aspects: checking availability and reviewing code at different levels (string, file, directory). This count is neither too thin nor excessive.

Completeness4/5

The tool surface covers the core workflows for code review using CLIs: checking availability and performing reviews. A minor gap exists in not providing tools to manage or configure the CLIs (e.g., setting API keys), but agents can work around this as the essential review operations are fully supported.

Maintenance

ActivityInactive
ResponsivenessSyncing

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    C
    quality
    F
    maintenance
    Connects AI assistants like Claude to the Codex CLI for code analysis, editing, and execution. Supports file references with @ syntax, sandboxed code execution with approval workflows, and structured code changes for automated refactoring and documentation.
    8
    198
    179
    MIT
  • A
    license
    B
    quality
    D
    maintenance
    Enables Claude to collaborate with Gemini for code reviews, second opinions, and iterative software development. It facilitates multi-step workflows including PRD creation and code generation through an AI orchestration framework.
    2
    18
    1
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Bridges Claude and OpenAI's Codex CLI for AI-powered code analysis, generation, and review, with support for session management, web search, and structured output.
    6
    765
    628
    ISC

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/je4550/review-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server