mcp-multi-agent-code-review
by Jk180603
README.md
# Multi-Agent Code Review MCP Server
An **MCP (Model Context Protocol) server** that runs a **multi-agent code review
pipeline** built with **LangGraph**. Point any MCP host (Claude Desktop, Cursor,
Claude Code) at it, hand it a git diff, and a graph of specialized LLM agents
reviews the change for security, quality, and logic issues, then a supervisor
agent aggregates their findings into a single pass/fail verdict.
This is not an API wrapper. It is a stateful multi-agent graph behind a
production-style guardrail stack.


## Why this exists
Most "AI code review" demos are a single prompt that says "review this code."
Real review is adversarial and multi-perspective: a security reviewer, a quality
reviewer, and a correctness reviewer each look for different things, and someone
senior reconciles them. This project models exactly that as a LangGraph state
graph, and exposes it over MCP so it plugs into the tools developers already use.
## Architecture
```
validate + rate-limit + audit (deterministic guardrails)
│
▼
prepare (redact secrets before any LLM call)
│
┌───────────────┼───────────────┐
▼ ▼ ▼
security quality logic (3 agents, run concurrently)
└───────────────┼───────────────┘
▼
supervisor (aggregate, gate on severity)
│
▼
pass / fail verdict
```
- **LangGraph** manages shared state and the fan-out / fan-in of the three
reviewer agents.
- Each agent returns **structured output** (a Pydantic `AgentReport`), so
malformed model output is rejected at the boundary, never propagated.
- The **supervisor** merges agent findings with a **deterministic secret scan**
and applies a configurable severity gate.
## Industry-level features (not a student toy)
- **Input validation** — size limits, encoding checks, file-count guards.
- **Deterministic secret pre-scan** — regex catch for AWS keys, private keys,
bearer/Slack tokens, API-key assignments. Runs *before* any LLM call.
- **Secret redaction** — matched secrets are redacted before the diff is ever
sent to an external model.
- **Rate limiting** — in-memory token bucket per configured window.
- **Audit logging** — every call logged with a SHA-256 of the input (never the
raw source), so the audit trail is not a secondary leak.
- **Config-driven severity gate** — a review fails only if a finding meets the
configured threshold; fully env-overridable via `pydantic-settings`.
- **Graceful degradation** — if one agent's model output is unparseable, that
agent is skipped, not the whole review.
- **Tested** — the deterministic core (security + supervisor logic) has a full
pytest suite that runs with no API key.
## Tech stack
`Python` · `MCP (FastMCP)` · `LangGraph` · `LangChain` · `Pydantic` · `pytest`
## Setup
```bash
pip install -r requirements.txt
export CODEREVIEW_LLM_PROVIDER=anthropic
export ANTHROPIC_API_KEY=... # or OPENAI_API_KEY with LLM_PROVIDER=openai
```
## Run the tests (no API key needed)
```bash
pytest
```
## Run the server
```bash
python -m src.server
```
## Register with Claude Code / Cursor
```bash
claude mcp add code-review -- python -m src.server
```
Then in your MCP host, call the `review_diff` tool with a unified git diff.
## Configuration
All settings are environment variables prefixed `CODEREVIEW_`:
| Variable | Default | Meaning |
| --------------------------------- | ------------------- | -------------------------------- |
| `CODEREVIEW_LLM_PROVIDER` | `anthropic` | `anthropic` or `openai` |
| `CODEREVIEW_FAIL_ON_SEVERITY` | `high` | Gate: fail at/above this level |
| `CODEREVIEW_MAX_DIFF_BYTES` | `100000` | Reject larger diffs |
| `CODEREVIEW_RATE_LIMIT_PER_MINUTE`| `20` | Max reviews/minute |
## Note on MCP versioning
`requirements.txt` pins `mcp>=1.28,<2` on purpose: MCP SDK v1.x is the stable
production line, and v2 is a breaking rework. Pinning below v2 is the
recommended practice for anything depending on `mcp`.
This server cannot be deployed
Maintenance
ActivitySlowing
ResponsivenessNo issues