code-review-agent
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@code-review-agentReview this Python function for issues: def foo(a, b): return a / b"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Code Review Agent
A code review agent for AI-generated code. When Claude Code / Codex / Cursor writes code, who checks it before merge? This agent does — triple-engine review (rule engine + AST structural analysis + LLM semantic review) with cross-validation, catching the patterns AI coding tools most commonly get wrong: hallucinated imports, eval() injections, shell=True, swallowed exceptions, and more. Returns structured reports with per-dimension scores, deterministic metrics, SARIF export, and directly applicable fix code. Provides REST API and 10 MCP tools.
中文 | English
Submission for X-Agent AI MCP Hackathon 2026 · Open Innovation Challenge.
Why: AI-generated code needs a different kind of review
AI coding tools (Claude Code, Codex, Cursor, GitHub Copilot) are fast — but they repeat the same mistakes:
AI pattern | What happens | Rule that catches it |
Hallucinated imports |
|
|
| AI uses |
|
| AI builds shell strings instead of arg lists |
|
Swallowed exceptions |
|
|
| AI writes |
|
Hardcoded secrets | AI inlines API keys instead of using env vars |
|
This agent's rule engine includes 6 dedicated AI-pattern rules (AI-H001–AI-H006) that target these hallucination patterns. The triple-engine design means: rule engine catches deterministic patterns (ms-level, free), AST analyzer catches structural errors (undefined vars, duplicate defs), and LLM confirms/denies rule hits to reduce false positives — the cross-validation that a single-engine tool can't do.
Related MCP server: CodePeel MCP Server
Triple-Engine Architecture
┌─────────────────────────────────────────────────────────┐
│ Input: Code / Diff / Multi-file / PR URL │
└───────────────┬─────────────────────────────────────────┘
▼
┌──────────────────────────┐ ┌─────────────────────────────┐
│ ① Rule Engine (regex) │ │ ② AST Analysis (Python) │
│ · 40 cross-language rules │ │ · Syntax errors (exact) │
│ · Python/JS/TS/Java/Go/ │ │ · Undefined variables │
│ Rust/C/C++/Shell │ │ · Unused imports │
│ · Security/Perf/AI/style │ │ · Duplicate definitions │
│ · Zero-cost, ms-level │ │ · Empty stub functions │
└──────────┬───────────────┘ └──────────┬──────────────────┘
▔▔▔▔▔▔▔▔┬────────────────────▘
▼
┌──────────────────────────┐
│ ③ LLM Semantic Analysis │
│ · Receives rule + AST │
│ pre-scan results │
│ · Confirms/denies hits │
│ · Semantic issues │
│ · Scores & fix_code │
└──────────┬───────────────┘
▼
┌───────────────────────────────────────────────────────────┐
│ ④ Cross-Validation Merge (merge_findings) │
│ · rule / ast / llm / confirmed (both agree → +0.3 conf) │
└───────────────────────────────┬───────────────────────────┘
▼
┌───────────────────────────────────────────────────────────┐
│ ⑤ Output: 5-dimension scores + metrics + SARIF + fixes │
│ · correctness/security/performance/maintainability/best │
│ · quality metrics (cyclomatic complexity, function length) │
│ · SARIF 2.1.0 export (VS Code / GitHub Code Scanning) │
│ · each issue includes fix_code (copy-paste ready) │
└───────────────────────────────────────────────────────────┘Features
Triple-engine review — Regex rule engine (40 rules, 9 languages, 6 AI-pattern rules) + AST-level static analysis (Python syntax/undefined vars/unused imports/duplicate defs) + LLM semantic review with cross-validation
5-dimension scoring — Correctness / Security / Performance / Maintainability / Best Practice, each 0–100, weighted composite score
Deterministic quality metrics — Cyclomatic complexity (McCabe), function length distribution, comment ratio, long lines — zero LLM cost, instant
SARIF 2.1.0 export — Standards-compliant output for VS Code (Sarif Viewer) and GitHub Code Scanning, CI-ready
GitHub PR/commit URL review — Paste a PR or commit URL, auto-fetch diff and review
Directly applicable fix code — Rule engine auto-generates
fix_codefor 8 key rule types, LLM covers complex scenariosFour review modes — Single file code, Unified Diff, Multi-file batch, GitHub PR URL
CLI one-click review —
python cli.pyreads git diff directly, no pasting neededMCP toolset — 10 tools: review / diff review / multi-file / PR review / security scan / metrics / SARIF export / rule explanation / fix generation / rule listing
Interactive workbench — Live demo with Metrics, SARIF, Rules, and PR URL tabs (free & instant, no LLM needed)
CLI One-Click Review (Recommended)
python cli.py # Review uncommitted changes (git diff)
python cli.py --staged # Review staged changes (git diff --cached)
python cli.py --commit HEAD~1 # Review the last commit
python cli.py src/utils.py # Review a single file
python cli.py --remote # Use remote Vercel deployment (no local server needed)
python cli.py --format json # Output JSON (machine-readable, for pipes/CI)
python cli.py --sarif out.sarif # Export SARIF (GitHub Code Scanning format)Exit codes: 0 no serious issues | 2 critical/major found (CI gate) | 1 runtime error
Auto-reads git diff → calls API → outputs structured report with severity icons, dimension scores, and fix code.
CI/CD Integration
GitHub Actions (PR Auto-Review)
Includes .github/workflows/code-review.yml, auto-triggers on PR to main:
Gets PR diff → calls Code Review Agent API
Fails Action if critical issues found (blocks merge)
Exports SARIF and uploads to GitHub Code Scanning (issues annotated on PR diff lines)
pre-commit hook
# .git/hooks/pre-commit
python cli.py --staged --remote || exit 1 # Blocks commit if critical/major foundSARIF + GitHub Code Scanning
python cli.py --sarif results.sarif --remote
# Then upload in GitHub Action with github/codeql-action/upload-sarif@v3API Overview
Method | Endpoint | Description |
|
| Review source code, return structured report |
|
| Review Unified Diff (PR changes) |
|
| Multi-file batch review (cross-file architecture analysis) |
|
| Review GitHub PR/commit URL (auto-fetch diff) |
|
| Generate complete fixed version for problematic code |
|
| Deterministic code quality metrics (no LLM) |
|
| SARIF 2.1.0 export (VS Code / GitHub Code Scanning) |
|
| List all rule engine rules |
|
| View single rule details and fix guidance |
|
| Health check, returns deployment commit |
|
| Deployment proof (slug + commit) |
|
| Interactive demo page |
Quick start (local)
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt
cp .env.example .env # Fill in LLM_API_KEY
uvicorn app.main:app --reloadOpen http://127.0.0.1:8000 for the demo page, or http://127.0.0.1:8000/docs for Swagger.
Example: Review Code
curl -X POST http://127.0.0.1:8000/v1/review \
-H "Content-Type: application/json" \
-d '{"code": "result = eval(user_input)", "language": "python"}'Response (abridged):
{
"report": {
"score": 68,
"grade": "C",
"dimension_scores": {
"correctness": 88, "security": 35,
"performance": 90, "maintainability": 80, "best_practice": 75
},
"issues": [
{
"severity": "critical",
"category": "security",
"line": 1,
"title": "Using eval() to execute arbitrary code",
"description": "eval() executes arbitrary strings as code, posing a severe injection risk.",
"suggestion": "Use ast.literal_eval() or a dedicated parser.",
"fix_code": "result = ast.literal_eval(user_input)",
"source": "confirmed",
"rule_id": "PY-S001",
"confidence": 1.0
}
],
"engine_info": {
"rule_count": 0, "llm_count": 0, "confirmed_count": 1,
"total_rules_run": 3, "engines": ["rule", "llm"]
}
}
}Example: Review a Diff
curl -X POST http://127.0.0.1:8000/v1/review_diff \
-H "Content-Type: application/json" \
-d '{"diff": "--- a/x.py\n+++ b/x.py\n@@ -1,3 +1,4 @@\n def f():\n- return 1\n+ return eval(data)", "language": "python"}'Response includes files_changed / added_lines / removed_lines change metadata with the full report.
Example: Multi-file Review
{
"context": "User service module",
"files": [
{"filename": "utils.py", "content": "import os\napi_key = os.environ['KEY']", "language": "python"},
{"filename": "main.py", "content": "from utils import *\nresult = eval(req.body)", "language": "python"}
]
}Returns per-file file_reports (rule scan) and one overall_report (LLM cross-file architecture review).
MCP Usage
Local stdio (Claude Code / Codex / Cursor)
python -m app.mcp_server # stdio transportRegister in client config:
{
"mcpServers": {
"code-review-agent": {
"command": "python",
"args": ["-m", "app.mcp_server"]
}
}
}Remote streamable HTTP (same deployment, no local Python needed)
After deployment, access https://<your-host>/mcp, configure in MCP client:
{
"mcpServers": {
"code-review-agent": {
"command": "npx",
"args": ["-y", "@anthropic-ai/mcp-client", "https://<your-host>/mcp"]
}
}
}The remote MCP endpoint and REST API share the same server. After deployment,
/mcpprovides streamable HTTP protocol,/v1/*provides REST.
Usage Guide (for Agents)
Free quick scan first: Use
detect_security/list_rules/explain_issue(no LLM call, ms-level response)Deep review: Use
review_code/review_diff/review_files, defaultdetail="brief"(saves context, returns title-level issues only)Full report when needed:
detail="full"returns complete description / suggestion / fix_code for each issueFix: Use
suggest_fixto get directly replaceablefixed_code
MCP Tools
Tool | Parameters | LLM | Description |
|
| ✅ | Review source code ( |
|
| ✅ | Review Unified Diff |
|
| ✅ | Multi-file batch review (structured params, not JSON string) |
|
| ✅ | Review GitHub PR/commit URL (auto-fetch diff) |
|
| ❌ | Rule engine security scan only, instant response |
|
| ❌ | Deterministic quality metrics (complexity, function length) |
|
| ❌ | Explain a rule (definition/severity/fix guidance) |
|
| ✅ | Return fixed code (fixed_code + change explanation) |
|
| ❌ | SARIF 2.1.0 export (VS Code / GitHub Code Scanning) |
| — | ❌ | List all rules |
review_filesfilesparameter is a structured array, each element{filename, content, language?}. Agents don't need to manually compose JSON strings.
Real MCP tool calls (captured output)
The following are real tool outputs from the deployed agent — non-LLM tools called locally, LLM tools called against the live endpoint.
Agent → detect_security (instant, no LLM):
Agent: detect_security(code="import os\napi_key='sk-1234567890abcdef'\nresult = eval(user_input)\nos.system('rm -rf /tmp/x')\ndata = pickle.loads(raw_data)", language="python")
Tool → {"ok": true, "total_findings": 3, "findings": [
{"rule_id": "PY-S001", "severity": "critical", "line": 3,
"title": "使用 eval() 执行任意代码",
"suggestion": "避免使用 eval()。如需解析表达式,使用 ast.literal_eval() 或专用解析器。",
"confidence": 0.95},
{"rule_id": "PY-S004", "severity": "major", "line": 2,
"title": "硬编码密钥/密码",
"suggestion": "使用环境变量或密钥管理服务:api_key = os.environ['API_KEY']",
"confidence": 0.8},
{"rule_id": "PY-S005", "severity": "major", "line": 5,
"title": "使用 pickle 反序列化不可信数据",
"suggestion": "使用 JSON 等安全格式序列化数据",
"confidence": 0.9}
]}Agent → analyze_metrics (instant, no LLM):
Agent: analyze_metrics(code="def process_data(items):\n result = []\n for i in range(len(items)):\n for j in range(len(items)):\n ...", language="python")
Tool → {"ok": true, "metrics": {
"complexity": {"average": 4.0, "max": 4,
"most_complex": [{"name": "process_data", "line": 1, "complexity": 4}]},
"functions": {"count": 1, "average_length": 7.0, "max_length": 7},
"lines": {"total": 7, "code": 7, "comment": 0, "blank": 0}
}}Agent → list_rules (instant, no LLM):
Agent: list_rules()
Tool → {"total": 40, "rules": [
{"id": "PY-S001", "severity": "critical", "category": "security", "title": "使用 eval() 执行任意代码"},
{"id": "PY-S002", "severity": "critical", "category": "security", "title": "使用 exec() 执行任意代码"},
{"id": "PY-S003", "severity": "critical", "category": "security", "title": "命令注入风险"},
{"id": "PY-S004", "severity": "major", "category": "security", "title": "硬编码密钥/密码"},
{"id": "PY-S005", "severity": "major", "category": "security", "title": "使用 pickle 反序列化不可信数据"},
... (35 more)
]}Agent → review_pr (LLM, live endpoint https://code-review-agent-ashy-six.vercel.app/v1/review_pr):
Agent: review_pr(url="https://github.com/kestarsheng/code-review-agent/commit/952fa21", language="python")
Tool → {"ok": true, "files_changed": 5, "added_lines": 240, "model": "deepseek-chat",
"report": {
"score": 46, "grade": "D",
"dimension_scores": {"correctness": 36, "security": 9, "performance": 85,
"maintainability": 88, "best_practice": 54},
"issues": [
{"severity": "critical", "source": "rule", "rule_id": "PY-AST-S001",
"line": 104, "title": "Python 语法错误,代码无法解析"},
{"severity": "critical", "source": "rule", "rule_id": "PY-S001",
"line": 229, "title": "使用 eval() 执行任意代码",
"fix_code": "result = ast.literal_eval(x)"},
{"severity": "major", "source": "llm",
"line": 78, "title": "fetch_diff 跟随重定向且未校验最终主机,存在 SSRF 风险"},
... (9 more)
],
"engine_info": {"rule_count": 8, "llm_count": 4, "confirmed_count": 0,
"engines": ["rule", "ast", "llm"]}
}
}The
review_prcall demonstrates the full pipeline: GitHub URL → diff fetch → triple-engine review → structured report with cross-engine attribution (source: "rule"vssource: "llm") and auto-generatedfix_code.
Rule Engine
Built-in 40 cross-language rules covering Python / JavaScript / TypeScript / Java / Go / Rust / C/C++ / Shell / PHP:
Category | Count | Examples |
Security | 15 |
|
Performance | 6 | Nested loops O(n²), dict iteration without |
AI Pattern | 6 | Hallucinated imports ( |
Maintainability / Best Practice | 13 | TODO/FIXME, bare |
8 key rule types have auto fix code generation (eval→ast.literal_eval, innerHTML→textContent, hardcoded secret→os.environ, etc.).
Configuration (Environment Variables)
Var | Default | Description |
|
| OpenAI-compatible base URL |
| — | API key (required) |
|
| Model name |
|
| LLM request timeout |
|
| Max characters per review |
|
| Deployment commit, returned by /health and verification file |
Deployment
Vercel (current):
vercel.jsonconfigured for Serverless service; push after setting env vars in Vercel projectDocker:
docker build -t code-review-agent . && docker run -p 8000:8000 code-review-agentRender: Use
render.yaml, push repo and set env vars
Post-deploy verification:
curl https://<your-host>/health
curl https://<your-host>/.well-known/xagent-verification.jsonTesting
python -m pytest tests/ -v93 unit tests covering rule engine (40 rules), AST analysis, diff parsing, 5-dimension scoring, fix code generation, multi-file review, PR URL review, metrics, SARIF export, and full triple-engine flow.
License
UNLICENSED — submission-only use for X-Agent AI MCP Hackathon 2026.
This server cannot be deployed
Maintenance
Related MCP Connectors
Security reviews for coding agents: diffs checked against your org policy and live infrastructure.
Agentic code review, no signup to try: reality gates + frontier-model review, with veto.
Code intelligence platform for AI agents. 20 tools for architecture, security & impact analysis.
Lints + auto-fixes how AI coding agents discover any new product. 24 rules, 6 tools, score 0-100.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceEnables AI-powered, zero-trust code review with multiple models, supporting single files, git diffs, and multiple files, with security, performance, and architecture checks across 10+ languages.15MIT

CodePeel MCP Serverofficial
FlicenseAqualityDmaintenanceEnables AI agents to review code diffs for bugs, security issues, and bad patterns, and generate fixes.4-- AlicenseNot gradedqualityDmaintenanceEnables AI-powered code review and improvement, including analysis, refactoring suggestions, and automatic test generation, with an optional agentic loop for iterative refinement.MIT
- AlicenseNot gradedqualityAmaintenanceEnables AI agents to scan code for security and quality issues and receive machine-readable reports with suggested fixes and verification criteria.136 npm2MIT