ai-review
README.md
# AI Code Review Agent
An autonomous GitHub pull-request review agent built on **Claude + MCP + RAG**.
When a PR is opened or updated, a webhook enqueues a background job. A worker
loads the repo's config, pulls the diff, retrieves relevant context from a RAG
index of the codebase, and runs a Claude tool-use loop that inspects files and
submits a structured review. Comments are filtered (severity threshold, ignored
paths, previously-dismissed suggestions) and posted back to the PR as a single
batched GitHub review. The same five review tools are also exposed over a
**Model Context Protocol (MCP)** stdio server so you can drive them from Claude
Code interactively.
The whole system **runs and tests green with zero external API keys** — Claude,
OpenAI embeddings, ChromaDB, Redis, and Postgres are all optional. When a
dependency is absent the code degrades gracefully (local embedder, in-memory
vector store, in-memory idempotency, and a static regex-based fallback review).
---
## Architecture
```
GitHub PR event
│ (HMAC-signed webhook)
▼
┌───────────────────────────────┐
│ Webhook (FastAPI) │ src/ai_review/webhook/
│ • verify signature │
│ • idempotency (delivery_id) │
│ • enqueue Celery task │
└───────────────┬───────────────┘
│ Redis queue "reviews"
▼
┌───────────────────────────────┐
│ Worker (Celery) │ src/ai_review/worker/
│ 1. load .github/ai-review.yml│
│ 2. fetch PR + diff (GitHub) │
│ 3. RAG context prefetch │──► RAG service (src/ai_review/rag/)
│ 4. Claude agent loop ────────┼──► tools.py (5 shared tools)
│ ↳ fallback: static rules│ ▲
│ 5. filter (threshold / │ │ same tools
│ ignore / learning) │ │
│ 6. post batched review │ ┌─────┴─────────────────┐
└───────────────────────────────┘ │ MCP stdio server │
│ (Claude Code) │
└───────────────────────┘
Cross-cutting: settings, models, github_client, learning store,
Prometheus metrics, structlog JSON logging.
```
Key idea: the five review tools (`github_get_pr`, `github_get_file`,
`github_get_repo_structure`, `github_post_review_comment`, `rag_query_repo`)
live in one shared module, [`src/ai_review/tools.py`](src/ai_review/tools.py).
Both the MCP server and the worker's agent loop call the *same* functions, so
behaviour is identical whether Claude drives them interactively over MCP or the
worker drives them autonomously.
---
## Requirements
- **Python 3.11+**
- Optional for full deployment: Redis, Postgres, ChromaDB, Docker
- Optional API keys: `ANTHROPIC_API_KEY` (real reviews), `OPENAI_API_KEY`
(OpenAI embeddings). Without them the system uses offline fallbacks.
---
## Setup
```bash
python -m venv .venv
source .venv/Scripts/activate # Windows (Git Bash); use .venv/bin/activate on macOS/Linux
# Install the extras you need. For everything + dev tooling:
pip install -e ".[webhook,worker,mcp,rag,dev]"
```
Optional extras, if you only want part of the system:
| Extra | Pulls in | For |
|-----------|--------------------------------------------|-----------------------------|
| `webhook` | fastapi, uvicorn, redis, celery | the HTTP webhook service |
| `worker` | celery, redis, anthropic, httpx, crypto | the review worker |
| `mcp` | mcp, httpx | the MCP stdio server |
| `rag` | chromadb, openai | persistent RAG / OpenAI emb.|
| `dev` | pytest, ruff, mypy | tests + linting |
Copy the environment template and edit as needed:
```bash
cp .env.example .env
```
Everything in `.env` is optional for a local run — the defaults use SQLite, a
local embedder, and in-memory fallbacks.
---
## Running the tests
```bash
pytest -q
```
50 tests, all hermetic (no network, no API keys). Lint with:
```bash
ruff check .
```
---
## Running the services locally
### Webhook (FastAPI)
```bash
uvicorn ai_review.webhook.app:app --reload --port 8000
```
Endpoints:
| Path | Purpose |
|-----------------------|------------------------------------------------|
| `POST /webhook/github`| GitHub webhook receiver (HMAC-verified) |
| `GET /health` | liveness |
| `GET /ready` | readiness |
| `GET /metrics` | Prometheus metrics |
Set `GITHUB_WEBHOOK_SECRET` to the same value configured on the GitHub App/repo
webhook. Requests with a missing or bad `X-Hub-Signature-256` get `401`.
### Worker (Celery)
```bash
celery -A ai_review.worker.tasks worker -Q reviews --loglevel=info --concurrency=4
```
Needs Redis (`REDIS_URL`). Without `ANTHROPIC_API_KEY` the worker still runs and
posts a **static fallback review** (regex rules for leaked secrets, `verify=False`,
`eval(...)`, SQL string concatenation, etc.), so the pipeline is testable
end-to-end offline.
### MCP server (Claude Code)
```bash
python -m ai_review.mcp_server
```
This speaks MCP over stdio. To wire it into Claude Code, point your MCP config at
it — see [`examples/claude_mcp_config.json`](examples/claude_mcp_config.json):
```json
{
"mcpServers": {
"ai-review": {
"command": "python",
"args": ["-m", "ai_review.mcp_server"],
"env": {
"GITHUB_APP_ID": "123456",
"GITHUB_PRIVATE_KEY_PATH": "/absolute/path/to/app-private-key.pem",
"GITHUB_TOKEN": "",
"EMBEDDING_PROVIDER": "local"
}
}
}
}
```
For quick local experiments you can set `GITHUB_TOKEN` to a personal access
token instead of configuring a full GitHub App.
---
## Choosing an LLM provider
The review agent defaults to Anthropic, but the same agent loop runs against any
**OpenAI-compatible** endpoint through a built-in adapter — so you can use free
providers like **NVIDIA NIM**, Groq, OpenRouter, or a fully local Ollama. Only
env vars change; no code changes.
**NVIDIA NIM** (free credits at [build.nvidia.com](https://build.nvidia.com)):
```bash
LLM_PROVIDER=nvidia
NVIDIA_API_KEY=nvapi-...
# Defaults (override if you like):
# LLM_BASE_URL=https://integrate.api.nvidia.com/v1
# LLM_MODEL=meta/llama-3.3-70b-instruct
```
**Any other OpenAI-compatible provider:**
```bash
LLM_PROVIDER=openai_compat
LLM_API_KEY=...
LLM_BASE_URL=https://api.groq.com/openai/v1 # or openrouter / ollama / …
LLM_MODEL=llama-3.3-70b-versatile
```
| Provider | `LLM_BASE_URL` | Example model |
|--------------|-----------------------------------------|-----------------------------------|
| NVIDIA NIM | `https://integrate.api.nvidia.com/v1` | `meta/llama-3.3-70b-instruct` |
| Groq | `https://api.groq.com/openai/v1` | `llama-3.3-70b-versatile` |
| OpenRouter | `https://openrouter.ai/api/v1` | `meta-llama/llama-3.3-70b-instruct` |
| Ollama (local)| `http://localhost:11434/v1` | `llama3.3` |
> **Tool-calling matters.** The agent works by having the model call a
> `submit_review` tool. Large models (Llama 3.3 70B and up) do this reliably;
> very small local models may not, in which case the pipeline falls back to the
> static review. NVIDIA's 70B-class models are a safe default.
The Anthropic path is unchanged — leave `LLM_PROVIDER=anthropic` (the default)
and set `ANTHROPIC_API_KEY` to use Claude. With no provider key set at all, the
worker still posts the static fallback review.
---
## Docker deployment
```bash
cp .env.example .env # fill in real values first
docker compose up --build
```
Compose brings up: `redis`, `postgres`, `chromadb`, the `webhook` service
(port 8000), 3 `worker` replicas, `prometheus` (9090), and `grafana` (3000).
The webhook and worker images install only the extras each service needs
(see [`docker/webhook.Dockerfile`](docker/webhook.Dockerfile) and
[`docker/worker.Dockerfile`](docker/worker.Dockerfile)).
Expose the webhook publicly (reverse proxy / tunnel) and register that URL as
the GitHub App's webhook, using the same `GITHUB_WEBHOOK_SECRET`.
---
## Per-repository configuration
Repos control the agent via `.github/ai-review.yml` (see the sample
[`.github/ai-review.yml`](.github/ai-review.yml) in this repo):
```yaml
review:
enabled: true
severity_threshold: warning # info | warning | error
categories: # per-category on/off
security: true
correctness: true
performance: true
architecture: true
testing: true
style: false
ignore_paths: ["*.lock", "vendor/**", "**/*.generated.*", "**/migrations/**"]
max_comments_per_pr: 50
require_approval_for: [] # e.g. ["security:error"]
custom_rules: | # extra natural-language rules for the agent
- "All API endpoints must have rate limiting"
learning:
enabled: true
dismissed_suggestion_ttl_days: 30
```
Config is fetched from the target repo, validated with pydantic, and cached with
a TTL. Missing/invalid config falls back to sane defaults.
---
## Environment variables
| Variable | Default | Notes |
|----------|---------|-------|
| `GITHUB_APP_ID` | – | GitHub App ID (production auth) |
| `GITHUB_PRIVATE_KEY` / `GITHUB_PRIVATE_KEY_PATH` | – | App PEM (inline or file) |
| `GITHUB_WEBHOOK_SECRET` | – | HMAC secret; **required** for the webhook |
| `GITHUB_TOKEN` | – | PAT fallback for local dev (not for prod) |
| `ANTHROPIC_API_KEY` | – | Enables real Claude reviews; empty → static fallback |
| `ANTHROPIC_MODEL` | `claude-sonnet-5` | Review model (Anthropic provider) |
| `LLM_PROVIDER` | `anthropic` | `anthropic`, `nvidia`, or `openai_compat` |
| `NVIDIA_API_KEY` | – | NVIDIA NIM key (when `LLM_PROVIDER=nvidia`) |
| `LLM_MODEL` | provider default | Overrides the model; NVIDIA default `meta/llama-3.3-70b-instruct` |
| `LLM_BASE_URL` | provider default | OpenAI-compatible endpoint; NVIDIA default `https://integrate.api.nvidia.com/v1` |
| `EMBEDDING_PROVIDER` | `local` | `local` (offline) or `openai` |
| `OPENAI_API_KEY` | – | Required only when provider is `openai` |
| `REDIS_URL` | `redis://localhost:6379/0` | Celery broker / idempotency |
| `DATABASE_URL` | `sqlite:///./aireview.db` | Learning store |
| `CHROMA_HOST` / `CHROMA_PORT` | `localhost` / `8001` | Empty host → embedded persistent client |
| `CHROMA_PERSIST_DIR` | `./.chroma` | Local Chroma path when no host |
| `MAX_COMMENTS_PER_PR` | `50` | Hard cap on posted comments |
| `REVIEW_TIMEOUT_SECONDS` | `180` | Per-review budget |
| `LOG_LEVEL` / `LOG_JSON` | `INFO` / `true` | structlog config |
See [`.env.example`](.env.example) for the full list.
---
## How a review is decided
1. **Categories & priority** (highest → lowest): security → correctness →
performance → architecture → testing → style.
2. **Severity threshold**: comments below the repo's threshold are dropped.
3. **Line targeting**: inline comments only attach to lines actually added in
the diff.
4. **Learning suppression**: suggestions a maintainer previously dismissed are
fingerprinted and suppressed for `dismissed_suggestion_ttl_days`.
5. **Quality score**: `100 − Σ(category_weight × severity_weight)`, clamped to
`[0, 100]`, plus the top concerns, are included in the summary comment.
Every posted review carries an AI-generated disclaimer.
---
## Runbook
**A PR opened but no review appeared.**
- Check the webhook received it: webhook logs + `GET /metrics`
(`webhook_events_total`). A `401` means the signature/secret is wrong.
- Check the worker consumed it: worker logs for the `reviews` queue, and that
Redis is reachable (`REDIS_URL`).
- If Claude is unreachable/no key, a static fallback review should still post —
its absence points at GitHub auth (App token or `GITHUB_TOKEN`) or the
target-repo permissions.
**Duplicate reviews.**
- Idempotency is keyed on the GitHub `X-GitHub-Delivery` id via Redis `SET NX`.
If Redis is down, the in-memory fallback only dedupes within a single process,
so duplicates are possible across replicas — restore Redis.
**Reviews are too noisy / too quiet.**
- Tune `review.severity_threshold` and `review.categories` in the repo's
`.github/ai-review.yml`. Lower the threshold to `info` for more comments,
raise to `error` for only the most serious.
**A recurring false positive keeps coming back.**
- Dismiss the comment on GitHub; the deletion webhook records it in the learning
store and it's suppressed for `dismissed_suggestion_ttl_days`. Confirm
`learning.enabled` is true.
**Metrics / dashboards.**
- Prometheus scrapes the webhook `/metrics`; Grafana is preconfigured in
`docker-compose.yml` (admin password via `GRAFANA_PASS`). Key series:
`webhook_events_total`, `reviews_total`, `mcp_calls_total`,
`review_duration_seconds`.
---
## Project layout
```
src/ai_review/
settings.py # env config (pydantic-settings)
models.py # ReviewComment, ReviewResult, PRRef, weights
diff_utils.py # unified-diff parsing → changed/added lines
github_client.py # httpx GitHub client (PRs, files, tree, post review)
github_auth.py # GitHub App installation tokens (PyJWT) + PAT fallback
tools.py # the 5 shared review tools
config_repo.py # .github/ai-review.yml loader + validation
learning.py # dismissed/accepted suggestion store (SQLAlchemy)
metrics.py # Prometheus metrics
logging_config.py # structlog JSON logging
webhook/ # FastAPI app, signature verify, idempotency
worker/ # Celery task, Claude agent loop, fallback, pipeline,
# scoring, filtering, prompts, rendering
rag/ # chunking, embeddings, ingest, service (Chroma/in-mem)
mcp_server/ # MCP stdio server exposing tools.py
tests/ # 50 hermetic tests
docker/ # webhook + worker Dockerfiles
docker-compose.yml # full local stack
examples/ # Claude Code MCP config
```
---
## Design notes
- **Offline-first.** Every external dependency has a fallback so the system is
runnable and testable without accounts or keys. Swap in real Claude / OpenAI /
Chroma / Redis / Postgres by setting env vars — no code changes.
- **One tool implementation.** MCP and the autonomous worker share `tools.py`,
eliminating drift between interactive and automated review.
- **Graceful degradation over failure.** If Claude is unavailable the worker
posts a static rule-based review rather than nothing (PRD NFR 4.2).
## License
MIT.
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues