Skip to main content
Glama
README.md
# CodePecker

An **MCP server** that reviews a piece of code across four dimensions — **security,
standards, production readiness, sustainability** — automatically **fixes** the
issues, **verifies** the fix by running the code's tests, and reports what it did.
Any MCP-capable agent (Claude Code, Codex, Copilot) can call it as a tool; there's
also a CLI for local demos.

```
review → remediate → run tests → repeat (bounded)   →   findings + scorecard + fixed code + diff + citations
```

## Quickstart (30 seconds)

```bash
git clone <this repo> && cd CodePecker
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt && pip install -e . --no-deps
cp .env.example .env                              # add a free Groq key — see Setup
codepecker examples/urlshortener_sample/store.py  # review a flawed sample file
```

No key yet? `pytest -q` runs the whole suite fully offline. Full details, other
providers, and MCP client wiring are below.

## How it works

For each of the four dimensions, CodePecker gathers findings two ways:

- **Deterministic checks** (regex/code) for rules that must be caught reliably —
  hardcoded secrets, `eval`/unsafe deserialization, bare `except`, missing tests.
  No LLM, so they never "forget".
- **An LLM judge** for the nuanced rules (input validation, logging, timeouts,
  N+1 queries, …), with guardrails: it may only cite rules from the batch it was
  given, any evidence it quotes must appear in the code, and severity/dimension come
  from the rule metadata — hallucinated findings are dropped in code.

Each rule lives in a markdown file in `codepecker-skill/rules/` (RAG), tagged
`deterministic: true|false` so it's enforced by *exactly one* path. A hand-written,
bounded agent loop then asks the model to remediate and **re-runs the tests** — a fix
that resolves a finding but breaks the tests is not accepted.

The rule corpus is packaged as an **Agent Skill**: `codepecker-skill/` is a valid
skill (a `SKILL.md` entry point over the same `rules/` folder). So the *same* corpus
serves two surfaces from one source of truth — a Claude agent can load it as a skill
to **suggest** fixes at the desk, and the MCP server reads the same `rules/` to
**enforce** them (deterministic checks, guardrails, test-verified remediation). One
corpus, no drift: the skill suggests, the MCP tool guarantees.

## Setup

Requires **Python 3.10+**.

```bash
git clone <this repo> && cd CodePecker
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
pip install -e . --no-deps            # makes `codepecker` + `python -m codepecker.*` work
cp .env.example .env                  # then add your key (below)
```

Verify the install (fully offline — no key needed):

```bash
pytest -q                             # the test suite should pass
```

**Keys — the default needs just one.** Text runs on **Groq** (fast Llama 3.3 70B) and
embeddings run **locally** (no key). Put your Groq key in `.env`:

```
GROQ_API_KEY=...        # free key at https://console.groq.com/keys
```

Provider-agnostic via [LiteLLM] — switch model or provider with **no code change**,
e.g. `CODEPECKER_TEXT_MODEL=openai/gpt-4o` (one key does both), or go **fully offline**
with `CODEPECKER_TEXT_MODEL=ollama/llama3.1`. See `.env.example`.

## Usage

### CLI (local demo)

```bash
codepecker examples/urlshortener_sample/store.py
# or:  python -m codepecker.cli examples/urlshortener_sample/store.py
```

Prints the findings, a **rule-coverage scorecard** (how the code did against *all*
rules per dimension — passes included, so you can attest coverage, not just
violations), the remediated code, a unified diff, the rules cited, and a metrics
summary. Each run is appended to `metrics.jsonl`.

If the file imports sibling modules, pass them with `--support` (repeatable) so the
sandbox can run the code's tests instead of failing on the import:

```bash
python -m codepecker.cli examples/urlshortener_sample/service.py \
  --support examples/urlshortener_sample/store.py \
  --support examples/urlshortener_sample/shortener.py
```

Useful flags:

| Flag | What it does |
|---|---|
| `--tests PATH` | run the code's test file (it should import the code as `solution`); also silences the "no tests" finding |
| `--tests-dir PATH` | confirm tests exist (rule RDY-03) *without* running them — the way to attest coverage in detect-only mode |
| `--support PATH` | add a sibling module the code imports (repeatable) |
| `--no-remediate` | **detect-only**: report findings but never call the model to rewrite the code — bounds token burn when you just want a report |

**A realistic, multi-file example.** `examples/auth_sample/` is an auth module split
across several files (hardcoded secrets, weak crypto, missing validation) — closer to
real code than a single snippet. Review its entry point, bringing the siblings it
imports so the sandbox can run the tests:

```bash
python -m codepecker.cli examples/auth_sample/auth.py \
  --support examples/auth_sample/crypto.py \
  --support examples/auth_sample/db.py
```

See `examples/GROUND_TRUTH.md` for the exact issues this auth sample is seeded with.

> **Reliable live demo:** the loop makes many LLM calls, so a free tier's
> tokens-per-minute cap can throttle a full run. The loop is resilient — a mid-run
> rate limit is recorded and the review still completes (with degraded coverage
> noted) rather than crashing. For a *smooth* end-to-end demo, use a higher-limit
> tier or run the text model locally: `CODEPECKER_TEXT_MODEL=ollama/llama3.1`
> (no key, no limits).

### MCP server (in a coding agent)

CodePecker speaks MCP over stdio — no ports, no daemon. Every client points at the
**same command**; only the config file and its shape differ:

- **command** `/absolute/path/to/CodePecker/.venv/bin/python`
- **args** `["-m", "codepecker.server"]`
- **env** `GROQ_API_KEY` (or whichever provider key your model IDs need)

Use the absolute path to the **venv's** Python so the agent inherits CodePecker's
dependencies. Sanity-check that it launches (it waits on stdio; Ctrl-C to exit):

```bash
python -m codepecker.server
```

<details>
<summary><b>Claude Code</b> (CLI)</summary>

```bash
claude mcp add codepecker \
  --env GROQ_API_KEY=your-key \
  -- /absolute/path/to/CodePecker/.venv/bin/python -m codepecker.server
```

Add `--scope project` to share it with your team via a checked-in `.mcp.json`.
</details>

<details>
<summary><b>GitHub Copilot</b> (VS Code, Agent mode)</summary>

Create **`.vscode/mcp.json`** in the workspace — note the top-level `servers` key and
the `type` field (VS Code's shape differs from the `mcpServers` one below):

```json
{
  "servers": {
    "codepecker": {
      "type": "stdio",
      "command": "/absolute/path/to/CodePecker/.venv/bin/python",
      "args": ["-m", "codepecker.server"],
      "env": { "GROQ_API_KEY": "your-key" }
    }
  }
}
```

Open Copilot Chat → switch to **Agent** mode → `codepecker` shows up in the tools
picker. (To avoid hardcoding the key, use VS Code's `"inputs"` secret prompt.)
</details>

<details>
<summary><b>Cursor</b> · <b>Windsurf</b> · <b>Claude Desktop</b> (shared JSON shape)</summary>

Identical `mcpServers` block; only the file location differs:

- **Cursor** — `.cursor/mcp.json` (project) or `~/.cursor/mcp.json` (global)
- **Windsurf** — `~/.codeium/windsurf/mcp_config.json`
- **Claude Desktop** — `claude_desktop_config.json` (macOS: `~/Library/Application Support/Claude/`)

```json
{
  "mcpServers": {
    "codepecker": {
      "command": "/absolute/path/to/CodePecker/.venv/bin/python",
      "args": ["-m", "codepecker.server"],
      "env": { "GROQ_API_KEY": "your-key" }
    }
  }
}
```
</details>

<details>
<summary><b>OpenAI Codex</b> (CLI)</summary>

Add to **`~/.codex/config.toml`** (TOML, not JSON):

```toml
[mcp_servers.codepecker]
command = "/absolute/path/to/CodePecker/.venv/bin/python"
args = ["-m", "codepecker.server"]
env = { GROQ_API_KEY = "your-key" }
```
</details>

> MCP config conventions move fast. If a client has renamed a key or moved its config
> file, check that client's own MCP docs — only the `command` / `args` / `env` values
> above are CodePecker-specific.

The server exposes one tool:

```
review_and_remediate(code, language="python", tests="", tests_dir="", support_files={})
```

- **code** — the source to review.
- **tests** *(optional)* — a separate test file; the code should import as `solution`
  (`from solution import ...`). Passing it runs the tests and suppresses the "no
  tests" finding.
- **tests_dir** *(optional)* — path to the code's test directory; confirms tests exist
  (rule RDY-03) *without* running them. This is how coverage is attested in detect-only
  mode (`CODEPECKER_REMEDIATE=false`), which skips test execution.
- **support_files** *(optional)* — `{"sibling.py": "<source>", …}` for modules the code
  (or its tests) imports, so they resolve in the sandbox instead of crashing test
  collection.

Detect-only vs. remediate is controlled by the `CODEPECKER_REMEDIATE` env var (default
`true`); set it `false` to report findings without ever calling the model to rewrite
code — the same behaviour as the CLI's `--no-remediate`.

### Evaluation

```bash
python eval/run_eval.py
```

Runs CodePecker over the labeled golden set (`eval/golden/`) and reports
precision/recall/F1 per dimension, remediation resolution + test-pass rates, and mean
iterations/latency; writes `eval/report.json`. This is the "how do I know it's good?"
evidence and is meant to run in CI. (It drives the full loop over every sample, so use
a decent rate-limit tier.)

## Design decisions (the short "why")

| Decision | Why |
|---|---|
| **MCP server**, not a bot/CI check | Reusable across agents, and reviews *in the loop* rather than post-hoc |
| **Hand-written loop**, no LangChain | Bounded task; transparent and testable control flow |
| **RAG** over fine-tuning for rules | Rules stay editable, auditable, and citable (markdown files) |
| **Deterministic** secrets/eval/except vs **LLM** for nuance | Reliability where it's non-negotiable, flexibility where it's fuzzy |
| **Tests gate success** | A fix that breaks behavior is a failure, not a fix |
| **Judge guardrails** (constrained citations + evidence grounding) | Hallucinated findings are dropped by code, not trusted |
| **One LLM seam** (LiteLLM behind `LLMClient`) | Swapping provider — or going offline — is a config change |
| **Sandboxed test run** (subprocess + timeout) | Executing untrusted code is a security boundary |


## Project layout

```
src/codepecker/
  config.py            env-driven model IDs + tuning constants
  types.py             LLM Protocols (DIP/ISP) + the Finding type
  llm_client.py        the only module that talks to a provider (LiteLLM)
  vector_store.py      ChromaDB adapter (RAG index)
  knowledge/loader.py  parse + embed the markdown knowledge banks
  tools/
    deterministic_checks.py   code checks, keyed by rule id
    judge.py                  batched, guardrailed LLM judge
    run_tests.py              sandboxed pytest runner
  agent.py             review_and_remediate() — the bounded loop
  metrics.py           append-only metrics log + summary
  evaluation.py        pure detection/remediation metrics (used by eval/run_eval.py)
  cli.py               local demo runner
  server.py            FastMCP server (stdio)
codepecker-skill/      the corpus as an Agent Skill (one source of truth)
  SKILL.md             agent-facing entry point (the "suggest" surface)
  rules/               the rules: security/ standards/ readiness/ sustainability/
eval/                  golden samples + run_eval.py
tests/                 the test suite
```

## Testing

```bash
pytest -q                       # 97 offline tests (local embeddings, faked LLM)
pytest -m "live or not live"    # + the 1 live acceptance test (needs GROQ_API_KEY)
```

The default suite is fully offline and deterministic; the one live test is opt-in.

## Non-goals / next steps

MVP simplifications, called out honestly:

- **Sandbox** is a subprocess + timeout, not a container — production wants
  gVisor/a microVM with no network and resource limits.
- **Deterministic checks** are regex-based — production would use AST analysis.
- **Local embeddings** (all-MiniLM-L6-v2) trade recall for zero keys — swap in a
  hosted embedder for higher-quality retrieval at scale.
- Not yet: metadata-routed retrieval for very large rule sets, a metrics dashboard,
  real GitHub integration, runtime energy profiling, remote HTTP/Cloud Run deploy.

[LiteLLM]: https://docs.litellm.ai/

TDQS

A4.6/5.0

Scored across 1 tool

Disambiguation5/5

With only one tool, there is no possibility of selecting the wrong tool. The tool's purpose is clearly defined, covering review and remediation in a single action.

Naming Consistency5/5

The single tool name 'review_and_remediate' follows a clear verb_noun pattern, and consistency is trivially maintained with only one tool.

Tool Count3/5

The server has just one tool, which is below the typical 3-15 range. However, the tool is comprehensive, encapsulating review, remediation, verification, and reporting, so the thin count is acceptable for a focused purpose.

Completeness4/5

The tool covers the full lifecycle from code review to remediation to test verification, and returns a detailed scorecard and diff. Minor gaps include the inability to separate review-only from remediation, but an environment variable supports detect-only mode, so core workflows are covered.

Maintenance

ActivitySlowing
ResponsivenessNo issues