Skip to main content
Glama
README.md
<p align="center">
  <img src="assets/logo.svg" width="360" alt="Leveret logo">
</p>

# Leveret

A leveret is a young hare — small, fast, and born with its eyes open.

Leveret is a self-hosted, hybrid engine for private code reviews: the successor to
hosted AI review bots for teams that want reviews to run on their own infrastructure.
It combines a deterministic static-analysis layer, a code graph built into every
checkout, a graded noise filter with durable memory, and adversarial agent
contracts — driven by the AI you bring (BYOAI: your provider and model — Anthropic
or OpenAI by API key or subscription, or a local OpenAI-compatible endpoint). The
engine layer itself never calls an LLM, and Leveret operates no hosted analysis
service: reviewed-checkout analysis and provider credentials are never routed
through Leveret-operated infrastructure. The model provider you configure does
receive the prompts and source excerpts the runner sends it; use a local endpoint to
keep even that on your network. An optional webhook relay, if you choose one, sees
webhook payloads and a scoped token (see [DESIGN.md](DESIGN.md)); self-hosting
without it involves no relay.

Priorities are working end-to-end reviews, accurate and actionable reports,
consequential defect recall with independent verification and beyond-diff evidence,
and honest coverage reporting. The target policy is no arbitrary file-count limits
by default, with optional client limits. Today the Java path on `main` still rejects
more than 2,000 files outside its configured roots; [#80](https://github.com/leveret-dev/leveret/pull/80)
removes that default and is not yet merged. Evidence-pack file-detail caps also
remain and must be evaluated against coverage quality, not service economics. See
[DESIGN.md](DESIGN.md#current-priorities-and-deployment-reality-owner-decision-2026-10-02).

## How a review works

```mermaid
flowchart TD
    D[/"📄 PR diff"/]:::gh
    S["🔍 scan<br>engines + delta vs base<br>+ profile + memory"]:::core
    R["🐇 review agent<br>five lenses,<br>cross-file blast radius"]:::agent
    V["⚖️ verification agent<br>refute or evidence,<br>three grades"]:::agent
    T[/"📋 tiered findings<br>+ walkthrough report"/]:::gh
    D --> S
    D --> R -- concerns --> V -- survivors only --> T
    S -- bounded post-walk leads --> V
    classDef gh fill:#6ea8fe,stroke:#3d6fd9,color:#111
    classDef tun fill:#ffc86b,stroke:#cc8f22,color:#111
    classDef core fill:#7ed6a2,stroke:#3d9e6a,color:#111
    classDef agent fill:#c9a0f5,stroke:#9059d1,color:#111
    classDef store fill:#9fd8e3,stroke:#4d9aab,color:#111
```

1. **Deterministic first pass.** Engines run only against what the change touches:
   semgrep (registry security + per-language rulesets, offline-capable), gitleaks
   (secrets over the commit range), shellcheck, ruff, actionlint, zizmor (workflow
   security), osv-scanner (lockfile CVEs), typos, jscpd (duplication, corpus-gated),
   custom semgrep/ast-grep rule packs, and any SARIF-emitting command via profile
   `custom:` entries ([recipes](docs/recipes.md): psalm taint, hadolint, trivy, …).
   **Delta scanning** is on by default with a base ref: findings already present at
   the base tree are dropped as pre-existing — counted, never silent — with multiset
   identity (a *copy* of a known-bad line still surfaces), rename tracking, and
   surfaced base-pass failures. A code graph is generated into the checkout at the
   exact reviewed commit, so agents query structure instead of grepping for it.
2. **Accounted filter.** After discovery completes, surviving deterministic leads
   enter verification as one bounded, routed stream. Every supplied concern and lead
   ends as `actionable`, `priced-noise`, `false-positive`, or `dropped`; profile and
   memory suppression happens before routing. Nothing is silent: suppression,
   exact-mechanism deduplication, overflow IDs/bytes, verifier rationale, and
   publication state remain separate structured accounting.
3. **Memory that learns from humans.** `.leveret/memory.jsonl`, versioned in the
   reviewed repo: fingerprint verdicts (optionally anchored to a source line — the
   memory dies when the line changes) plus **conventions** — free-text rulings
   taught by maintainers via `learn`, injected into the agent prompts as repo case
   law, able both to suppress noise and to *raise* findings that violate them.
4. **Adversarial contracts.** The discovery walk runs five lenses (correctness and
   hostile inputs, contract conformance, test honesty, blast radius, and explicitly
   deferred leads triage) without deterministic lead material, and traces changed
   symbols to call sites *outside* the diff. After the walk, the verification agent
   tries to refute every concern and routed lead; claims it can neither refute nor
   ground in executed evidence are dropped, not published.
5. **Reporting.** Findings publish in importance tiers (`critical / major / minor /
   nit`, distinct from engine severity), out-of-diff findings appear with their
   stated correlation to the change, pre-existing defects adjacent to edited lines
   return as reminders, and every review carries a walkthrough: per-lens outcomes
   (clean included), per-file verdicts, post-walk lead/overflow metrics, the engine
   table, and a run-configuration line naming the harness, model, and thinking level
   that produced the review.

## Ways to run it

**GitHub App (autonomous).** A self-hosted App layer receives PR webhooks, checks
out the head, builds the code graph, runs the scan, drives the standardized runner,
and posts the review — inline comments plus walkthrough. The App holds only a GitHub
App key and webhook secret; model credentials live exclusively in the runner. Human
replies on findings feed `learn`. Getting started + diagram: [docs/app.md](docs/app.md).

**Standardized runner.** `leveret-runner-pi` drives the review/verify contracts
through a pinned [Pi](https://github.com/earendil-works/pi) runtime. Leveret supplies
the system prompt, phase-specific terminal submission schema, and read-only review
tools; Pi supplies the provider/model runtime and trusted host resources. Models
submit phase results through `leveret_submit_phase`; assistant text is not parsed as
JSON. Host-installed Pi/OMP extensions and hooks, Pi/Claude/
Codex skills, prompt templates, and context are loaded. The reviewed checkout is
never Pi's working directory, so its settings, hooks, skills, prompts, MCP
configuration, and context cannot extend the session. You choose provider, model,
and effort (`--model` / `--effort` / `--provider`, or the matching
`LEVERET_RUNNER_*` env vars; defaults `openai/gpt-5.6-sol` at `high`). Every
walkthrough records the effective client, model, prompt hash, capabilities, and tool
metrics. A custom `LEVERET_RUNNER` remains the bring-your-own-harness escape hatch.

Autonomous reviews retain a private, owner-controlled audit trace by default:
Pi's native per-attempt sessions, normalized harness events, App/scanner/subprocess
activity, exact failed output, checksums, and a verified zstd-or-gzip archive under
`LEVERET_DATA`. Raw content never enters default stdout. See
[Private audit traces](docs/app.md#private-audit-traces) for policy, retention,
export, security, and inspection controls.

**Interactive (MCP).** Register the server in any MCP-capable client and drive
reviews yourself — the served `review`/`verify` prompts arrive with your repo's
accumulated rulings substituted in (getting started + diagram:
[docs/interactive.md](docs/interactive.md)):

```sh
npm install && npm run build
claude mcp add leveret -- node /path/to/leveret/dist/server.js
```

MCP tools: `scan`, `ast_search` (structural search via ast-grep), `context`
(per-function complexity, churn, recency — prioritization signal, not findings),
`remember` (persist a graded verdict), `memory` (inspect the store), `learn`
(persist a human-taught convention); MCP prompts: `review`, `verify`.

## The reviewer toolbelt

The engines and structural indexes are capabilities of the reviewer, not the
reviewed repository: install them beside Leveret. Full belt: `codegraph`,
`graphify`, `semgrep`, `gitleaks`, `shellcheck`, `ruff`, `actionlint`, `zizmor`,
`osv-scanner`, `typos`, `jscpd`, `ast-grep`, `lizard`, and a pre-staged Serena
LSP bundle for semantic navigation. From a clone, build one with
`node dist/runner/prefetch-serena.js --bundle /opt/leveret/serena-bundle` and run
with `LEVERET_SERENA_BUNDLE` set to that path (the installed package also exposes
`leveret-prefetch-serena`). Runtime downloads are refused.

Before autonomous model work, Leveret builds and validates exact-checkout
CodeGraph and code-only Graphify indexes, then warms one Serena symbol query per
detected packaged language. Missing indexes fail closed by default; set
`LEVERET_REQUIRE_INDEXES=0` only for an explicitly degraded run.

Optional Java reference evidence requires the installed Inspect distribution,
cache-only artifacts, Linux Bubblewrap and `prlimit`, and paired
`LEVERET_INSPECT_JAVA_CONFIG` / `LEVERET_INSPECT_JAVA_CONFIG_SHA256` host values.
It supports one Java module with main/test sources, distinguishes exact
overloads and method-reference expressions, and reports unresolved or skipped
scope separately. A complete response is not proof a change is safe.
See [Inspect's supported host and configuration](inspect/README.md#internal-java-reference-worker).

```sh
npm test        # integration suite; exercises the real tools
```

## Repository layout

This is the canonical Leveret monorepo. Inspect is Leveret's JVM inspection
module; it retains an independent Gradle build:

| Path | Component | Toolchain |
|---|---|---|
| repository root (`src/`, `test/`, `docs/`, …) | Leveret: engine, MCP server, runner, GitHub App | Node.js, TypeScript, npm |
| [`inspect/`](inspect/README.md) | Inspect: Java classpath and checked method-reference evidence for Leveret | JDK 25+, Gradle wrapper |

With a hash-pinned host configuration outside the reviewed checkout, the
TypeScript reviewer calls Inspect's isolated JVM worker for exact Java method
references outside a pinned diff. Without it, the capability is reported
unavailable; structural matches do not become checked calls. Inspect has no
separate service or product lifecycle. Its original private history and
research are not imported into this public monorepo.

| Command | Runs |
|---|---|
| `npm run build` / `npm test` | TypeScript build / test suite (Leveret) |
| `npm run build:inspect` / `npm run test:inspect` | `./inspect/gradlew -p inspect build` / `test` |
| `npm run build:all` / `npm run test:all` | Both components, TypeScript first |

The JVM build requires JDK 25 or newer. `build:inspect` and `build:all` also run the
JVM tests, including Linux-only classpath oracles requiring Bubblewrap, Git, Maven,
JDK 21, and network access for cold fixture caches. See
[Inspect build prerequisites](inspect/README.md#build-and-test). Inspect can also be built from
`inspect/` with `./gradlew build`.

Once per clone, run `sh scripts/setup-hooks.sh` from the repository root to activate the
tracked Git hooks. Inspect has no separate hook setup.

## Design and status

[DESIGN.md](DESIGN.md) holds the architecture and decisions: the three-grade
filter, memory and learnings, runner standardization, the GitHub App split, and the
validation benchmark that gates replacing a hosted review bot with Leveret.

## License

[AGPL-3.0-or-later](LICENSE).

TDQS

A4.3/5.0

Scored across 5 tools

Disambiguation5/5

Each tool serves a distinct function: context for prioritization, scan for finding issues, ast_search for syntactic search, remember for storing verdicts, and memory for listing stored verdicts. There is no overlap or ambiguity in their purposes.

Naming Consistency4/5

Tool names are concise and lowercase, but ast_search uses an underscore while others are single words. This is a minor deviation, but the naming style is otherwise consistent and intuitive.

Tool Count5/5

With 5 tools, the server is well-scoped for a code review assistant. Each tool covers a necessary step in the workflow without redundancy, and the count is neither too sparse nor excessive.

Completeness4/5

The tools cover the core review lifecycle: contextual prioritization, scanning, code search, verdict memory, and memory inspection. Minor gaps exist (e.g., no explicit update/delete for memory entries), but the surface is otherwise complete and functional.

Maintenance

ActivityMaintained
ResponsivenessResponsive