Context Zero Engine
by Classevelabs
README.md
# Context Zero Engine
[](https://github.com/Classevelabs/context-zero-engine/releases/latest) [](LICENSE)
**A local code-intelligence engine for AI agents.** ContextZero indexes a
repository into a PostgreSQL-backed code graph and serves structured,
token-budgeted context — symbols, dependencies, effects, contracts, similar
code, and blast radius — over MCP and HTTP. The engine and its database run
locally and do not require an external analysis or embedding API. If an
operator enables repository validation commands, those commands inherit the
repository's own behavior and may access the network.
Built by [ClassEve](https://classeve.com). Licensed under Apache-2.0.
> **Official repository.** This is the only official repository for Context Zero Engine.
> ClassEve's complete list of official accounts is at [classeve.com/official](https://classeve.com/official).
> The GitHub account `github.com/ClassEve` is an unrelated third party, not affiliated with ClassEve.
---
## The Problem
Coding agents and developer tools usually inspect source one file at a time.
On a non-trivial codebase that means opening dozens of files, re-reading the
same code across tasks, manually tracing transitive effects — and still
missing contract assumptions or behaviorally similar code elsewhere in the
repository.
ContextZero indexes the repository once and answers the same investigation
with targeted queries: *give me this symbol with its dependencies and
contracts*, *what breaks if I change it*, *where else does this logic exist*,
*which tests cover it*.
An AI assistant asked to change one function has to see three things: the
function, the code it uses from other files, and the code that calls it.
Measured across 955 jobs on 18 public repositories, the typical job costs:
| | Searching and reading files | ContextZero |
|---|---:|---:|
| Files opened | 2 | 1 request |
| Lines of code to read | 1,413 | **135** |
| Tokens paid for | 11,577 | **1,927** |
**83% fewer tokens** — 10.1× fewer across the whole run. And it costs less
*without* knowing less: given the same tokens to spend, searching contains the
function you asked about only **1 time in 9**, while ContextZero has it **every
time**; it finds the code that calls it **50% of the time against 25%**, **6 in
10** of the helpers the code uses from other files against almost none, and a
covering test for **1 job in 4**.
Searching for a name finds the places that mention it, not the things it needs.
[BENCHMARKS.md](BENCHMARKS.md) has the method, a worked example, a second
codebase, and what the engine still does badly. Reproduce on your own repository
with `node scripts/bench-context-quality.mjs`.
---
## What It Computes
| Capability | Description |
|-----------|-------------|
| **Context Capsules** | Everything you need to understand a symbol in one call — source, dependencies, contracts, effects — inside a token budget you set. When the budget is tight it drops detail in five defined steps rather than truncating. |
| **Blast Radius** | What breaks if you change this. Scored across five kinds of coupling — structural, behavioral, contract, similar-code, and what has historically changed alongside it — with a severity and a confidence for each. |
| **Behavioral Profiling** | Functions are classified as pure / read_only / read_write / side_effecting. TS/JS external effects are **type-resolved** through the compiler. The shipped, author-designed fixture suite measured 100% precision and recall; this is regression evidence, not a claim of perfect accuracy on arbitrary repositories (see [BENCHMARKS.md](BENCHMARKS.md)). |
| **Effect Signatures** | What a function actually touches: nine typed effects (reads, writes, opens, throws, calls_external, logs, emits, normalizes, acquires_lock), each labelled as the function's own effect or one inherited through a call chain, with the hop count. |
| **Contract Extraction** | Input/output types, error contracts, security contracts, guard clauses, derived invariants — mined from the code itself. |
| **Homolog Detection** | Finds code elsewhere in the repository that does the same job, even when it shares no text with the original. Seven independent signals vote, and disagreement between them is reported rather than averaged away. |
| **Smart Context** | One call: source + blast radius + callers + tests + contracts. Replaces 8+ separate lookups. |
| **Dispatch Resolution** | Which implementation a call actually reaches — through inheritance, interfaces, and overrides — rather than just the name at the call site. |
| **Concept Families** | Groups symbols that solve the same kind of problem, names the clearest example of each group, and flags the members that break the pattern. |
| **Temporal Intelligence** | Git-derived co-change analysis, temporal risk scoring, churn metrics. |
| **Symbol Lineage** | Cross-snapshot identity tracking through renames and refactors. |
| **Transactional Editing** | 9-state change lifecycle with DB-backed rollback and 6-level progressive validation. |
| **Semantic Search** | Find code by what it does rather than what it is called. Runs locally on TF-IDF and MinHash similarity — no external API, no embedding service, no key to buy. |
| **Uncertainty Tracking** | Where extraction is unsure — a recovered parse, an unresolved type, dynamic dispatch — the symbol is flagged, and those flags aggregate into a snapshot-level confidence. It surfaces what it is *not* sure about instead of presenting every answer as equally solid. |
| **Self-Maintaining Index** | A file watcher folds each changed file into the existing snapshot within seconds of hitting disk — no scheduled job, no editor plugin. Repository-wide analysis is deferred under load and settled while you are idle, and whatever is outstanding is reported rather than assumed. |
## Languages
TypeScript, JavaScript, Python, C, C++, CUDA-flavored `.cu`/`.cuh`, Go, Rust,
Java, C#, Ruby, Kotlin, Swift, PHP, Bash — 32 file extensions across 13
parsers, since C, C++ and CUDA share the C++ parser.
TypeScript and JavaScript use full AST analysis through the TypeScript
Compiler API. Python uses LibCST with 60+ behavioral patterns. The remaining
languages use tree-sitter with language-specific walkers. CUDA files are
indexed for structure; kernel-specific semantics are not modelled separately.
## How It Works
```
MCP-compatible client (Claude Desktop, Claude Code, Codex, Cursor, ...)
|
| MCP protocol (stdio) HTTP clients
| |
ContextZero MCP Bridge (61 tools) REST API (62 routes)
| |
+------------------+------------------+
|
+-- Ingestor (13 language parsers, delta ingestion)
+-- 13 Analysis Engines
| Behavioral | Contract | Deep Contract | Blast Radius
| Effect | Dispatch | Concept Families | Temporal
| Symbol Lineage | Runtime Evidence | Uncertainty
| Structural Graph | Capsule Compiler
+-- Semantic Engine (TF-IDF, MinHash LSH, cosine similarity)
+-- Homolog Engine (7-dimensional scoring)
+-- Transactional Editor (opt-in constrained validation, rollback)
+-- Service Layer (transport-agnostic services)
+-- Database Driver (circuit breaker, batch loader, advisory locks)
|
PostgreSQL (all data local, nothing leaves your machine)
```
The `scg_` prefix on tools and environment variables comes from the engine's
internal name for its data model — the structural code graph.
Deep dives: [ARCHITECTURE.md](ARCHITECTURE.md) (subsystems and tool
registry) and [TECHNICAL_DESIGN.md](TECHNICAL_DESIGN.md) (data structures,
algorithms, engine internals).
---
## Install
Node.js 20 or newer, and nothing else. If the machine has no PostgreSQL, ContextZero creates one for
itself — in your user data directory, on a port nobody else uses, with a password it generates and
nobody has to type.
```bash
git clone https://github.com/Classevelabs/context-zero-engine.git
cd context-zero-engine
npm run setup -- --install-mcp=claude
```
That installs dependencies, provisions the database, applies the schema, builds, writes the MCP
config for your client (`claude`, `codex`, `cursor`, or `all`), and checks the result. Restart the
client and the tools are there.
Python source analysis also wants Python 3 with `libcst`; everything else works without it.
### Using a PostgreSQL you already run
Set `DB_HOST`, `DB_USER` and `DB_PASSWORD` before running setup and ContextZero uses that database
and never provisions one. It needs PostgreSQL 14 or newer, the `pg_trgm` extension, and a UTF-8
database — source code is UTF-8, and a database created with a machine's code page rejects it:
```bash
createdb -E UTF8 -T template0 scg_v2
psql -d scg_v2 -c "CREATE EXTENSION IF NOT EXISTS pg_trgm;"
```
### The database ContextZero provisions
```bash
npm run db:status # where it is, and whether it is running
npm run db:stop # stop it
npm run db:up # start it (the MCP client does this for you)
```
Its data, binaries and log live under `%LOCALAPPDATA%\ContextZero` on Windows,
`~/Library/Application Support/ContextZero` on macOS, and `~/.local/share/contextzero` on Linux.
Set `CONTEXTZERO_HOME` to put them somewhere else.
---
## Quickstart
### 1. Wire it into an MCP client
The bundled installer writes the config (with a timestamped backup of the
existing file) for Claude Desktop, Codex, or Cursor:
```bash
npm run mcp:install -- --client claude
```
Or generate config snippets without touching client files
(`npm run mcp:config`), or register manually — for example with the
Claude Code CLI:
```bash
claude mcp add contextzero -s user \
-e CONTEXTZERO_ENV_FILE=/absolute/path/to/context-zero-engine/.env \
-- node /absolute/path/to/context-zero-engine/scripts/mcp-start.mjs
```
Any MCP client that speaks stdio works: the server is
`node scripts/mcp-start.mjs` with the `DB_*`/`SCG_*` environment (or a single
`CONTEXTZERO_ENV_FILE` pointing at your `.env`). It starts the database when that
database is the one ContextZero provisioned, then becomes the bridge.
MCP uses a trusted local stdio child-process boundary; it is not a remote
network authentication layer. The 43 read tools are listed by default. The 18
tools that ingest, edit, run retention cleanup, or validate are listed only
after a local operator sets `SCG_MCP_MUTATIONS_ENABLED=true`; until then the
session is told once, at connect, that they exist and where the switch is.
Validation commands additionally require `SCG_ALLOW_UNSANDBOXED_EXECUTION=true`
and should run only on trusted repositories under a restricted
operating-system account.
### 2. Index a repository
From the MCP client, call:
```text
scg_health_check → should report status: healthy
scg_register_repo / scg_ingest_repo → index a repo under SCG_ALLOWED_BASE_PATHS
```
Then start asking: `scg_smart_context`, `scg_blast_radius`,
`scg_compile_context_capsule`, `scg_find_homologs`,
`scg_semantic_search`, ...
Three native tools (`scg_native_codebase_overview`,
`scg_native_symbol_search`, `scg_native_search_code`) work immediately
without a database — they analyze the filesystem directly.
### 3. Keep it current
```bash
npm run watch
```
Watches every registered repository and folds each change into its snapshot as
it happens, so the graph describes the code as it is rather than as it was at
the last ingest. Set `SCG_WATCH=true` to start it with the MCP server instead.
It watches the filesystem and nothing else — the same behaviour whether the code
is edited by an IDE, a coding agent, a script, or a branch switch.
### 4. Or run it as an HTTP server
```bash
npm run build
npm start # HTTP server on port 3100
```
```bash
curl http://localhost:3100/health
curl -X POST http://localhost:3100/scg_codebase_overview \
-H "X-API-Key: <your key>" -H "Content-Type: application/json" \
-d '{"repo_id": "..."}'
```
62 routes (9 GET + 53 POST) mirror the MCP tool surface plus health,
readiness, Prometheus metrics, cache, and admin endpoints. All non-health
routes require API-key authentication (`X-API-Key` or `Authorization:
Bearer`). State-changing, repository-registration, and validation-command
routes require a distinct `SCG_ADMIN_API_KEYS` credential.
### Docker (self-hosted server + bundled PostgreSQL)
```bash
cp .env.docker.example .env
# Set DB_PASSWORD, SCG_API_KEYS, and a distinct SCG_ADMIN_API_KEYS value.
docker compose up -d
```
When registering repositories from Docker, use paths under `/repos` — that
is where `SCG_REPOS_PATH` is mounted inside the container.
---
## MCP Tool Surface (61 tools)
| Category | Count | Examples |
|----------|------:|----------|
| Core | 8 | `scg_health_check`, `scg_ingest_repo`, `scg_incremental_index`, `scg_codebase_overview` |
| Symbol Intelligence | 8 | `scg_resolve_symbol`, `scg_read_source`, `scg_semantic_search`, `scg_get_tests` |
| Behavioral & Contract | 8 | `scg_get_behavioral_profile`, `scg_get_invariants`, `scg_get_effect_signature` |
| Impact Analysis | 8 | `scg_blast_radius`, `scg_compile_context_capsule`, `scg_smart_context`, `scg_find_homologs` |
| Change Planning | 4 | `scg_plan_change`, `scg_prepare_change`, `scg_apply_propagation` |
| Code Graph | 8 | `scg_get_class_hierarchy`, `scg_get_symbol_lineage`, `scg_get_co_change_partners` |
| Transactional Editing | 6 | `scg_create_change_transaction`, `scg_validate_change`, `scg_rollback_change` |
| Data Management | 3 | `scg_list_snapshots`, `scg_batch_embed`, `scg_ingest_runtime_trace` |
| Native Workspace (no DB) | 3 | `scg_native_codebase_overview`, `scg_native_symbol_search`, `scg_native_search_code` |
| Admin | 5 | `scg_admin_run_retention`, `scg_admin_db_stats`, `scg_admin_system_info` |
A session lists the 43 read tools by default; the 18 that mutate appear once
`SCG_MCP_MUTATIONS_ENABLED=true` is set. The complete registry is in
[ARCHITECTURE.md](ARCHITECTURE.md).
---
## Security
- **Local by design** — no telemetry or required external analysis APIs; opt-in repository commands retain their own network capabilities
- **SQL injection protection** — parameterized queries plus table/column allowlists for dynamic queries
- **5-layer path traversal protection** — null bytes, URL encoding, backslash handling, symlink escape checks, base-path boundary enforcement
- **Fail-closed authentication** — timing-safe comparison, 32-character minimum keys, per-IP brute-force lockout, and separate production admin credentials for privileged HTTP routes
- **Constrained validation runner** — disabled by default; applies time/output/resource limits, process groups, SIGKILL escalation, and environment sanitization, but does not isolate filesystem or network access
- **Hardened HTTP surface** — per-route rate limits and body-size limits, input validation on every route, sanitized error responses (no stack traces, paths, or SQL)
See [SECURITY.md](SECURITY.md) for the deployment hardening checklist and
how to report a vulnerability.
---
## Testing
```bash
npm test # full unit suite
npm run test:db # opt-in integration test against a real PostgreSQL
npm run test:ci # with coverage
npm run typecheck # TypeScript strict mode
npm run lint
```
---
## Documentation
| Document | Description |
|----------|-------------|
| [docs/INSTALL.md](docs/INSTALL.md) | Install paths, MCP client configuration, diagnostics |
| [docs/OPERATIONS.md](docs/OPERATIONS.md) | Day-to-day operation, indexing, network server mode |
| [ARCHITECTURE.md](ARCHITECTURE.md) | System architecture, subsystems, tool registry |
| [TECHNICAL_DESIGN.md](TECHNICAL_DESIGN.md) | Data structures, algorithms, engine internals |
| [BENCHMARKS.md](BENCHMARKS.md) | Benchmark methodology and results |
| [SECURITY.md](SECURITY.md) | Hardening checklist and vulnerability reporting |
---
## About
Built and maintained by [ClassEve](https://classeve.com) — engineering for AI agents and developer tooling. Project page: [classeve.com/public/context-zero-engine](https://classeve.com/public/context-zero-engine).
## License
Apache License 2.0 — see [LICENSE](LICENSE). Copyright 2026
[ClassEve](https://classeve.com).
This server cannot be deployed
Maintenance
ActivityActive
ResponsivenessNo issues