Atlas
by Brair-Mpagi
README.md
# Atlas
**A local-first codebase intelligence layer for AI coding agents.**
AI coding agents waste a large share of their context budget re-deriving facts
about a codebase they've already seen where a symbol is defined, who calls it,
what breaks if it changes. Atlas builds a persistent, queryable model of a
repository and exposes it to any agent through a small, typed tool surface (an
**MCP server** plus a **CLI**), so an agent queries for structure instead of
reading dozens of files to orient itself, and opens source only for the handful
it actually needs to edit.
This repository is the **Phase 0 proof of concept** from [`project.md`](project.md):
single-language (Python), local, no embeddings the cheap, deterministic
"structure discovery" core the larger product is built on.
> **Scope honesty.** Phase 0 deliberately implements a slice, not the whole plan.
> What's here vs. what's planned is spelled out in [Status](#status) below. Where
> a shortcut is taken (name-based call resolution, lexical search), the tools say
> so in their output rather than overclaiming a core design principle of the
> plan (§11).
---
## What it does
- **Parses** Python with [tree-sitter](https://tree-sitter.github.io/) into a
graph of definitions (modules, classes, functions, methods, module-level
variables), call sites, and imports.
- **Resolves** the call graph using the receiver's class (`self.m()`), local
`x = Foo()` bindings, and explicit imports before falling back to name
matching with a **confidence** on every edge, a hard cap on calls through
unknown receiver types, and **unresolved call sites kept** rather than dropped.
- **Ranks** every symbol by importance with a personalized-PageRank engine, and
fits query results to a token budget the Aider insight that a raw graph is
still too big to hand an LLM.
- **Caches** incrementally in SQLite keyed by content hash: re-indexing only
re-parses files that actually changed.
- **Overlays** uncommitted edits: queries re-parse touched files in-memory on top
of the committed index, so answers reflect the live working tree.
- **Provenance + freshness on every answer:** each result points back to an exact
`path:line` (and commit) and is tagged `committed` / `working-tree` / `stale`
so an agent can verify a fact or fall back to reading source.
## Documentation
| Doc | For |
|---|---|
| **[USAGE.md](USAGE.md)** | Using the CLI every command, freshness/confidence, CI, troubleshooting |
| **[AGENTS.md](AGENTS.md)** | Wiring into Claude Code / Cursor, agent workflows, trust signals, limitations |
| [project.md](project.md) | The full product & technical plan this implements |
## Install
```bash
python -m venv .venv && source .venv/bin/activate
pip install -e ".[mcp,dev]" # mcp = server extra, dev = pytest
```
Requires Python ≥ 3.10. The tree-sitter Python grammar ships as a wheel no
network needed at runtime.
## CLI
```bash
atlas index # build/update the index for the repo
atlas find authenticate # locate a symbol definition
atlas callers authenticate # who calls it (ranked, confidence-flagged)
atlas callees login_endpoint # what it calls
atlas impact verify # blast radius + tests to run
atlas search "auth flow" # lexical intent search (Phase 0)
atlas context "add rate limiting" --budget 2000 # budget-fitted bundle
atlas read 'pkg/core.py::authenticate' # raw source escape hatch
atlas stats # index health + how much is guesswork
```
Add `--json` to any query command for machine-readable output. Freshness is shown
as `✓` committed, `●` working-tree, `⚠` stale.
The index lives in `.atlas/index.db` at the repo root (add it to
`.gitignore`).
## MCP server
Expose the same index to any MCP client (Claude Code, Cursor, …):
```bash
atlas mcp --root /path/to/repo # speaks MCP over stdio
```
Example Claude Code / client config:
```json
{
"mcpServers": {
"atlas": {
"command": "atlas",
"args": ["mcp", "--root", "/path/to/repo"]
}
}
}
```
### Tool surface
Deliberately small and composable (plan §7) not one "dump the graph" call:
| Tool | Purpose |
|---|---|
| `find_symbol(name)` | Locate a definition, with provenance and a purpose summary |
| `find_callers(symbol)` / `find_callees(symbol)` | One hop of the call graph, ranked |
| `impact_of_change(symbol)` | Symbols + tests affected if this changes |
| `search_intent(query)` | Lexical relevance search (semantic arrives with embeddings) |
| `get_context(task, token_budget)` | Ranked, budget-fitted bundle for a specific task |
| `read_source(node_id)` | Explicit escape hatch back to raw file content |
Every response carries a `freshness` field and a `source` pointer.
## Architecture
```
repo → parser (tree-sitter) → graph (defs/refs/calls/imports)
→ ranking (PageRank) → store (SQLite, content-hash keyed, incremental)
→ indexer (+ working-tree overlay, provenance, freshness)
→ query engine → { CLI, MCP server }
```
| Module | Responsibility |
|---|---|
| [`model.py`](src/atlas/model.py) | Portable data model + schema version |
| [`parser.py`](src/atlas/parser.py) | tree-sitter walk → symbols, calls, imports |
| [`graph.py`](src/atlas/graph.py) | Graph + approximate call resolution with confidence |
| [`ranking.py`](src/atlas/ranking.py) | Global + personalized PageRank (power iteration) |
| [`store.py`](src/atlas/store.py) | SQLite persistence, incremental, schema- and analyzer-versioned |
| [`gitinfo.py`](src/atlas/gitinfo.py) | Commit stamping + dirty-file detection |
| [`indexer.py`](src/atlas/indexer.py) | Orchestration + working-tree overlay + freshness |
| [`query.py`](src/atlas/query.py) | The six-tool query surface |
| [`cli.py`](src/atlas/cli.py) / [`mcp_server.py`](src/atlas/mcp_server.py) | Human + agent front-ends |
## Testing
```bash
pytest
```
## Measured context cost
```bash
python benchmarks/token_savings.py
```
Measured on this repo (~23k tokens to read cold) with the real `cl100k_base`
tokenizer. `oracle` is a deliberately unfair baseline it assumes the agent
already knows which files hold the answer, which is exactly what Atlas
supplies so it is the conservative floor:
| Question | Atlas | vs oracle | vs cold |
|---|---:|---:|---:|
| Who calls `build_graph`? | 2,199 | 60.1% | 90.5% |
| What breaks if I change `parse_source`? | 5,954 | 69.0% | 74.4% |
| Where is relevance ranking implemented? | 1,053 | 65.7% | 95.5% |
| What does `Indexer.index` call? | 3,228 | 68.4% | 86.1% |
| Context to add caching to the query engine | 1,739 | 81.8% | 92.5% |
| **Total** | **14,173** | **70.2%** | **87.8%** |
> **This does not validate the §9 claim.** It measures context cost only. §9
> requires token cost *and task success rate together* over a fixed task suite
> fewer tokens at lower accuracy is a product failure. Quote these as a
> necessary-but-insufficient signal, never as "70–90% token reduction".
`get_context` budgets are enforced against the real payload and verified to stay
within budget under `cl100k_base` (a 1500-token budget emits ~1376 real tokens;
the internal estimator is deliberately conservative).
## Status
Implemented (Phase 0): Python parsing, def/call/import graph, confidence-scored
call resolution, global + personalized PageRank, incremental content-hash cache,
working-tree overlay, provenance + freshness, CLI, MCP server.
Deliberately **not** in Phase 0 (see [`project.md`](project.md) §8 roadmap):
multi-language grammars, semantic embeddings, SCIP import/export, non-code
artifacts, access control, CI integration, and the benchmark suite that must
validate the token-savings claim before it's made publicly (§9). `search_intent`
is lexical until the embedding layer lands, and the call graph is approximate by
design for a dynamic language both are surfaced in tool output rather than
hidden.
This server cannot be deployed
Maintenance
ActivitySlowing
ResponsivenessNo issues