Cartograph
# Cartograph
**Agent-native code intelligence.** Turn any repository into a queryable code graph and serve it to coding agents over [MCP](https://modelcontextprotocol.io) — so an agent can ask *"what breaks if I change this?"* instead of grepping and hoping.
tree-sitter + SQLite. No embeddings, no vector store, no API keys, no server, no cost.
**[→ Live demo](https://gokulraj2210.github.io/cartograph-mcp/)** — generated from a real index of this repo on every push.
[](https://github.com/GokulRaj2210/cartograph-mcp/actions/workflows/ci.yml)


---
## The problem
Give a coding agent a large unfamiliar repo and watch what it does: `grep`, read a file, `grep` again, read another file. It burns context reconstructing structure that a parser could have told it in one call — and it still misses the caller three modules away that its change just broke.
The usual fix is RAG: embed the codebase, retrieve "similar" chunks. But *"who calls this function?"* is not a similarity question. It has an exact answer, and that answer lives in the call graph.
Cartograph builds the graph, then hands agents ten tools shaped for how they actually work.
```console
$ cartograph blast src/cartograph/graph/store.py
## Blast radius — file `src/cartograph/graph/store.py`
17 dependent file(s), 31 affected symbol(s), 7 test file(s).
**Tests to run first**
- `tests/test_cli.py`
- `tests/test_docs.py`
- `tests/test_incremental.py`
- `tests/test_mcp.py`
- `tests/test_resolver.py`
- `tests/test_traversal.py`
- `tests/test_views.py`
**Dependent files** (by import distance)
- `src/cartograph/graph/resolver.py` · d1
- `src/cartograph/indexer/pipeline.py` · d1
- `src/cartograph/service.py` · d1
- `src/cartograph/cli.py` · d2
…
```
One call, before the edit. Not seven greps after the test suite goes red.
---
## Quickstart
```bash
uv tool install cartograph-mcp # or: pipx install cartograph-mcp
cartograph index ~/code/my-repo # builds .cartograph/cartograph.db
cartograph arch # modules, layers, cycles, hotspots
cartograph blast src/auth/token.py # what a change here could break
cartograph callers validate_token # reverse call tree
```
### Wire it into an agent
Claude Code:
```bash
claude mcp add cartograph -- cartograph serve /path/to/repo
```
Or any MCP client, via `mcp.json`:
```json
{
"mcpServers": {
"cartograph": {
"command": "cartograph",
"args": ["serve", "/path/to/repo"]
}
}
}
```
`serve` indexes on first run if no index exists. Then ask your agent *"what would break if I changed the token validator?"* and it will call `blast_radius` instead of guessing.
---
## The ten tools
| Tool | Answers |
|---|---|
| `find_symbol` | Where is X defined? (ranked by structural importance) |
| `search_code` | Full-text over names, signatures, docstrings (BM25) |
| `get_symbol` | One symbol: signature, doc, members, callers, callees, source |
| `who_calls` | Reverse call tree — before you change a signature |
| `what_it_calls` | Forward call tree — understand code without reading every file |
| `blast_radius` | What a change could break, **and which tests to run** |
| `related_symbols` | "What else should I read?" via personalized PageRank |
| `file_summary` | What a file defines, imports, and who imports it |
| `architecture_overview` | Modules, layering, import cycles, hotspots, entry points |
| `index_stats` | Index health and the edge-resolution breakdown by rule |
Plus MCP resources (`cartograph://architecture`, `cartograph://stats`) and an `orient` prompt for a graph-first first pass at an unfamiliar repo.
**Languages:** Python, TypeScript, TSX, JavaScript, Go.
---
## Design decisions worth arguing about
### 1. Confidence is a first-class column
Without a type checker you cannot *know* that `store.who_calls()` means `GraphStore.who_calls`. You can only rank hypotheses. So rather than pretending, every edge records the rule that produced it and a confidence:
| Rule | Confidence | Intuition |
|---|---|---|
| `inferred-type` | 0.97 | the receiver's type was inferred locally |
| `receiver-type` | 0.96 | `Foo.bar()` where `Foo` is a known container |
| `inherited` | 0.955 | the owner's **base class** defines it |
| `same-file` | 0.95 | the definition is right there in scope |
| `import` | 0.90 | the file explicitly imported this name |
| `same-module` | 0.75 | sibling file in the same package |
| `unique-global` | 0.60 | exactly one repo symbol has this name, bare call |
| `name-only` | 0.45 | one match, but on an **untyped receiver** |
| `ambiguous` | ≤0.40 | N candidates, kept as N edges at 1/N each |
| `external` | 0.00 | rooted at a third-party/stdlib import |
| `unresolved` | 0.00 | genuinely unknown (dynamic, or unannotated) |
The table is ordered the way the cascade tries the rules, and that is enforced:
a rule that runs earlier must also be trusted more, or resolution would be
overruling stronger evidence with weaker.
Callers then choose their own operating point. `who_calls` defaults to ≥0.5 — **precision first**, because an agent *acts* on the answer. `blast_radius` drops to 0.3 — **recall first**, because a missed impacted test is the expensive mistake and a false positive only costs a reviewer a glance.
That `name-only` tier exists because of a real bug. `seen.add(...)` on a builtin `set` was resolving to a repo class's `add` method, purely because the name happened to be unique — and it showed up as a confident caller. A method name on a receiver you cannot type is not evidence, so it now lands below the precision line. ([test](tests/test_resolver.py))
`external` exists for honesty about metrics: on most repos the "unresolved" bucket is dominated by `typer.Option` and `sqlite3.execute`. Lumping those in makes coverage look far worse than it is, so Cartograph reports **internal resolution** — of the call sites that *could* hit a repo symbol, how many did.
### 2. Local type inference, and the point where it refuses
The single largest class of unresolvable call is `receiver.method()`. You cannot know that `self.conn.execute(...)` means `sqlite3.Connection.execute` without knowing what `conn` is — so Cartograph works it out from three shallow, high-confidence sources:
| Source | Example |
|---|---|
| annotation | `def f(store: GraphStore)`, `conn: Conn`, `s *Store` |
| construction | `store = GraphStore(...)`, `new Engine()`, `&Engine{}` |
| alias | `self.store = store`, chased transitively |
Bindings are scoped the same way definitions are, so a value assigned in `__init__` is visible from every method, and a struct field is visible from every method on that struct. Attribute chains resolve one hop: `s.conn.Exec()` is a `Conn.Exec` when `s` is a `Store` and `Store.conn` is a `Conn`.
**What it deliberately does not do** is return-type propagation, generics, control-flow merging, or anything cross-module. Those need a real type checker. The rule the [tests](tests/test_bindings.py) defend is that inference must be *right or absent*: a wrong inferred type produces a confident edge pointing at the wrong function, and an agent will act on it — strictly worse than the honest `unresolved` the tool produced before inference existed. So a variable reassigned to a second type keeps its first binding, a genuine union (`Foo | Bar`) binds nothing, and a type expression that is not a plain identifier is dropped rather than mangled.
This is also why Go gains the most per line of query: `func (s *Store) Run()` names its receiver's type outright, and that same declaration is what finally lets Cartograph attach a Go method to the struct that owns it — without it, `Store.Run` and `Cache.Run` are both just `Run` forever.
### 3. Locality is not evidence about ownership
Being in the same file or package tells you *where* a name lives, not *whose* method it is. Cartograph used to break those ties with a deterministic sort, which produced a confident, arbitrary and often wrong edge. On gin that was 435 call sites; `w.Header().Get` resolved to `Params.Get`, and `b.Bind` to `xmlBinding.Bind`.
The fix is a guard: if the receiver's owner could not be named, and the surviving candidates are members of more than one container, that is ambiguity and the graph says so. The candidates are all kept as sub-threshold edges, so recall-first `blast_radius` still finds them — only the false certainty is gone.
How wrong were those ties? Edges that local inference now types at 0.97 make a usable labelled set. Measured against it on gin:
| Candidate set | n | old rules agreed with the inferred target |
|---|--:|--:|
| name owned by exactly one container | 658 | **100.0%** |
| name owned by more than one | 296 | **66.2%** |
So the guard leaves the provably-correct half alone and withdraws from a population with a measured **one-in-three error rate** — which is well below what a ≥0.5 edge is supposed to promise.
### 4. Inheritance, using edges the graph already had
`self.method()` where `method` lives on a base class is the single most common call shape in class-heavy Python, and it used to fall through to the locality rules — that is, to a name match. But the `inherits` edges needed to answer it properly were already being extracted.
So once a rule identifies the receiver's owner and finds no member on it directly, Cartograph walks that owner's declared bases breadth-first. That approximates Python's C3 linearisation and matches it exactly for single inheritance; a diamond can pick a different branch than the interpreter would, which is why `inherited` sits *below* a direct hit rather than beside one. A visited set makes a cyclic hierarchy — `class A(B)` / `class B(A)`, which a half-saved file really does produce — terminate instead of loop.
The more interesting half is what this licenses the resolver to *conclude*. Once a pinned owner's whole ancestry has been searched, where that ancestry **ends** is the answer:
| Ancestry ends... | Verdict |
|---|---|
| entirely inside the repo | `unresolved` — the member is genuinely not ours |
| at a third-party base | `external` — the member is `unittest`'s, or Pydantic's |
| at a name we cannot account for | keep going — absence proves nothing |
The `external` case is why django gains 4.9 points from this alone: `self.assertEqual` in a Django test reaches `unittest.TestCase` through several in-repo base classes, and 15,360 call sites move from "no idea" to a correct, specific answer.
That verdict is delicate enough to have needed one prerequisite fix: `from .sansio.app import App` inside `flask/app.py` was being resolved against the *module* rather than its package, so flask's own internals were reported as third-party imports — which made the ancestry walk stop in the wrong place. Relative imports now resolve against the importer's package, except in a `__init__.py`, which *is* its package. ([test](tests/test_resolver.py))
If the owner is pinned, its entire ancestry is present in the repo, and none of it defines the member, then the answer is genuinely not here — and matching an unrelated module-level function of the same name would be inventing one. That refusal is deliberately restricted to a **pinned** owner, meaning exactly one container in the repo answers to that name. gin has both a `render.JSON` struct and a `binding.JSON` variable; treating the struct as the owner of `JSON.Bind` and then *suppressing* the answer is worse than offering a weak one, because nothing downstream can correct it. ([test](tests/test_resolver.py))
Worth **+6.9 points** on django and **+5.8** on flask.
### 5. Counting resolution per call site, not per edge
An ambiguous call site emits one edge per candidate. Counting *edges* therefore scored six guesses as six successes, and the headline rate went **up** the less the resolver knew. Cartograph now counts **call sites**, and a site counts as resolved only if its winning edge is at or above the 0.5 precision line. ([test](tests/test_resolver.py))
That correction is why the numbers below are lower than earlier versions of this README claimed — flask was never at 87%; it was at 28%, and the extra was ambiguity being counted as success.
### 6. Parsing is incremental; resolution never is
A file is reparsed only when its sha256 moves. But raw references are stored as *facts* in a `refs` table, and `edges` is recomputed as a pure function of (refs × symbols) whenever anything changed.
This is what makes "reindex after every edit" trustworthy. If resolution were also incremental, editing one file could leave an edge in *another* file pointing at a symbol that had moved. Global re-resolution makes that structurally impossible. ([test](tests/test_incremental.py))
The cost is real, so there is exactly one safe shortcut: if no file was added, reparsed, or removed, both input tables are unchanged and resolution is provably identical — so it is skipped. That took a no-op reindex of Django from **7.5s to 0.67s** with a byte-identical graph.
### 7. PageRank instead of embeddings
"Which `get` did you mean?" is a *structural* question. The `get` that forty call sites depend on is the one the agent wants, and the call graph already knows that. So symbol ranking is weighted PageRank over the call graph — stable, explainable, and free. No model, no index build, no vector store.
`related_symbols` extends the same idea: personalized PageRank seeded on one symbol, treating the graph as undirected, because when you are about to change a function both its callers and its callees are relevant context. It is the structural analogue of semantic search, and it needs no embeddings.
### 8. Tools return Markdown, not JSON, under a token budget
The consumer is a context window. A 40-symbol JSON array spends thousands of tokens on braces and repeated keys, and the model reformats it anyway. Every view here is compact Markdown with a hard token budget.
Critically, **every truncation is announced**. An agent handed 20 of 87 callers with no marker will confidently conclude the other 67 do not exist, and then delete something.
### 9. Traversal runs in SQLite, not Python
`who_calls` at depth 4 is a recursive CTE, so the whole traversal stays inside SQLite's C loop. On Django's 252k-edge graph that is ~5ms. Pulling the edge table into Python to walk it would not be.
---
## Benchmarks
Real repositories, M-series laptop, single process. Cold = full index from scratch; warm = no-op reindex.
| Repo | Files | KLOC | Symbols | Edges | Cold | Warm | DB | Internal resolution |
|---|--:|--:|--:|--:|--:|--:|--:|--:|
| [django](https://github.com/django/django) | 2,977 | 534 | 45,508 | 241,278 | 12.5s | 1.82s | 79 MB | 42.7% |
| [gin](https://github.com/gin-gonic/gin) (Go) | 98 | 24 | 1,610 | 10,006 | 0.45s | 0.04s | 2.7 MB | 74.0% |
| [flask](https://github.com/pallets/flask) | 83 | 18 | 1,624 | 4,175 | 0.25s | 0.04s | 1.7 MB | 37.3% |
Query latency (median of 5, warm):
| Repo | `find_symbol` | `who_calls` d3 | `blast_radius` | `architecture_overview` |
|---|--:|--:|--:|--:|
| django | 13.3ms | 10.9ms | 11.7ms | 246.4ms |
| gin | 0.4ms | 1.5ms | 2.3ms | 8.5ms |
| flask | 0.4ms | 0.5ms | 1.0ms | 4.0ms |
Reproduce with `scripts/bench.py`.
### What the resolver changes bought
Every column uses the per-call-site metric described above, so they are comparable to each other but *not* to the inflated figures earlier versions of this README quoted.
| Repo | v0.1 | + inference | + inheritance | + external ancestry |
|---|--:|--:|--:|--:|
| **cartograph** (self) | 48.1% | 63.6% | 63.7% | **63.7%** |
| django | 29.7% | 30.9% | 37.8% | **42.7%** |
| flask | 28.3% | 31.4% | 37.2% | **37.3%** |
| gin (Go) | 82.2% | 74.7% | 74.0% | **74.0%** |
The spread is the finding, not a disappointment.
**Inference pays in proportion to how much a codebase writes its types down.** Cartograph annotates everything and gains 15 points from it. Flask and Django annotate almost nothing, so `app.config.get(...)` stays unknowable and inference alone barely moves them — a limit of the source, not of the parser.
**Inheritance pays in proportion to how much a codebase uses classes**, which is the mirror image: it is worth ~6–7 points on Django and Flask and almost nothing on the fully-annotated, composition-heavy cartograph. The two features cover each other's blind spot, which is why both are here.
**Knowing where a hierarchy *ends* is worth as much as knowing what is in it.** Django's `self.assertEqual` reaches `unittest.TestCase` through several in-repo base classes; once the walk can see that the ancestry runs out into third-party code, 15,360 call sites stop reading as "no idea" and start reading as "that one is unittest's" — worth another 4.9 points, entirely by reclassification rather than by guessing harder.
**Go goes down, on purpose.** It gains 955 `inferred-type` edges and still loses ground, because the ambiguity guard withdrew 563 same-package name matches that had been counted as confident. Measured against the inference-labelled set, one in three of those pointed at the wrong receiver. Trading 563 confident-and-often-wrong edges for 955 exact ones and an honest `ambiguous` is the trade this tool exists to make.
---
## Architecture
```mermaid
flowchart LR
subgraph index["cartograph index"]
W[walker<br/>git ls-files] --> P[tree-sitter<br/>+ .scm queries]
P --> X[extract<br/>defs · refs · imports]
end
X --> DB[(SQLite<br/>symbols · refs<br/>edges · FTS5)]
DB --> R[resolver<br/>rule cascade]
R --> DB
DB --> RK[PageRank<br/>Tarjan SCC]
RK --> DB
DB --> S[service facade]
S --> V[views<br/>token-budgeted MD]
V --> M[MCP server<br/>10 tools]
V --> C[CLI]
M --> A((coding agent))
```
| Module | Responsibility |
|---|---|
| `indexer/walker.py` | File discovery — defers to `git ls-files` for correct `.gitignore` semantics |
| `indexer/languages.py` | One adapter per language: extensions, queries, docstrings, module keys, import resolution |
| `indexer/extract.py` | AST → symbols/references/imports, language-agnostic |
| `indexer/bindings.py` | Local type inference: the scoped `name → type` environment |
| `queries/*.scm` | tree-sitter capture patterns — the per-language knowledge, as data |
| `graph/schema.sql` | The graph: `files`, `symbols`, `refs`, `edges`, `imports`, FTS5 |
| `graph/resolver.py` | The confidence cascade |
| `graph/algorithms.py` | PageRank, personalized PageRank, iterative Tarjan SCC, layering |
| `graph/store.py` | Recursive-CTE traversal, ranked lookup, aggregates |
| `service.py` | One facade so the CLI and MCP server cannot drift |
| `views.py` | Token-budgeted Markdown |
### Scoping without combinatorial queries
The trick that keeps `queries/*.scm` small: scope is never encoded in the query. Every captured definition is indexed by its tree-sitter node id, and a reference's enclosing symbol is found by walking its `parent` chain until it hits one. That is O(tree depth) per reference and handles closures, methods, inner classes, and arrow functions for free — no per-shape patterns.
### Adding a language
Subclass `LanguageAdapter` (~40 lines) and drop in a `.scm` file. `GoAdapter` is the shortest complete example. `tests/test_queries.py` then automatically compiles your queries against the grammar and asserts they capture something.
---
## Development
```bash
git clone https://github.com/GokulRaj2210/cartograph-mcp && cd cartograph-mcp
uv sync
uv run pytest -q # 254 tests
uv run ruff check .
uv run mypy # strict
```
CI runs the suite on Python 3.11/3.12/3.13 (plus macOS), then **dogfoods**: it indexes this repo, fails on import cycles, asserts a no-op reindex reparses nothing, and drives the MCP server over real stdio. It also installs the built wheel into a clean venv and indexes with it, because packaged `.scm` files are easy to leave out of a wheel and impossible to notice locally.
The cycle gate has already earned its keep — it caught a `store → resolver → store` cycle that I introduced in this repo, which was fixed by moving the offending helper rather than by relaxing the gate.
### Notable tests
- `tests/test_queries.py` — every `.scm` compiles against **every** grammar that loads it, and captures something. A pattern valid in JavaScript (`(class_heritage (identifier))`) is an *Impossible pattern* in TypeScript, which wraps supertypes in `extends_clause`. That one line silently produced zero TypeScript symbols.
- `tests/test_incremental.py` — no stale edges after edits, deletions, or a symbol moving between files.
- `tests/test_resolver.py` — every rule fires, none over-claims its confidence, and "resolved" keeps meaning one call site with one trustworthy answer.
- `tests/test_bindings.py` — local type inference is *right or absent*: reassignment, real unions and unannotated parameters must all infer nothing.
- `tests/test_cli.py` — a reader and an indexer can hold the database at once.
- `tests/test_docs.py` — the generated demo page is well-formed HTML with balanced tags, which is how the Markdown renderer's crossed-tag bug on `min_confidence` was caught.
---
## Limitations
Stated plainly, because a code-intelligence tool that oversells its precision is worse than useless:
- **Type inference is local only.** Annotations, constructors and aliases, one attribute hop, within a file. No return-type propagation, no generics, no cross-module dataflow. A receiver whose type is never written down — `app.config.get(...)` — stays unresolved, which is most of what remains on untyped Python.
- **Dynamic dispatch is invisible.** `getattr(obj, name)()`, decorator registries, and DI containers do not appear as edges.
- **Cross-language edges are not tracked.** A TypeScript frontend calling a Python endpoint is two disconnected subgraphs.
- **Definitions only, not every reference.** A symbol used as a value (passed as a callback) is weaker in the graph than one that is called.
- **Inheritance is walked, but not linearised.** Base classes are searched breadth-first, which matches Python's C3 order for single inheritance but can pick a different branch of a diamond than the interpreter would.
- **Interface dispatch is not modelled.** A call through a Go interface or an abstract base resolves to no single implementation, and is reported as ambiguous or unresolved rather than fanned out over implementers.
Roadmap: fanning interface calls out over their implementers, Rust and Java adapters, optional LSP enrichment for exact resolution where a language server is available, and a `--changed-since <ref>` mode for PR-scoped blast radius.
---
## Why this exists
I wanted to know whether a coding agent's biggest weakness on large repos — no structural model of the code — could be fixed with static analysis and a well-shaped tool surface rather than a bigger model or a vector database. Mostly, it can.
## License
MIT
TDQS
Scored across 10 tools
Tool purposes are largely distinct and descriptions explicitly route agents to the right one, but find_symbol/get_symbol and who_calls/blast_radius have adjacent responsibilities that could occasionally cause misselection. Overall, the overlap is minor and well-documented.
All names are readable snake_case, but the set mixes verb-object names (find_symbol, search_code, get_symbol), question-style names (who_calls, what_it_calls), and noun-phrase names (blast_radius, file_summary, architecture_overview). This is not chaotic, but it lacks a single consistent naming pattern.
Ten tools is a well-scoped surface for a code-graph analysis server. Each tool addresses a distinct job—search, symbol detail, call trees, impact analysis, overview, index health—without redundancy or bloat.
The toolchain covers symbol discovery, detailed lookup, dependency analysis, impact assessment, file outlining, architecture orientation, and index health, giving strong coverage of the code-understanding workflow. Minor gaps like direct raw-file access or listing all symbols in a file must be worked around via file_summary and get_symbol.