Skip to main content
Glama
README.md
# repopedia

**Code is a graph — calls, dependencies, inheritance — but today's AI doc tools pretend it's text.** repopedia recovers the structure first (AST → knowledge graph), then grows answers, wikis, and impact analysis from the graph. The graph says what's true (every claim cites file:line); the LLM makes it readable.

repopedia is MIT-licensed, local-first, and dependency-light: `pip install` and index any repo. No Docker, no server, no vendor lock-in.

## Why a graph, not chunks

Naive code understanding chunks source files, embeds them, and retrieves the top-k chunks most similar to a question. That works for *"where is the retry logic?"* — the answer sits in one chunk. It breaks on structural questions like:

> *If I change `Engine.run`, what breaks?*

No single chunk contains the answer. You need the call graph: who calls `run`, who calls *them*, and so on. Vector similarity can't do graph traversal; a knowledge graph can. repopedia builds that graph with tree-sitter (precise, language-aware parsing — not regexes), then answers from it.

## Architecture

```mermaid
flowchart LR
    A[Source repo<br/>.py / .ts / .tsx] --> B[tree-sitter<br/>extractors]
    B --> C[FileFacts<br/>symbols · calls · imports · inherits]
    C --> D[indexer<br/>cross-file resolution]
    D --> E[(SQLite graph<br/>nodes + edges)]
    E --> F[queries<br/>blast radius · incremental]
    E --> G[MCP server<br/>7 tools]
    E --> H[wiki generator<br/>pages + diagrams from graph]
```

## Querying the graph

```bash
cd ./myrepo
repopedia query callers pkg.mod.Service.run     # who calls it (file:line)
repopedia query callees pkg.mod.Service.run     # what it calls (unresolved shown as <unresolved: raw>)
repopedia query inherits pkg.mod.Service        # base classes, nearest first
repopedia query file pkg/mod.py                 # symbols defined in a file
repopedia blast-radius pkg.mod.Service.run --depth 3
repopedia search "retry handler"                 # BM25 + one graph hop
```

`blast-radius` answers *"if I change this, what breaks?"*: transitive
reverse `calls` traversal up to `--depth`, plus every file that imports
the symbol's file. Each hit cites `file:line` and how it was reached
(`calls(2)`, `imports`…). Add `--json` to any command for machine-readable
output — the same functions back the Week 3 MCP server.

Python API (all returns are plain JSON-serializable dicts/lists):

```python
from repopedia.store import open_store
from repopedia import query

with open_store("myrepo/.repopedia/graph.db") as s:
    for hit in query.blast_radius(s, "pkg.mod.Service.run", depth=3):
        print(hit["depth"], hit["via"], hit["qualified_name"],
              f"{hit['file']}:{hit['line_start']}")
```

## Incremental updates

```bash
repopedia update ./myrepo
# updated: +1 added, ~2 modified, -0 deleted
```

`update` diffs against the commit the graph was built from (stored as
`head_sha` in the DB) and re-extracts only changed files, re-resolving
their edges against the full store. Files that pointed *into* a changed
file (callers, importers) get their references repaired too. Non-git
directories fall back to a full reindex with a warning. One honest
limitation: previously *unresolvable* edges in files that point at
nothing changed stay unresolved until a full `repopedia index`.

## MCP server: the graph for AI coding agents

The graph's second consumer is not a human — it's your coding agent.
`repopedia mcp` serves the query API over the Model Context Protocol
(stdio), so Cursor, Claude Code, Windsurf, or any MCP client can traverse
the codebase structurally instead of guessing from text search.

```bash
repopedia mcp --repo ./myrepo
# repopedia MCP server: ./myrepo/.repopedia/graph.db (stdio)
```

The LLM is optional: only `ask_codebase` synthesizes prose, and only when
`REPOPEDIA_LLM_BASE_URL` / `REPOPEDIA_LLM_API_KEY` / `REPOPEDIA_LLM_MODEL`
are set (any OpenAI-compatible endpoint). Every other tool works
offline, straight from the graph.

| Tool | What it does |
|---|---|
| `find_symbol` | symbols by qualified/short name, each citing `file:line` |
| `get_callers` / `get_callees` | one `calls` hop, each way (unresolved sites marked) |
| `blast_radius` | transitive reverse calls + importing files (`depth` param) |
| `search_codebase` | BM25 + one graph hop, ranked with `via: lexical\|graph` |
| `get_file_symbols` | everything a file defines |
| `ask_codebase` | retrieval + grounded synthesis; without an LLM it returns the evidence block and says so |

Client configuration:

```bash
# Claude Code CLI
claude mcp add repopedia -- repopedia mcp --repo /path/to/repo
```

```jsonc
// Cursor / Windsurf / Claude Desktop — mcpServers
{
  "repopedia": {
    "command": "repopedia",
    "args": ["mcp", "--repo", "/path/to/repo"]
  }
}
```

## Wiki generator: documentation derived from the graph

```bash
repopedia wiki ./myrepo --out docs/
# wiki written to ./myrepo/docs
#   architecture.md
#   index.md
#   modules/auth.md
#   ...
```

This is the anti-DeepWiki move: instead of an LLM reading text chunks and
writing plausible prose, pages are **generated from graph edges** —
module tables from `defines`, dependency lists from `imports`, the
architecture diagram from the actual import graph, "most-called functions"
from real call counts. Every structural claim cites `file:line`; Mermaid
diagrams that exceed the node cap say so on the page.

Before generating, repopedia compares the graph's `head_sha` against git
HEAD. If the graph is stale, it prints a warning *and* embeds it in
`index.md` — a wiki that might be outdated says so loudly instead of
lying quietly. (Run `repopedia update` first; note that previously
unresolved references in unchanged files only heal on a full reindex.)

With an LLM configured (same env vars as above), `index.md` and module
pages get short prose summaries synthesized under a strict
cite-only-what-you're-given prompt. Without one, you get a deterministic
structural wiki — tables plus diagrams, no prose claims — which is
genuinely useful on its own: it answers "what lives where" with
citations. Same graph → byte-identical markdown, every time.

## Search: BM25 + graph, no vectors (deliberate)

Code identifiers are not natural language. `"where is retry handled"`
is answered better by matching the symbol `retry` / `handle_retry` and
then *walking the graph* (who calls it, what file it's in) than by
embedding the query into a vector space trained on prose. So
`repopedia search` ranks symbols with a hand-rolled BM25 (~40 lines, zero
dependencies) over qualified names, file paths, and kinds, then expands
one hop along `calls` / `inherits` / `imports` to pull in structurally
related symbols — marked `"via": "graph"` in the output. Structure beats
embeddings for "where" questions; vectors may come later as a supplement.

## Quickstart

```bash
pip install repopedia        # or: pip install -e .  (from source)
repopedia index ./myrepo
# indexed ./myrepo
#   db:      ./myrepo/.repopedia/graph.db
#   files:   132 parsed, 1 skipped
#   symbols: 841
#   edges:   2130
```

Limit to one language, or choose the DB location:

```bash
repopedia index ./myrepo --language py
repopedia index ./myrepo --db /tmp/graph.db
```

Python API:

```python
from repopedia.store import open_store

with open_store("myrepo/.repopedia/graph.db") as s:
    for sym in s.find_symbol("Engine.run"):
        print(sym["qualified_name"], sym["file"], sym["line_start"])
    for e in s.in_edges(sym["id"], kind="calls"):   # who calls Engine.run?
        print("called by:", e["src_qualified"], e["src_file"])
```

## What gets extracted

| Language | Extensions | Symbols | Edges |
|----------|-----------|---------|-------|
| Python | `.py` | classes, functions, methods (incl. decorated/async/nested) | `defines`, `imports` (abs + relative), `calls`, `inherits` |
| TypeScript | `.ts`, `.tsx` | classes, functions, methods (incl. exported) | `defines`, `imports` (relative), `calls` (+ `new`), `inherits` (`extends`) |

Skipped, never fatal: `node_modules/`, `.git/`, `venv/`, `__pycache__/`, `dist/`, `build/`, non-UTF8 files, and files with syntax errors (logged as warnings — tree-sitter error recovery is deliberately *not* trusted for partial extraction).

### Graph schema

```sql
nodes(id, kind, name, qualified_name, file, line_start, line_end, language)
-- kind: file | class | function | method
-- qualified_name: dotted path, e.g. "pkg.mod.Service.run"
-- file: repo-relative path; line_* are 1-based inclusive

edges(src, dst, kind, data)
-- kind: defines | imports | calls | inherits
-- dst is NULL when the reference can't be resolved in-repo
-- data (JSON) always keeps the raw written form:
--   imports  -> {"module": "pkg.util"}
--   calls    -> {"raw": "helper"} (+ "candidates": N when ambiguous)
--   inherits -> {"raw": "Base"}
```

**Resolution rules (heuristic, documented):**
- *calls*: resolved by short name when unambiguous in the repo. `self.x` / `this.x` match method `x`; `ClassName(...)` resolves to the class node (constructor call). Ambiguous or unknown names stay `dst=NULL` with the raw name preserved for later phases to refine.
- *imports*: resolved via normal module-resolution rules (`pkg/util.py`, `__init__.py`, `./util` → `util.ts(x)`/`index.ts(x)`). Stdlib, third-party, and bare specifiers stay unresolved with the module string recorded.
- *inherits*: resolved when the base class name is unique in the repo.

The schema is stable: Week 2 (query API, blast-radius = reverse `calls` traversal, git-diff incremental reindex) and Week 3 (MCP tools, wiki pages + Mermaid diagrams generated *from* the graph with file:line citations) build on exactly these tables.

## Roadmap

- **Week 1** ✅: tree-sitter indexer (Python + TypeScript), SQLite store, CLI
- **Week 2** ✅: query API (callers/callees, blast radius, inheritance, file symbols), git-diff incremental reindex with reference repair, BM25 + graph hybrid search
- **Week 3** ✅: MCP server (7 tools: find_symbol, get_callers/callees, blast_radius, search_codebase, get_file_symbols, ask_codebase with optional grounded LLM synthesis), wiki generator (index + per-module pages + architecture, all derived from the graph with file:line citations, deterministic output, staleness warnings)
- **Later**: more languages (Go, Java, Rust), vector embeddings as an optional supplement, hosted demo

## What repopedia is NOT

- **Not a vector RAG clone.** No embeddings in MVP — deliberately. Structure first, vectors later as a supplement.
- **Not a DeepWiki clone.** DeepWiki-style tools generate prose from text chunks; repopedia generates *from the graph*. Different foundation, different accuracy properties: our wiki pages are only as wrong as the parser, never as wrong as a hallucinating LLM.
- **Not an IDE.** The graph is built for two consumers: humans reading generated wikis, and AI coding agents querying via MCP.
- **Not magic.** Unresolved references (`dst=NULL`) are shown, not hidden. Stale graphs warn loudly. The LLM is instructed to cite only what the graph gave it — and when no LLM is configured, repopedia says so instead of faking synthesis.

## License

MIT — use it at work, ship it in products, no strings attached. See [LICENSE](LICENSE).