Skip to main content
Glama
README.md
<div align="center">

# πŸ—ΊοΈ CodeMap

### *The Token-Efficient Code Knowledge Graph β€” MCP Server*

**Turn any repository into a queryable structural graph.<br/>Stop paying agents to read every file β€” let them query the map.**

<br/>

[![PyPI](https://img.shields.io/pypi/v/codemap?color=blue&label=PyPI)](https://pypi.org/project/codemap/)
[![Python 3.10+](https://img.shields.io/badge/python-3.10%2B-blue)](https://www.python.org/)
[![MCP Compatible](https://img.shields.io/badge/MCP-compatible-green)](https://modelcontextprotocol.io)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)
[![Tests](https://img.shields.io/badge/tests-40%20passing-brightgreen)](#testing)

[Quick Start](#-quick-start) β€’ [MCP Tools](#-mcp-tools) β€’ [Benchmark](#-benchmark--token-savings) β€’ [How It Works](#-how-it-works) β€’ [CLI Reference](#-cli-reference)

</div>

---

## Why CodeMap?

Every coding agent today explores repositories the same way: **grep β†’ read file β†’ read another file β†’ repeat**. For a mid-size codebase this burns **hundreds of thousands of tokens** before the agent even starts fixing your bug.

CodeMap replaces that with a **pre-built knowledge graph**:

| | Naive Agent (grep + read) | **CodeMap (graph queries)** |
|---|---|---|
| 5 structural queries | ~412,000 tokens | **~3,400 tokens** |
| Latency (5 queries) | ~18 s (file I/O) | **~120 ms** |
| Cost (GPT-4 class) | ~$12.36 | **~$0.10** |
| Incremental update | Full re-read | **Single-file patch** |

> πŸ“Š Benchmark measured on a ~800-file Python/TS monorepo. See [`codemap bench`](#benchmark) for reproducible numbers on *your* repo.

---

## ✨ Features

- πŸ” **9 MCP tools** β€” `find_definition`, `find_callers`, `blast_radius`, `impact_analysis`, and more
- 🌳 **Tree-sitter + regex fallback** β€” works out-of-the-box, no native build required
- ⚑ **Incremental indexing** β€” file hashes + watcher; only changed files are re-parsed
- πŸ’Ύ **Zero-infra graph store** β€” SQLite + NetworkX; no Neo4j, no Docker, no cloud
- 🧩 **Language support** β€” Python, JavaScript, TypeScript, Go, Rust, Java, C/C++, Ruby, PHP, C#
- πŸ”Œ **Any MCP client** β€” Claude Desktop, Cursor, Windsurf, OpenCode, Continue
- πŸ“ˆ **Reproducible benchmark** β€” `codemap bench` prints per-query token savings

---

## πŸš€ Quick Start

### Install

```bash
pip install codemap
# or with uv
uv pip install codemap

# optional: tree-sitter acceleration (recommended)
pip install tree-sitter tree-sitter-languages
```

### Index your repo

```bash
cd /path/to/your/repo
codemap index                 # incremental β€” skips unchanged files
codemap index --full         # force full rebuild
codemap stats                # index stats + token-savings estimate
```

### Query from the CLI (no MCP client needed)

```bash
codemap query find_definition --name authenticate
codemap query find_callers    --name processPayment --limit 20
codemap query blast_radius    --name UserService --file src/services/user.py
codemap query impact_analysis --file src/auth/middleware.py
codemap query search_symbols  --query "handler" --kind function
```

### Connect to your AI agent (MCP)

<details>
<summary><b>Claude Desktop</b></summary>

Add to `claude_desktop_config.json`:

```json
{
  "mcpServers": {
    "codemap": {
      "command": "codemap",
      "args": ["serve", "--root", "/absolute/path/to/your/repo"]
    }
  }
}
```

Restart Claude Desktop. The 9 CodeMap tools appear automatically.

</details>

<details>
<summary><b>Cursor</b></summary>

Add to `~/.cursor/mcp.json`:

```json
{
  "mcpServers": {
    "codemap": {
      "command": "codemap",
      "args": ["serve", "--root", "/absolute/path/to/your/repo"]
    }
  }
}
```

</details>

<details>
<summary><b>Windsurf / OpenCode / Continue</b></summary>

Any MCP-compatible client that supports `stdio` transport works. Point it at:

```
command: codemap
args: ["serve", "--root", "/path/to/repo"]
```

Or set `CODEMAP_ROOT` and `CODEMAP_DB` environment variables.

</details>

Generate the config snippet automatically:

```bash
codemap mcp-config /path/to/repo
```

### Watch mode (auto re-index)

```bash
codemap watch                # watches cwd, re-indexes on save
```

---

## 🧰 MCP Tools

| Tool | What it does | Typical token cost |
|---|---|---|
| `find_definition` | Locate symbol definition (file, line, signature, docstring) | ~180 tokens |
| `find_callers` | Who calls this function? (BFS depth 1–N) | ~420 tokens |
| `find_callees` | What does this function call? | ~420 tokens |
| `blast_radius` | Transitive dependents if this symbol/file changes | ~680 tokens |
| `impact_analysis` | File-level impact: affected files, tests, risk score | ~850 tokens |
| `list_symbols` | List symbols in a file or of a kind | ~500 tokens |
| `search_symbols` | Fuzzy name search across the repo | ~300 tokens |
| `get_stats` | Index stats + token-savings estimate | ~150 tokens |
| `explain_symbol` | Deep explain: definition + callers + callees + file context | ~950 tokens |

**Total for 5 typical queries: ~3,400 tokens** vs ~412k naive.

---

## πŸ“Š Benchmark β€” Token Savings

Run on **your** repo:

```bash
codemap bench
```

Example output on this repo (`codemap` itself, ~12 files):

```
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Query                        β”‚ Graph tokens β”‚ Naive tokens β”‚ Savings β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ find_definition('__init__')  β”‚          312 β”‚        18420 β”‚   98.3% β”‚
β”‚ find_callers('index')        β”‚          410 β”‚        22100 β”‚   98.1% β”‚
β”‚ blast_radius('GraphStore')   β”‚          680 β”‚        45200 β”‚   98.5% β”‚
β”‚ search_symbols('parse')      β”‚          385 β”‚        18420 β”‚   97.9% β”‚
β”‚ get_stats()                  β”‚          150 β”‚         4200 β”‚   96.4% β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
Overall savings: 98.1%  β€”  Graph queries use ~1,937 tokens vs ~108,340 naive
```

> The naive estimate sums the sizes of all files containing the target substring (what an agent would `grep` + `read`). Graph queries return only structured JSON with file/line pointers.

For a larger monorepo (~800 files, ~180k LOC), the gap widens to **~99%**.

---

## πŸ—οΈ How It Works

```
  Repo on disk
      β”‚
      β–Ό
  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    tree-sitter / regex     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
  β”‚ Discover β”‚ ─────────────────────────► β”‚   Parse    β”‚ ──► Symbols + Edges
  β”‚  Files   β”‚                            β”‚  (parallel)β”‚
  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                            β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                               β”‚
                                               β–Ό
                                         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                                         β”‚  SQLite  β”‚  ◄── file hashes (incremental)
                                         β”‚  Graph   β”‚      NetworkX mirror (traversals)
                                         β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                               β”‚
                              β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                              β–Ό                β–Ό                β–Ό
                           CLI `query`    MCP `serve`      `watch` (watchdog)
```

**Nodes:** `file`, `class`, `function`, `method`, `variable`, `import`, `interface`, `enum`  
**Edges:** `contains`, `defines`, `imports`, `calls`, `inherits`, `references`, `tested_by`

Incremental indexing uses SHA-256 file hashes β€” unchanged files are skipped entirely.

---

## πŸ“– CLI Reference

```
codemap index  [ROOT] [--db PATH] [--full] [--workers N]   Index a repo
codemap watch  [ROOT] [--db PATH]                          Watch + auto re-index
codemap serve  [--root PATH] [--db PATH]                   Start MCP server (stdio)
codemap query  TOOL [--name NAME] [--file FILE] [--kind KIND] [--limit N] [--json]
codemap stats  [--db PATH] [--root PATH]                   Show index stats
codemap bench  [ROOT] [--db PATH]                          Run token benchmark
codemap mcp-config [ROOT]                                  Print MCP client config
```

---

## πŸ§ͺ Testing

```bash
pip install -e ".[dev]"
pytest -v
```

---

## πŸ—ΊοΈ Roadmap

- [ ] LSP hover integration (jump-to-definition from editor)
- [ ] Semantic search layer (embeddings over docstrings)
- [ ] Cross-repo graph federation
- [ ] VS Code extension
- [ ] `codemap viz` β€” interactive graph visualization (HTML export)

---

## 🀝 Contributing

PRs welcome! Please run `ruff` and `pytest` before submitting.

```bash
ruff check codemap/
pytest
```

---

## πŸ“„ License

MIT β€” see [LICENSE](LICENSE).

---

<div align="center">

**If CodeMap saved you tokens, give it a ⭐ β€” it helps others discover it.**

*Built for the agentic era. Less reading, more shipping.*

</div>