Skip to main content
Glama

🗺️ CodeMap

The Token-Efficient Code Knowledge Graph — MCP Server

Turn any repository into a queryable structural graph.Stop paying agents to read every file — let them query the map.

PyPI Python 3.10+ MCP Compatible License: MIT Tests

Quick Start • MCP Tools • Benchmark • How It Works • CLI Reference


Why CodeMap?

Every coding agent today explores repositories the same way: grep → read file → read another file → repeat. For a mid-size codebase this burns hundreds of thousands of tokens before the agent even starts fixing your bug.

CodeMap replaces that with a pre-built knowledge graph:

Naive Agent (grep + read)

CodeMap (graph queries)

5 structural queries

~412,000 tokens

~3,400 tokens

Latency (5 queries)

~18 s (file I/O)

~120 ms

Cost (GPT-4 class)

~$12.36

~$0.10

Incremental update

Full re-read

Single-file patch

📊 Benchmark measured on a ~800-file Python/TS monorepo. See codemap bench for reproducible numbers on your repo.


Related MCP server: codebase-rag

✨ Features

  • 🔍 9 MCP tools — find_definition, find_callers, blast_radius, impact_analysis, and more

  • 🌳 Tree-sitter + regex fallback — works out-of-the-box, no native build required

  • ⚡ Incremental indexing — file hashes + watcher; only changed files are re-parsed

  • 💾 Zero-infra graph store — SQLite + NetworkX; no Neo4j, no Docker, no cloud

  • 🧩 Language support — Python, JavaScript, TypeScript, Go, Rust, Java, C/C++, Ruby, PHP, C#

  • 🔌 Any MCP client — Claude Desktop, Cursor, Windsurf, OpenCode, Continue

  • 📈 Reproducible benchmark — codemap bench prints per-query token savings


🚀 Quick Start

Install

pip install codemap
# or with uv
uv pip install codemap

# optional: tree-sitter acceleration (recommended)
pip install tree-sitter tree-sitter-languages

Index your repo

cd /path/to/your/repo
codemap index                 # incremental — skips unchanged files
codemap index --full         # force full rebuild
codemap stats                # index stats + token-savings estimate

Query from the CLI (no MCP client needed)

codemap query find_definition --name authenticate
codemap query find_callers    --name processPayment --limit 20
codemap query blast_radius    --name UserService --file src/services/user.py
codemap query impact_analysis --file src/auth/middleware.py
codemap query search_symbols  --query "handler" --kind function

Connect to your AI agent (MCP)

Add to claude_desktop_config.json:

{
  "mcpServers": {
    "codemap": {
      "command": "codemap",
      "args": ["serve", "--root", "/absolute/path/to/your/repo"]
    }
  }
}

Restart Claude Desktop. The 9 CodeMap tools appear automatically.

Add to ~/.cursor/mcp.json:

{
  "mcpServers": {
    "codemap": {
      "command": "codemap",
      "args": ["serve", "--root", "/absolute/path/to/your/repo"]
    }
  }
}

Any MCP-compatible client that supports stdio transport works. Point it at:

command: codemap
args: ["serve", "--root", "/path/to/repo"]

Or set CODEMAP_ROOT and CODEMAP_DB environment variables.

Generate the config snippet automatically:

codemap mcp-config /path/to/repo

Watch mode (auto re-index)

codemap watch                # watches cwd, re-indexes on save

🧰 MCP Tools

Tool

What it does

Typical token cost

find_definition

Locate symbol definition (file, line, signature, docstring)

~180 tokens

find_callers

Who calls this function? (BFS depth 1–N)

~420 tokens

find_callees

What does this function call?

~420 tokens

blast_radius

Transitive dependents if this symbol/file changes

~680 tokens

impact_analysis

File-level impact: affected files, tests, risk score

~850 tokens

list_symbols

List symbols in a file or of a kind

~500 tokens

search_symbols

Fuzzy name search across the repo

~300 tokens

get_stats

Index stats + token-savings estimate

~150 tokens

explain_symbol

Deep explain: definition + callers + callees + file context

~950 tokens

Total for 5 typical queries: ~3,400 tokens vs ~412k naive.


📊 Benchmark — Token Savings

Run on your repo:

codemap bench

Example output on this repo (codemap itself, ~12 files):

┌──────────────────────────────┬──────────────┬──────────────┬─────────┐
│ Query                        │ Graph tokens │ Naive tokens │ Savings │
├──────────────────────────────┼──────────────┼──────────────┼─────────┤
│ find_definition('__init__')  │          312 │        18420 │   98.3% │
│ find_callers('index')        │          410 │        22100 │   98.1% │
│ blast_radius('GraphStore')   │          680 │        45200 │   98.5% │
│ search_symbols('parse')      │          385 │        18420 │   97.9% │
│ get_stats()                  │          150 │         4200 │   96.4% │
└──────────────────────────────┴──────────────┴──────────────┴─────────┘
Overall savings: 98.1%  —  Graph queries use ~1,937 tokens vs ~108,340 naive

The naive estimate sums the sizes of all files containing the target substring (what an agent would grep + read). Graph queries return only structured JSON with file/line pointers.

For a larger monorepo (~800 files, ~180k LOC), the gap widens to ~99%.


🏗️ How It Works

  Repo on disk
      │
      ▼
  ┌──────────┐    tree-sitter / regex     ┌────────────┐
  │ Discover │ ─────────────────────────► │   Parse    │ ──► Symbols + Edges
  │  Files   │                            │  (parallel)│
  └──────────┘                            └────────────┘
                                               │
                                               ▼
                                         ┌──────────┐
                                         │  SQLite  │  ◄── file hashes (incremental)
                                         │  Graph   │      NetworkX mirror (traversals)
                                         └──────────┘
                                               │
                              ┌────────────────┼────────────────┐
                              ▼                ▼                ▼
                           CLI `query`    MCP `serve`      `watch` (watchdog)

Nodes: file, class, function, method, variable, import, interface, enum
Edges: contains, defines, imports, calls, inherits, references, tested_by

Incremental indexing uses SHA-256 file hashes — unchanged files are skipped entirely.


📖 CLI Reference

codemap index  [ROOT] [--db PATH] [--full] [--workers N]   Index a repo
codemap watch  [ROOT] [--db PATH]                          Watch + auto re-index
codemap serve  [--root PATH] [--db PATH]                   Start MCP server (stdio)
codemap query  TOOL [--name NAME] [--file FILE] [--kind KIND] [--limit N] [--json]
codemap stats  [--db PATH] [--root PATH]                   Show index stats
codemap bench  [ROOT] [--db PATH]                          Run token benchmark
codemap mcp-config [ROOT]                                  Print MCP client config

🧪 Testing

pip install -e ".[dev]"
pytest -v

🗺️ Roadmap

  • LSP hover integration (jump-to-definition from editor)

  • Semantic search layer (embeddings over docstrings)

  • Cross-repo graph federation

  • VS Code extension

  • codemap viz — interactive graph visualization (HTML export)


🤝 Contributing

PRs welcome! Please run ruff and pytest before submitting.

ruff check codemap/
pytest

📄 License

MIT — see LICENSE.


If CodeMap saved you tokens, give it a ⭐ — it helps others discover it.

Built for the agentic era. Less reading, more shipping.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    Provides a semantic understanding of your codebase by parsing with tree-sitter and building a graph of symbols and dependencies. Enables AI assistants to navigate code, analyze changes, and discover architecture using 18 tools with minimal context overhead.
    18 npm
    1
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables LLM agents to efficiently understand and navigate a codebase by providing semantic search over symbols and a reference graph, replacing expensive grep/glob calls with structured tools like definition lookup, caller/callee queries, and change-impact analysis.
    3
    MIT