Skip to main content
Glama

repo2graph

Give coding agents trustworthy, cited answers about unfamiliar codebases.

Ask a repository a question; get back the actual source that answers it, every block stamped with the file and line range it came from.


⚡ What is repo2graph?

An agent dropped into a codebase it has never seen has two bad options. Grep for a word and it either floods its context with whole matching files, or finds nothing because the code spells the idea differently than you did. Guess from training data and it writes something confident and wrong. Either way you cannot tell which of the two just happened.

repo2graph answers questions about a repository with the repository's own source. Ask "how does a request get authenticated" and you get back the function that does it, the functions that call it and the ones it calls — each block headed [cite: path:start-end], so every claim in the answer is one click from the line it came from. If the answer is wrong, the citation shows you where it went wrong. That is the whole point.

It gets there by reading the code rather than searching it: one parse pass records who calls whom, who imports what, and which class extends which, and retrieval follows those links instead of matching more text. The result is served straight into Claude Code, Cursor or any Model Context Protocol client, or packed into a markdown context with a hard token ceiling for any other LLM.

flowchart LR
    A[your code] --> B[tree-sitter<br/>reads the code]
    B --> C[graph<br/>dots + arrows]
    C --> D[graph.html<br/>the picture]
    C --> E[chunks.jsonl<br/>pieces for an AI]
    C -->|MCP stdio| F[Claude / Cursor /<br/>any MCP client]

No project setup, no language server, no build step — point it at a folder and it works.

Interactive canvas, zoomed

Filter & inspector controls

graph.html is one self-contained file — no server, no internet, drag to pan, scroll to zoom, click a node to inspect its code and neighbours.

Related MCP server: Lore MCP Server

👥 Who it's for

🧭 Joining a new codebase

The problem: week one goes on reading files to find out which ones matter.

Build once, open the map, and start from the hub files instead of the root directory. Then ask whole questions — "how does a request get from the router to the handler" — and read the answer as source, with the callers and callees already attached.

uvx repo2graph build . -o .r2g
open .r2g/human/graph.html
repo2graph rag "how does routing work" -o .r2g

🤖 Driving a coding agent

The problem: the agent greps, pulls in three whole files, and still edits the wrong one.

Point Claude Code, Cursor or any MCP client at the repo. The agent gets cited blocks under a hard 12k-token ceiling instead of raw file dumps, and can walk from a symbol to its callers in one hop. Secrets are excluded from agent replies unconditionally — no flag turns that off.

claude mcp add repo2graph -- \
  uvx --from "repo2graph[mcp]" repo2graph-mcp .

🔍 Reviewing a pull request

The problem: the diff is 40 lines; the blast radius is unknown.

Ask the graph what touches the changed symbol — callers, importers, subclasses — and what the repository's own history says usually changes alongside it (CO_CHANGE, mined from git). That last one catches the test file or the config the diff forgot.

repo2graph build . -o .r2g --git-history 500
repo2graph explain node "sym:src/auth.py::verify" -o .r2g

🌱 Maintaining a project

The problem: every new contributor asks the same "where do I start" question.

Commit a fresh graph on every push with the GitHub Action, and publish graph.html to a branch contributors can browse. The job summary reports hub files, co-change hotspots and the graph delta since the last build, so architectural drift shows up in the run.

- uses: Srinivasan-78/repo2graph@v2
  with: { git-history: "500", commit-branch: graph }

Both of those are still in the box — repo2graph seeds every query with BM25, and dense vectors are an opt-in fusion. The difference is what happens after the first match.

grep / ripgrep

Embedding search

repo2graph

Finds

the exact string

text that reads similarly

the symbol, then everything wired to it

Different words than the code uses

returns nothing

handles it

BM25 seeds, then graph hops reach code the query never named

"What calls this?"

can't answer — a match in a comment ranks like the definition

can't answer — neighbours aren't in the embedding

CALLS edges, with direction and a confidence score

"What breaks if I change this?"

you read every hit by hand

not represented

callers, importers and subclasses in one hop

What comes back

matching lines, or whole files an agent then dumps into context

top-k similar chunks, callers unretrieved

the source that answers it, each block headed [cite: path:start-end]

Token cost

unbounded — the agent decides how much file to read

unbounded

hard ceiling on the whole pack, re-measured before it returns

"Which files keep changing together?"

—

—

CO_CHANGE, mined from git history

Setup

none

index build + an embedding model (~90 MB)

one parse pass, no model, no API key, no language server

Ranking is explainable

n/a

a cosine number

repo2graph explain retrieval "<q>" names the seed and the edge that pulled each block in

Use grep when you want every occurrence of a literal string — a config key, an error message, a TODO. repo2graph has no special knowledge of string literals and will not beat it. Use repo2graph when the question is about relationships: what calls this, what breaks if I change it, how does data get from A to B. Longer version: docs/why-graph.md.

🚀 Four ways to run repo2graph

Same graph, same chunk format, same .r2g output — pick the interface for where you're standing right now.

Local dev, scripting, ad-hoc questions from a terminal.

Jump in ↓

A fresh graph committed next to your code on every push, zero Python setup.

Jump in ↓

Give Claude, Cursor or any MCP client live, cited access to the codebase.

Jump in ↓

Enterprise-ready, read-only, non-root container deployment.

Jump in ↓

🐍 1. Python / CLI

Requires Python 3.10+. The fastest way to see what it does — a bundled example repo, indexed and interrogated, with nothing to configure and no repository of your own:

uvx repo2graph demo

That writes a small orders service to a scratch directory, builds a real graph over it, and answers five questions against it, each one citing files and line ranges. Then point it at your own code:

uvx repo2graph build . -o .r2g && open .r2g/human/graph.html

Or install it properly:

pip install repo2graph
repo2graph build /path/to/project -o .r2g --git-history 200
repo2graph query "how does routing match a path" -o .r2g

Two minutes end to end, with the expected output at each step: docs/quickstart.md.

Five questions to start with

The ones that show what a graph gives you over a text search. Copy any of them onto your own repo — or run repo2graph demo to watch each answered against the bundled fixture.

Ask your repo

What comes back that grep cannot give you

Where is authentication enforced?

the guard itself, plus every route that calls it

What calls <function>?

CALLS edges in, so callers come back even when the name is shadowed

What tests cover <module>?

IMPORTS edges from the test module back to the code under test

What would be affected by changing <api>?

the blast radius: direct callers and what they are called from

Trace <a request> from route to persistence.

a whole path across modules, each block cited to file and line

repo2graph rag "What would be affected by changing pack_context?" -o .r2g

Something not working? repo2graph doctor . checks the environment, the index and your MCP client config, and prints a fix for anything it finds.

One pass over this repository — 195 files, 2,552 nodes, 11,118 edges — takes about three seconds and needs no configuration file, no language server and no API key. Ask it something, and the answer comes back as source you can check, not a summary you have to trust:

Full flag tables, budget accounting and the Python API: docs/cli.md · docs/python-api.md.

⚙️ 2. GitHub Action

Published on the GitHub Marketplace — one step, no Python setup on the runner:

- uses: actions/checkout@v4
  with: { fetch-depth: 0 }   # full history, so CO_CHANGE edges are meaningful

- uses: Srinivasan-78/repo2graph@v2
  with:
    path: .                  # or: repo: some-org/other-repo
    git-history: "500"       # commits scanned for CO_CHANGE edges (0 = skip)
    artifact-name: repo-graph

@v2 follows every 2.x release; pin an exact tag (@v2.2.0) to upgrade by hand instead. It never calls an LLM — --answer is deliberately not exposed — and it writes a job-summary table (hub files, CO_CHANGE hotspots, the graph delta since the last build) straight from the artifacts, so the shape of the map shows up in the run without downloading anything.

Also pack a cited context for a fixed question, and push the map to a browsable branch:

- uses: Srinivasan-78/repo2graph@v2
  with:
    query: "how does auth middleware validate a token"
    commit-branch: graph     # force-pushed; this repo's own /graph branch is built this way

All inputs/outputs, private-repo tokens and the vector-embedding step: docs/github-action.md.

🔌 3. MCP server

repo2graph-mcp is a stdio MCP server. It builds its own index on the first call if one doesn't exist yet — nothing to run ahead of time.

Claude Code

claude mcp add repo2graph -- uvx --from "repo2graph[mcp]" repo2graph-mcp /path/to/project

Claude Desktop (claude_desktop_config.json) and Cursor (.cursor/mcp.json) — same block:

{
  "mcpServers": {
    "repo2graph": {
      "command": "uvx",
      "args": ["--from", "repo2graph[mcp]", "repo2graph-mcp", "/path/to/project"]
    }
  }
}

Any other stdio-based MCP client (Windsurf, Zed, generic clients) takes the same command/args pair — see docs/mcp.md for config file locations per platform and client.

That is the hop grep cannot do: one symbol in, and its definer, its callers and its callees come back with file and line — the relationship, not a text match that happens to contain the name.

🐳 4. Docker

For enterprise and shared deployments, an official Dockerfile is provided. It's a multi-stage build running as a non-root user (10000:10000), fully compatible with a read-only root filesystem and dropped capabilities.

docker build -t repo2graph .
docker run --rm \
  --read-only \
  --cap-drop=ALL \
  --security-opt=no-new-privileges \
  --network=none \
  -v /path/to/repo:/repo:ro \
  -v repo2graph-index:/repo/.r2g \
  repo2graph build /repo -o /repo/.r2g

See docs/ENTERPRISE_DEPLOYMENT.md for full container hardening and HTTP server instructions, and docs/deployment-security.md for the trust boundary and supported/not-recommended verdict per deployment shape — including whether HTTP without TLS is safe (short answer: only on loopback) and a worked hardened reverse-proxy example.

✨ Key features

Deterministic graph, not embeddings-only search

Callers, callees, imports and class hierarchies resolved from the actual AST — not a nearest-neighbour guess.

Hybrid retrieval

BM25 + graph-neighbour expansion by default; optional dense vector fusion (repo2graph embed) with zero required extra dependencies.

Hard token ceilings, enforced twice

pack_context()'s budget bounds the entire rendered markdown, not just chunk text — and the MCP server clamps and re-measures before returning.

17 grammars, full treatment

Python, JS, TS, TSX, Go, Rust, Java, Ruby, C, C++, C#, PHP, Kotlin, Swift, Scala, Bash and Lua get functions/classes/calls — 29 file extensions in all. Everything else still appears as files on the map. See our strategic language scorecard and framework relationship RFC.

CI-native

Published as a GitHub Action — commit a fresh graph next to your code on every push.

Local by default

build, query, rag and the MCP server over stdio make zero network calls — asserted by socket-level tests. rag --answer is the only path that ever sends your code anywhere, and it prints the provider + hostname first. No telemetry.

Export to real graph tooling

graph.graphml (yEd, Gephi, NetworkX) and graph.cypher (Neo4j, Memgraph) come out of every build, no extra step.

⚖️ What it does — and what it does not

A retrieval tool that oversells itself is worse than no retrieval tool, because you stop checking its answers. So, plainly:

It does

  • Return the source that answers a question, cited to path:start-end, inside a token budget it enforces rather than requests.

  • Resolve callers, callees, imports and class hierarchies from a real parse of the code, and let you walk them in either direction from any symbol.

  • Mine CO_CHANGE from git history — the files that keep being edited together, which no parser can tell you.

  • Run entirely locally, with no model, no account and no network call, in the CLI, in CI and over MCP.

  • Degrade gracefully: an unparsed language still appears as file nodes and is still retrievable as text; a missing vector index falls back to BM25 rather than failing.

It does not

Limitation

What that means in practice

Resolve calls by type

Calls are matched by name, with same-class / same-file / import scoping to break ties. When scoping can't isolate one target, the call fans out to up to 5 candidate edges at confidence = 1/n, flagged ambiguous. Filter to confidence == 1.0 when you need certainty over recall — 4.6%–21% of CALLS edges are ambiguous across our five benchmark repos.

See dynamic dispatch

A string-keyed lookup, a plugin registry, getattr-style dispatch, a virtual call resolved at runtime — none of it is written down as syntax, so no edge is drawn. No arrow does not prove no call.

See reflection or computed imports

importlib.import_module(name), Java reflection, a dynamic import() with a computed specifier. Nothing literal to resolve, so nothing to link.

Follow dependency injection to the implementation

A DI container wires an interface to a concrete class at runtime. The call site names the interface method, so the edge lands on the declaration (or fans out across every same-named implementation), never on the class the container actually injected. Walk INHERITS to enumerate the candidates.

Distinguish generated code

A .pb.go, a bundled .js, a codegen'd client — all indexed exactly like hand-written code, with no marker. They can dominate a symbol count without representing a line anyone maintains. Exclude them with --exclude.

Notice that your files changed

The index is a snapshot of the tree you built it from. Nothing watches the filesystem: edit a file and the graph keeps describing the old one. Rebuild (build --incremental re-parses only what moved), or let the GitHub Action rebuild on every push. repo2graph doctor checks index integrity and vector drift — not whether your working tree moved on.

Cross a language boundary

Python calling into C++ through generated bindings becomes a CALLS_EXTERNAL edge, not a link to the C++ function. That is a structural limit of source-only analysis, not a matching bug.

Parse macro-heavy C/C++ cleanly

tree-sitter emits ERROR nodes around unexpanded macros; a cpp preprocessor fallback recovers some. Expect a non-trivial parse_errors count in stats.json and read it as a floor on missed symbols.

Every one of these is measured, not asserted — the rates, the repositories they were measured on, and the reproduction commands are in docs/limitations.md. See also the dynamic pattern evaluation in BENCHMARK.md and the framework relationship proposal in docs/rfcs/rfc-framework-relationship-graph.md.

🆚 How it compares to other graph tools

Several tools build a graph out of a codebase. The thing that separates them is what comes back when you ask a question — a picture, a subgraph, or the code itself.

repo2graph

Graphify

Code Graph (Obsidian)

grep / embedding RAG

What a query returns

the source, packed — every block headed [cite: path:start-end]

a scoped subgraph, a path, or a concept explanation to traverse

a force-directed picture to read

matching lines, or nearest-neighbour chunks

How hits are ranked

BM25 seeds, then k-hop graph expansion; optional dense fusion

graph traversal (explicitly not a vector index)

n/a — it is a view

lexical only, or vectors only

Token budget

hard cap on the whole pack, re-measured before returning (12k ceiling over MCP)

not a packing layer

n/a

usually unbounded

Edges from git history

CO_CHANGE, from --git-history

—

—

—

Runs with no assistant, no model, no account

yes — CLI, MCP, or the GitHub Action

code pass is local; the docs/media pass uses a model

needs Obsidian desktop 1.7.2+

varies

Corpus

code in 17 parsed grammars, every other file as text

code in ~40 languages, plus docs, PDFs, images, video

TS/TSX/JS/Python parsed, imports-only for 8 more

anything

Reach for Graphify when the graph itself is the product: community detection, shortest path between two concepts, and your PDFs and design docs in the same graph as the code. Reach for the Obsidian plugin when a human wants to read the graph beside their notes. Reach for repo2graph when an agent needs cited source inside a fixed token budget, when it has to run in CI with no model and no account, or when "which files keep changing together" is part of the answer.

Longer version, with the trade-offs each choice implies: docs/comparison.md.

🛠️ MCP tools exposed

Five tools. Three answer questions about the code; two report on the server itself.

Tool

Arguments

What comes back

repo_map

none

Languages, hub files, and top entry points. Stable across calls — read this first.

repo_search

query, optional k (default 8, max 50), hops (default 1, max 4), budget_tokens (default 6000, max 12000)

Seed chunks plus graph neighbours, each block headed [cite: path:start-end].

repo_neighbours

node_id, optional hops (default 1, max 4), limit (default 20, max 50)

One graph hop from a symbol/file/dir id: callers, callees, base classes, defining file.

repo_cache_stats

none

Result-cache counters: hits, misses, size, max_size, ttl_s, evictions, hit_rate. Never itself cached.

repo_build_status

task_id

Progress of a background --async-build: building, ready, failed or unknown, with progress_pct and eta_s.

The three content tools exclude secrets unconditionally — no flag turns that off — and every numeric argument is clamped in the handler, so a caller cannot widen a bound by asking. Full contract, argument ceilings and client configs: docs/mcp.md. Running it shared, over HTTP, with bearer or OIDC auth and an audit log: docs/ENTERPRISE_DEPLOYMENT.md.

📐 Architecture & token economics

The retrieval layer is a GraphRAG pipeline: tree-sitter parses the source into a typed graph, BM25 picks the seed chunks, and the graph — not further text similarity — decides what else is worth spending the budget on. Optional dense vectors (repo2graph embed) fuse into the seed ranking; nothing downstream requires them.

  • Nodes: repo, dir, file, symbol (function/method/class/struct/trait/interface/type), module (external dependency), external (an unresolved call target).

  • Edges: CONTAINS, DEFINES, IMPORTS, CALLS (carries count + confidence), CALLS_EXTERNAL, INHERITS, CO_CHANGE (from --git-history, requires 3+ co-edits).

  • Call resolution is name-based, not type-based — a deliberate trade-off that keeps repo2graph language-agnostic and setup-free. Ambiguous calls fan out to up to 5 candidate edges at confidence = 1/n; filter to confidence == 1.0 when you need certainty over recall.

  • Two budget models, on purpose: Index.retrieve()'s budget_chars bounds only the chunks' own text (a back-compat surface); Index.pack_context()'s budget_chars bounds the entire rendered markdown — citation headers, separators, everything. New retrieval code should be built on pack_context().

  • Chunking: roughly one chunk per function/class, cut at ~4000 characters with 8 lines of overlap so nothing is lost at a seam; each chunk's header names its callers and callees, which is what makes graph-expanded retrieval better than plain top-k text search.

Full breakdown of every node/edge kind and the chunk schema: docs/reference.md. The pipeline, the Python API, and where the graph guesses (and why): TECHNICAL.md.

📊 See it on real repositories

Not a toy demo — five real, large, public repositories, each indexed at a pinned commit, with the generated graph committed and the exact reproduction command recorded. Every number is measured, from benchmarks/results.json, not estimated.

Repository

Language(s)

Scope

Nodes

Edges

Kubernetes

Go

scoped (controllers, scheduler, API server)

14,451

110,246

TensorFlow

C++ / Python

scoped (Python/C++ boundary)

21,380

115,984

Django

Python

full repository

55,810

303,339

VS Code

TypeScript

scoped (src/vs/)

113,080

656,158

Linux kernel

C

scoped (extreme-scale)

136,219

256,413

See examples/README.md for the full index and reproduction commands, docs/benchmarks.md for methodology, and docs/limitations.md for what running against five real repositories actually surfaced (parse-error rates on macro-heavy C/C++, call-name ambiguity, cross-language resolution limits).

🧪 Reproducible benchmark & regression suite

Beyond full-scale public codebase indexing, repo2graph includes a reproducible 25-task benchmark across 5 application archetypes (benchmarks/corpus/: TypeScript app, Python backend, modular monolith, React frontend, and dynamic patterns) comparing repo2graph against ripgrep and agent-baseline search:

Metric

repo2graph (GraphRAG)

ripgrep Search

Agent Baseline Search

Query Correctness

100% (25/25)

80.0% (20/25)

56.0% (14/25)

Citation Precision

97.9% (46/47)

80.9% (38/47)

55.3% (26/47)

Mean Query Latency

1.82 ms

12.44 ms

31.84 ms

Token Budget Compliance

100% (Clamped)

0% (Unbounded)

72%

Full methodology, ground truth evidence, failure cases, and reproduction scripts: BENCHMARK.md. Continuous benchmark regression checking is enforced in CI via .github/workflows/benchmark.yml.

📖 CLI & server reference

Command

Does

repo2graph build <path> -o .r2g [--git-history N]

Parse a local repo into a graph + chunks.

repo2graph github <owner/repo> -o <dir>

Fetch, build, and clean up — no local clone needed.

repo2graph query "<question>" -o .r2g

Lexical search + one-hop graph expansion.

repo2graph rag "<question>" -o .r2g [--vectors] [--answer]

Budget-bounded GraphRAG pack; --answer sends it to an LLM (opt-in, network).

repo2graph embed -o .r2g [--verify-rag]

Compute/verify dense vectors for hybrid search.

repo2graph map -o .r2g [--viz-nodes N]

Regenerate graph.html with a different node cap.

repo2graph stats -o .r2g [--format text]

Node/edge/function counts for an existing index; --format text for a quality summary.

repo2graph index-status -o .r2g [--json] [--check]

Indexed commit/branch, build time, counts, languages, skipped paths, size, and freshness. --check exits 1 when stale.

repo2graph doctor [path]

Diagnose environment, dependencies, permissions, and index integrity.

repo2graph bug-report -o .r2g [--category CAT]

Privacy-preserving diagnostic bundle for an issue. No source code, no paths by default.

repo2graph explain-path <path> [-r <repo>]

Say whether a path would be indexed, and which precedence rule decided.

`repo2graph explain <edge

node

repo2graph impact -i .r2g [--base main]

PR / diff architectural blast radius & impact analysis.

repo2graph completion [shell]

Print shell tab completion setup script (bash, zsh, fish).

repo2graph-mcp <path> [--no-auto-build] [--async-build]

stdio MCP server over .r2g.

Environment variables (only read by rag --answer, in this precedence order): GEMINI_API_KEY → OPENAI_API_KEY → ANTHROPIC_API_KEY → OLLAMA_HOST. --model overrides the provider's best-effort default. No other command makes a network call or reads these. Full flag tables and budget accounting: docs/cli.md.

🔐 Security

  • Secure-by-default secret exclusion: build, github, auto-building query/rag, the GitHub Action and MCP all skip credential files automatically — .env*, private keys, certificates, .ssh, .aws, .gnupg, .kube, credentials/, secrets/ and more. --include-secrets opts out; the MCP tools have no equivalent, because an agent returning .env is a different problem from a human choosing to read it.

  • Content-aware secret scanning: Chunks are scanned for high-entropy tokens, cloud API keys (AWS, OpenAI, Google, Slack, GitHub), JWTs, DB URLs and private keys. Matches are redacted line-preservingly (--secret-policy redact-match|exclude-file|warn-only|off).

  • Sanitized logs and events: Audit logs and structured event sinks enforce cycle detection, container size limits, recursion depth ceilings, and scrub URL basic-auth credentials.

  • Local by default, and only one path ever sends your code: build, query, rag, map, stats and the MCP server over stdio open no socket at all — asserted by socket-level tests, not just by reading the code. Four commands can reach the network, and only the first sends anything of yours: rag --answer (uploads the pack; prints provider + hostname first), repo2graph github (clones), repo2graph embed (downloads an embedding model once), and repo2graph-mcp --auth-oidc-issuer (fetches public keys). No telemetry of any kind, and no setting to turn off.

Where every byte goes and how to delete it: docs/PRIVACY.md. What an attacker could try, and what is out of scope: docs/THREAT_MODEL.md. Hardened configurations to copy: docs/secure-configuration.md. Reporting a vulnerability: .github/SECURITY.md.

🤝 Contributing & community

git clone https://github.com/Srinivasan-78/repo2graph
cd repo2graph
python3 -m venv .venv && .venv/bin/pip install -e ".[dev]"
make lint format-check typecheck test    # the four gates CI runs

Branch from develop and open the PR against it — main is the release branch. make lint alone is only ruff check .; ruff format --check . is a separate gate and the one people miss.

  • docs/good-first-issues.md — seven starter tasks, each with a file and line to start from, acceptance criteria, and the catch that makes it harder than it looks.

  • .github/CONTRIBUTING.md — local setup, the CI gates, the branch model, the test house style, and the registry/Glama release process.

  • docs/ARCHITECTURE.md — the module map: what each module owns, which way dependencies run, and where a change of each kind goes.

  • docs/BACKLOG.md — deliberately deferred work and why; the closest thing to a roadmap.

  • LANGUAGE_SUPPORT.md — strategic language roadmap, scorecard generator, and ecosystem relationship extraction (routes, test links, DI, ORMs).

  • BENCHMARK.md — reproducible 25-task evaluation across 5 archetypes, comparison against ripgrep and agent baseline, and regression gating.

  • PR_IMPACT.md — flagship PR & diff architectural blast radius analysis, reverse caller exposure, test coverage mapping, and CI workflow.

  • docs/ROADMAP_LANGUAGE_ISSUES.md — prioritized tracking issues (LANG-01 to LANG-11) for deep language and framework support.

  • AGENTS.md — this codebase's non-obvious conventions (Windows encoding, text slicing, the two budget models) before editing repo2graph/.

  • POSITIONING.md — what repo2graph claims, what it deliberately does not claim, and the copy for every outward-facing surface. Read it before changing any of them.

  • CODE_OF_CONDUCT.md — Contributor Covenant v2.1.

  • Found a bug or have a feature idea? Open an issue.

License

MIT. See LICENSE.


Found repo2graph useful? Star the repo — it's the easiest way to help other people find it.

Available Tools

6 tools
repo_build_statusA
Read-onlyIdempotent

Query progress and status of a background index build task under --async-build. Read-only check of in-memory background worker. When to use: use when polling build progress after an async index build was started. When NOT to use: do not use when building synchronously or when queries already succeed. Once completed, use repo_search or repo_map to query code. Output: JSON object with task_id, status, parsed file progress, and error details.

ParametersJSON Schema
NameRequiredDescriptionDefault
task_idYesTask ID string returned by a previous tool call when an asynchronous build was initiated.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds useful behavioral context beyond annotations by stating it is a read-only check of an in-memory background worker and by specifying the output fields (task_id, status, parsed file progress, error details). This goes beyond what the annotations alone provide, though it does not disclose edge cases like unknown task IDs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-organized with clear sections for usage, non-usage, follow-up, and output. Every sentence contributes useful information without redundancy, and the most important scoping detail (--async-build) is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter read-only tool with strong annotations, the description is complete: it explains the operation, when to call it, what not to do, what the output contains, and what to use afterward. The absence of an output schema is compensated by listing the expected JSON fields.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the task_id parameter is already well documented as the ID returned by a previous async build call. The description does not add significant meaning beyond the schema, so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool as a status/progress query for a background index build task under --async-build, with a specific verb and resource. It also distinguishes itself from sibling tools by explicitly mentioning repo_search and repo_map as the post-completion alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit when-to-use guidance (polling after an async build was started) and when-not-to-use guidance (synchronous builds or when queries already succeed). It also names the correct follow-up tools, making the decision boundary unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

repo_cache_statsA
Read-onlyIdempotent

Retrieve runtime diagnostic counters for the tool result cache (hits, misses, size, max_size, ttl_s). Read-only, in-memory diagnostics, zero side effects. When to use: use when evaluating cache hit rate or debugging server performance. When NOT to use: do not use to search repository contents or inspect code structure; use repo_map or repo_search instead. Output: JSON object with cache metrics.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is covered. The description adds meaningful context beyond those: the data is in-memory, runtime diagnostics, and has zero side effects, which clarifies that the values are ephemeral and the call cannot influence server state.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the action and metric list, and each remaining sentence adds distinct value: side-effect declaration, when to use, when not to use, and expected output. It is concise enough for a zero-parameter tool while providing genuinely useful routing information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given zero parameters, no output schema, and rich annotations, the description covers everything an agent needs: what the tool does, which metrics it returns, the side-effect profile, and explicit alternatives. The output format is also stated ('JSON object with cache metrics'), so no critical information is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, and the schema description coverage is effectively 100% (empty schema), so there are no parameter semantics for the description to add. The baseline for zero-parameter tools is 4, and the description appropriately emphasizes that the tool is a straightforward read of counters without needing inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Retrieve runtime diagnostic counters for the tool result cache' and enumerates the exact metrics returned. It also explicitly contrasts with sibling tools by stating that repo_map and repo_search are for repository contents and structure, making differentiation clear.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit 'When to use' guidance (evaluating cache hit rate or debugging server performance) and explicit 'When NOT to use' guidance with named alternatives (repo_map or repo_search for repository content/code structure). This is the strongest possible usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

repo_impactA
Read-onlyIdempotent

Analyze PR or git diff impact against a base branch using the code graph. Detects changed symbols, affected public APIs, impacted callers, test coverage, and architectural blast radius with grounded citations. Read-only, deterministic, zero side effects. When to use: use when assessing PR risk, planning test execution, evaluating breaking API changes, or investigating diff blast radius. When NOT to use: do not use for generic lexical code search (use repo_search). Output: structured markdown impact report or PR summary.

ParametersJSON Schema
NameRequiredDescriptionDefault
baseNoBase ref or branch to compare against (default 'main').
diffNoOptional raw unified diff text. If provided, overrides git diff.
headNoHead ref or branch to compare (default 'HEAD' or current working tree).
formatNoReport format: 'markdown' (full report), 'pr-comment' (compact PR summary), or 'json'.
max_depthNoCaller traversal hops around changed symbols (default 2, max 4).

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds useful behavioral context beyond annotations: grounded citations, the use of a code graph, deterministic execution, and zero side effects, along with the structured output nature.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured and front-loaded, starting with the core purpose before moving to usage, exclusions, and output. Every sentence contributes value, and the 'When to use'/'When NOT to use' formatting makes it easy for an agent to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only analysis tool with no required parameters and full schema coverage, the description covers purpose, inputs, use cases, exclusions, and output format. Even without an output schema, it tells the agent what kind of report to expect and what content the report includes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all five parameters and the format enum. The description adds contextual behavior ('diff impact', 'PR summary') but does not need to elaborate on each parameter; this is appropriately at baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource: 'Analyze PR or git diff impact against a base branch using the code graph.' It lists concrete detection targets (changed symbols, affected APIs, callers, test coverage, blast radius) and explicitly distinguishes itself from generic lexical search via repo_search.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit 'When to use' guidance with concrete scenarios: assessing PR risk, planning test execution, evaluating breaking API changes, and investigating diff blast radius. It also gives a clear 'When NOT to use' exclusion and names the alternative tool, repo_search.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

repo_mapA
Read-onlyIdempotent

Retrieve a high-level structural map of the repository: languages, hub files, and top entry points. Read-only, deterministic, zero side effects. When to use: call this first at session start to understand codebase layout and identify entry points before detailed queries. Use when deciding where to investigate. When NOT to use: do not use to search code (use repo_search) or inspect call graphs (use repo_neighbours). Output: markdown summary of languages, hub files, and entry points.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, and the description reinforces this with 'Read-only, deterministic, zero side effects.' It adds useful context beyond annotations by specifying the output format ('markdown summary') and the intended session-start usage, which provide extra behavioral clarity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose, then safety, then usage guidance, and ends with output format. Every sentence adds value; the structure is clear and easy to scan. There is minimal redundancy beyond reinforcing the read-only nature.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a parameterless tool with no output schema, the description is complete: it states the artifact produced, the timing of use, the exclusions, and the alternatives. An agent has everything needed to select and invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool has zero parameters, so the schema carries no burden and there is nothing to document. The description compensates by clarifying what the tool returns ('markdown summary'), which is the relevant semantic context for a parameterless tool.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Retrieve a high-level structural map of the repository' and lists the content (languages, hub files, top entry points). It also distinguishes itself from siblings by explicitly naming repo_search and repo_neighbours as alternatives for other tasks.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit guidance on when to use it ('call this first at session start', 'Use when deciding where to investigate') and when not to use it, naming the correct sibling tools for search and call-graph inspection. This fully routes an agent to the right tool with no ambiguity.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

repo_neighboursA
Read-onlyIdempotent

Traverse code graph relationships from a known symbol or file node_id (callers, callees, base classes, definitions). Read-only, deterministic traversal, no side effects. When to use: use with a specific node_id (e.g. from repo_search citations) to inspect callers (CALLS in), callees (CALLS out), inheritance, or definitions. When NOT to use: do not use for text search across code (use repo_search) or repo overview (use repo_map). Output: markdown list formatted as - <EDGE_TYPE> <in|out>: <name> (<path:line>) [<node_id>].

ParametersJSON Schema
NameRequiredDescriptionDefault
hopsNoTraversal depth from node_id (default 1, max 4).
limitNoMaximum neighbor rows to return (default 20, max 50).
node_idYesTarget graph node identifier to expand from (e.g. 'sym:pkg/mod.py::func', 'file:pkg/mod.py', 'dir:pkg').

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description states 'Read-only, deterministic traversal, no side effects' and describes the output format as a markdown list with edge types. Annotations already convey read-only/idempotent/destructive hints, but the addition of 'deterministic traversal' and the exact output layout adds value beyond annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description front-loads the core purpose, then gives usage guidance, exclusions, and output format in compact sentences. Every sentence earns its place; no redundant filler or restating of schema fields.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no output schema and moderate complexity, the description covers invocation context, parameter source, exclusions, and output format. It could enumerate all possible edge types, but the format example plus 'callers, callees, base classes, definitions' is sufficient for an agent to use it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% with clear descriptions and examples for node_id, hops, and limit. The description adds the hint that node_id typically comes from repo_search citations, but this is contextual rather than necessary to parse the parameters. Baseline 3 is appropriate since the schema already handles semantics.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Traverse') and clearly identifies the resource ('code graph relationships') and scope ('from a known symbol or file node_id'). It also names the sibling tools it is not ('do not use for text search ... use repo_search') which solidifies differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit 'When to use' and 'When NOT to use' sections give concrete conditions and name alternatives (repo_search, repo_map). It also advises pairing with a node_id from repo_search citations, which is actionable guidance agents can follow directly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev2.2.0
    • Addedrepo_impact
  2. 3 tool updatesv1.5.1
    • Changedrepo_build_status1 field changed
      • changedInput schema / properties / task_id / description
        Previous value: -"the id a previous call returned"New value: +"Task ID string returned by a previous tool call when an asynchronous build was initiated."
    • Changedrepo_neighbours3 fields changed
      • changedInput schema / properties / hops / description
        Previous value: -"graph hops (default 1, max 4)"New value: +"Traversal depth from node_id (default 1, max 4)."
      • changedInput schema / properties / limit / description
        Previous value: -"neighbours (default 20, max 50)"New value: +"Maximum neighbor rows to return (default 20, max 50)."
      • changedInput schema / properties / node_id / description
        Previous value: -"e.g. sym:pkg/a.py::run"New value: +"Target graph node identifier to expand from (e.g. 'sym:pkg/mod.py::func', 'file:pkg/mod.py', 'dir:pkg')."
    • Changedrepo_search4 fields changed
      • changedInput schema / properties / budget_tokens / description
        Previous value: -"max 12000"New value: +"Maximum token ceiling for returned markdown pack (default 6000, max 12000)."
      • changedInput schema / properties / hops / description
        Previous value: -"graph hops (default 1, max 4)"New value: +"Graph traversal depth around seed chunks (default 1, max 4; 0 returns seeds only)."
      • changedInput schema / properties / k / description
        Previous value: -"seed chunks (default 8, max 50)"New value: +"Number of initial seed chunks retrieved via BM25 lexical scoring (default 8, max 50)."
      • changedInput schema / properties / query / description
        Previous value: -"the question"New value: +"Natural language question, search terms, or symbol identifier to search for (e.g. 'pack_context' or 'how does export work')."
  3. 5 tool updatesv0.1.0
    • First observedrepo_build_status
    • First observedrepo_cache_stats
    • First observedrepo_map
    • First observedrepo_neighbours
    • First observedrepo_search

TDQS

A4.7/5.0

Scored across 6 tools

Disambiguation5/5

Each tool has a clearly distinct purpose: repo_map for overview, repo_search for lexical search, repo_neighbours for graph traversal, repo_impact for diff/PR analysis, and the two status tools for operational diagnostics. The 'When NOT to use' guidance in each description further eliminates boundary confusion.

Naming Consistency5/5

All tools follow a consistent repo_ prefix with lowercase snake_case names, creating a predictable and recognizable pattern. Minor noun/verb variation (map vs search) is natural and does not undermine consistency.

Tool Count5/5

Six tools is a well-scoped count for a repository code-graph analysis server. Each tool adds a distinct capability without bloating the surface or leaving the set feeling thin.

Completeness5/5

The tool set covers the main analysis workflow: repo orientation, open-ended code search, symbol graph traversal, PR/diff impact analysis, and operational build/cache diagnostics. Since the server is read-only analysis tooling, missing CRUD operations are not a gap.

Maintenance

ActivityMaintained
ResponsivenessResponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    A
    maintenance
    An MCP code-intelligence server for AI agents with pre-indexed AST cache, 62 MCP tools, and TOON-compressed output, enabling token-efficient code analysis and project health grading entirely locally.
    9
    246 PyPI
    52
    MIT
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables LLM agents to query a codebase's structural knowledge (symbols, imports, call graphs, etc.) via MCP, reducing tokens and improving correctness compared to raw file access.
    47 npm
    7
    MIT
  • F
    license
    Not graded
    quality
    A
    maintenance
    Enables agents to build and query code knowledge graphs for repositories in a folder — finding shortest paths between concepts, explaining concepts with neighbours and community context, and visualizing per-repo graphs through MCP tools.
    -
  • F
    license
    Not graded
    quality
    C
    maintenance
    Enables coding agents to query a local-first code intelligence graph of Python repositories—covering functions, classes, modules, and their relationships—via MCP, supporting subgraph retrieval, caller lookup, and impact analysis without re-reading the codebase.
    -