repo2graph
This server provides read-only access to a repository's code graph, letting you map, search, and traverse code relationships with cited results.
repo_map: Get a high-level overview of the repository—languages, hub files, and entry points.
repo_search: Ask natural-language questions or search terms; returns cited code blocks
[cite: path:start-end]using BM25 plus graph neighbours, within a token budget.repo_neighbours: From a known symbol/file/dir
node_id, traverse callers, callees, base classes, and definitions up to 4 hops.repo_cache_stats: View runtime cache diagnostics (hits, misses, hit rate, TTL).
repo_build_status: Poll progress of a background async index build by
task_id.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@repo2graphFind the functions involved in user login and list their callers, callees, and neighbours."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
repo2graph
Give coding agents trustworthy, cited answers about unfamiliar codebases.
Ask a repository a question; get back the actual source that answers it, every block stamped with the file and line range it came from.
⚡ What is repo2graph?
An agent dropped into a codebase it has never seen has two bad options. Grep for a word and it either floods its context with whole matching files, or finds nothing because the code spells the idea differently than you did. Guess from training data and it writes something confident and wrong. Either way you cannot tell which of the two just happened.
repo2graph answers questions about a repository with the repository's own source. Ask "how
does a request get authenticated" and you get back the function that does it, the functions that
call it and the ones it calls — each block headed [cite: path:start-end], so every claim in the
answer is one click from the line it came from. If the answer is wrong, the citation shows you
where it went wrong. That is the whole point.
It gets there by reading the code rather than searching it: one parse pass records who calls whom, who imports what, and which class extends which, and retrieval follows those links instead of matching more text. The result is served straight into Claude Code, Cursor or any Model Context Protocol client, or packed into a markdown context with a hard token ceiling for any other LLM.
flowchart LR
A[your code] --> B[tree-sitter<br/>reads the code]
B --> C[graph<br/>dots + arrows]
C --> D[graph.html<br/>the picture]
C --> E[chunks.jsonl<br/>pieces for an AI]
C -->|MCP stdio| F[Claude / Cursor /<br/>any MCP client]No project setup, no language server, no build step — point it at a folder and it works.
Interactive canvas, zoomed | Filter & inspector controls |
graph.html is one self-contained file — no server, no internet, drag to pan, scroll to zoom,
click a node to inspect its code and neighbours.
Related MCP server: Lore MCP Server
👥 Who it's for
🧭 Joining a new codebase
The problem: week one goes on reading files to find out which ones matter.
Build once, open the map, and start from the hub files instead of the root directory. Then ask whole questions — "how does a request get from the router to the handler" — and read the answer as source, with the callers and callees already attached.
uvx repo2graph build . -o .r2g
open .r2g/human/graph.html
repo2graph rag "how does routing work" -o .r2g🤖 Driving a coding agent
The problem: the agent greps, pulls in three whole files, and still edits the wrong one.
Point Claude Code, Cursor or any MCP client at the repo. The agent gets cited blocks under a hard 12k-token ceiling instead of raw file dumps, and can walk from a symbol to its callers in one hop. Secrets are excluded from agent replies unconditionally — no flag turns that off.
claude mcp add repo2graph -- \
uvx --from "repo2graph[mcp]" repo2graph-mcp .🔍 Reviewing a pull request
The problem: the diff is 40 lines; the blast radius is unknown.
Ask the graph what touches the changed symbol — callers, importers, subclasses — and what the
repository's own history says usually changes alongside it (CO_CHANGE, mined from git). That
last one catches the test file or the config the diff forgot.
repo2graph build . -o .r2g --git-history 500
repo2graph explain node "sym:src/auth.py::verify" -o .r2g🌱 Maintaining a project
The problem: every new contributor asks the same "where do I start" question.
Commit a fresh graph on every push with the GitHub Action, and publish graph.html to a branch
contributors can browse. The job summary reports hub files, co-change hotspots and the graph delta
since the last build, so architectural drift shows up in the run.
- uses: Srinivasan-78/repo2graph@v2
with: { git-history: "500", commit-branch: graph }🔎 Why repo2graph instead of grep or vector search?
Both of those are still in the box — repo2graph seeds every query with BM25, and dense vectors
are an opt-in fusion. The difference is what happens after the first match.
grep / ripgrep | Embedding search | repo2graph | |
Finds | the exact string | text that reads similarly | the symbol, then everything wired to it |
Different words than the code uses | returns nothing | handles it | BM25 seeds, then graph hops reach code the query never named |
"What calls this?" | can't answer — a match in a comment ranks like the definition | can't answer — neighbours aren't in the embedding |
|
"What breaks if I change this?" | you read every hit by hand | not represented | callers, importers and subclasses in one hop |
What comes back | matching lines, or whole files an agent then dumps into context | top-k similar chunks, callers unretrieved | the source that answers it, each block headed |
Token cost | unbounded — the agent decides how much file to read | unbounded | hard ceiling on the whole pack, re-measured before it returns |
"Which files keep changing together?" | — | — |
|
Setup | none | index build + an embedding model (~90 MB) | one parse pass, no model, no API key, no language server |
Ranking is explainable | n/a | a cosine number |
|
Use grep when you want every occurrence of a literal string — a config key, an error message, a TODO. repo2graph has no special knowledge of string literals and will not beat it. Use repo2graph when the question is about relationships: what calls this, what breaks if I change it, how does data get from A to B. Longer version: docs/why-graph.md.
🚀 Four ways to run repo2graph
Same graph, same chunk format, same .r2g output — pick the interface for where you're
standing right now.
Local dev, scripting, ad-hoc questions from a terminal.
A fresh graph committed next to your code on every push, zero Python setup.
Give Claude, Cursor or any MCP client live, cited access to the codebase.
Enterprise-ready, read-only, non-root container deployment.
🐍 1. Python / CLI
Requires Python 3.10+. The fastest way to see what it does — a bundled example repo, indexed and interrogated, with nothing to configure and no repository of your own:
uvx repo2graph demoThat writes a small orders service to a scratch directory, builds a real graph over it, and answers five questions against it, each one citing files and line ranges. Then point it at your own code:
uvx repo2graph build . -o .r2g && open .r2g/human/graph.htmlOr install it properly:
pip install repo2graph
repo2graph build /path/to/project -o .r2g --git-history 200
repo2graph query "how does routing match a path" -o .r2gTwo minutes end to end, with the expected output at each step: docs/quickstart.md.
Five questions to start with
The ones that show what a graph gives you over a text search. Copy any of them
onto your own repo — or run repo2graph demo to watch each answered against
the bundled fixture.
Ask your repo | What comes back that grep cannot give you |
| the guard itself, plus every route that calls it |
| CALLS edges in, so callers come back even when the name is shadowed |
| IMPORTS edges from the test module back to the code under test |
| the blast radius: direct callers and what they are called from |
| a whole path across modules, each block cited to file and line |
repo2graph rag "What would be affected by changing pack_context?" -o .r2gSomething not working? repo2graph doctor . checks the environment, the index
and your MCP client config, and prints a fix for anything it finds.
One pass over this repository — 195 files, 2,552 nodes, 11,118 edges — takes about three seconds and needs no configuration file, no language server and no API key. Ask it something, and the answer comes back as source you can check, not a summary you have to trust:
Full flag tables, budget accounting and the Python API: docs/cli.md · docs/python-api.md.
⚙️ 2. GitHub Action
Published on the GitHub Marketplace — one step, no Python setup on the runner:
- uses: actions/checkout@v4
with: { fetch-depth: 0 } # full history, so CO_CHANGE edges are meaningful
- uses: Srinivasan-78/repo2graph@v2
with:
path: . # or: repo: some-org/other-repo
git-history: "500" # commits scanned for CO_CHANGE edges (0 = skip)
artifact-name: repo-graph@v2 follows every 2.x release; pin an exact tag (@v2.2.0) to upgrade by hand instead. It never
calls an LLM — --answer is deliberately not exposed — and it writes a job-summary table (hub
files, CO_CHANGE hotspots, the graph delta since the last build) straight from the artifacts, so
the shape of the map shows up in the run without downloading anything.
Also pack a cited context for a fixed question, and push the map to a browsable branch:
- uses: Srinivasan-78/repo2graph@v2
with:
query: "how does auth middleware validate a token"
commit-branch: graph # force-pushed; this repo's own /graph branch is built this wayAll inputs/outputs, private-repo tokens and the vector-embedding step: docs/github-action.md.
🔌 3. MCP server
repo2graph-mcp is a stdio MCP server. It builds its own index on the first call if one doesn't
exist yet — nothing to run ahead of time.
Claude Code
claude mcp add repo2graph -- uvx --from "repo2graph[mcp]" repo2graph-mcp /path/to/projectClaude Desktop (claude_desktop_config.json) and Cursor (.cursor/mcp.json) — same block:
{
"mcpServers": {
"repo2graph": {
"command": "uvx",
"args": ["--from", "repo2graph[mcp]", "repo2graph-mcp", "/path/to/project"]
}
}
}Any other stdio-based MCP client (Windsurf, Zed, generic clients) takes the same command/args
pair — see docs/mcp.md for config file locations per platform and client.
That is the hop grep cannot do: one symbol in, and its definer, its callers and its callees come back with file and line — the relationship, not a text match that happens to contain the name.
🐳 4. Docker
For enterprise and shared deployments, an official Dockerfile is provided. It's a multi-stage build running as a non-root user (10000:10000), fully compatible with a read-only root filesystem and dropped capabilities.
docker build -t repo2graph .
docker run --rm \
--read-only \
--cap-drop=ALL \
--security-opt=no-new-privileges \
--network=none \
-v /path/to/repo:/repo:ro \
-v repo2graph-index:/repo/.r2g \
repo2graph build /repo -o /repo/.r2gSee docs/ENTERPRISE_DEPLOYMENT.md for full container hardening and HTTP server instructions, and docs/deployment-security.md for the trust boundary and supported/not-recommended verdict per deployment shape — including whether HTTP without TLS is safe (short answer: only on loopback) and a worked hardened reverse-proxy example.
✨ Key features
Deterministic graph, not embeddings-only search | Callers, callees, imports and class hierarchies resolved from the actual AST — not a nearest-neighbour guess. |
Hybrid retrieval | BM25 + graph-neighbour expansion by default; optional dense vector fusion ( |
Hard token ceilings, enforced twice |
|
17 grammars, full treatment | Python, JS, TS, TSX, Go, Rust, Java, Ruby, C, C++, C#, PHP, Kotlin, Swift, Scala, Bash and Lua get functions/classes/calls — 29 file extensions in all. Everything else still appears as files on the map. See our strategic language scorecard and framework relationship RFC. |
CI-native | Published as a GitHub Action — commit a fresh graph next to your code on every push. |
Local by default |
|
Export to real graph tooling |
|
⚖️ What it does — and what it does not
A retrieval tool that oversells itself is worse than no retrieval tool, because you stop checking its answers. So, plainly:
It does
Return the source that answers a question, cited to
path:start-end, inside a token budget it enforces rather than requests.Resolve callers, callees, imports and class hierarchies from a real parse of the code, and let you walk them in either direction from any symbol.
Mine
CO_CHANGEfrom git history — the files that keep being edited together, which no parser can tell you.Run entirely locally, with no model, no account and no network call, in the CLI, in CI and over MCP.
Degrade gracefully: an unparsed language still appears as file nodes and is still retrievable as text; a missing vector index falls back to BM25 rather than failing.
It does not
Limitation | What that means in practice |
Resolve calls by type | Calls are matched by name, with same-class / same-file / import scoping to break ties. When scoping can't isolate one target, the call fans out to up to 5 candidate edges at |
See dynamic dispatch | A string-keyed lookup, a plugin registry, |
See reflection or computed imports |
|
Follow dependency injection to the implementation | A DI container wires an interface to a concrete class at runtime. The call site names the interface method, so the edge lands on the declaration (or fans out across every same-named implementation), never on the class the container actually injected. Walk |
Distinguish generated code | A |
Notice that your files changed | The index is a snapshot of the tree you built it from. Nothing watches the filesystem: edit a file and the graph keeps describing the old one. Rebuild ( |
Cross a language boundary | Python calling into C++ through generated bindings becomes a |
Parse macro-heavy C/C++ cleanly | tree-sitter emits |
Every one of these is measured, not asserted — the rates, the repositories they were measured on, and the reproduction commands are in docs/limitations.md. See also the dynamic pattern evaluation in BENCHMARK.md and the framework relationship proposal in docs/rfcs/rfc-framework-relationship-graph.md.
🆚 How it compares to other graph tools
Several tools build a graph out of a codebase. The thing that separates them is what comes back when you ask a question — a picture, a subgraph, or the code itself.
repo2graph | Code Graph (Obsidian) | grep / embedding RAG | ||
What a query returns | the source, packed — every block headed | a scoped subgraph, a path, or a concept explanation to traverse | a force-directed picture to read | matching lines, or nearest-neighbour chunks |
How hits are ranked | BM25 seeds, then k-hop graph expansion; optional dense fusion | graph traversal (explicitly not a vector index) | n/a — it is a view | lexical only, or vectors only |
Token budget | hard cap on the whole pack, re-measured before returning (12k ceiling over MCP) | not a packing layer | n/a | usually unbounded |
Edges from git history |
| — | — | — |
Runs with no assistant, no model, no account | yes — CLI, MCP, or the GitHub Action | code pass is local; the docs/media pass uses a model | needs Obsidian desktop 1.7.2+ | varies |
Corpus | code in 17 parsed grammars, every other file as text | code in ~40 languages, plus docs, PDFs, images, video | TS/TSX/JS/Python parsed, imports-only for 8 more | anything |
Reach for Graphify when the graph itself is the product: community detection, shortest path between two concepts, and your PDFs and design docs in the same graph as the code. Reach for the Obsidian plugin when a human wants to read the graph beside their notes. Reach for repo2graph when an agent needs cited source inside a fixed token budget, when it has to run in CI with no model and no account, or when "which files keep changing together" is part of the answer.
Longer version, with the trade-offs each choice implies: docs/comparison.md.
🛠️ MCP tools exposed
Five tools. Three answer questions about the code; two report on the server itself.
Tool | Arguments | What comes back |
| none | Languages, hub files, and top entry points. Stable across calls — read this first. |
|
| Seed chunks plus graph neighbours, each block headed |
|
| One graph hop from a symbol/file/dir id: callers, callees, base classes, defining file. |
| none | Result-cache counters: |
|
| Progress of a background |
The three content tools exclude secrets unconditionally — no flag turns that off — and every numeric argument is clamped in the handler, so a caller cannot widen a bound by asking. Full contract, argument ceilings and client configs: docs/mcp.md. Running it shared, over HTTP, with bearer or OIDC auth and an audit log: docs/ENTERPRISE_DEPLOYMENT.md.
📐 Architecture & token economics
The retrieval layer is a GraphRAG pipeline: tree-sitter
parses the source into a typed graph, BM25 picks the seed chunks, and the graph — not further text
similarity — decides what else is worth spending the budget on. Optional dense vectors
(repo2graph embed) fuse into the seed ranking; nothing downstream requires them.
Nodes:
repo,dir,file,symbol(function/method/class/struct/trait/interface/type),module(external dependency),external(an unresolved call target).Edges:
CONTAINS,DEFINES,IMPORTS,CALLS(carriescount+confidence),CALLS_EXTERNAL,INHERITS,CO_CHANGE(from--git-history, requires 3+ co-edits).Call resolution is name-based, not type-based — a deliberate trade-off that keeps repo2graph language-agnostic and setup-free. Ambiguous calls fan out to up to 5 candidate edges at
confidence = 1/n; filter toconfidence == 1.0when you need certainty over recall.Two budget models, on purpose:
Index.retrieve()'sbudget_charsbounds only the chunks' own text (a back-compat surface);Index.pack_context()'sbudget_charsbounds the entire rendered markdown — citation headers, separators, everything. New retrieval code should be built onpack_context().Chunking: roughly one chunk per function/class, cut at ~4000 characters with 8 lines of overlap so nothing is lost at a seam; each chunk's header names its callers and callees, which is what makes graph-expanded retrieval better than plain top-k text search.
Full breakdown of every node/edge kind and the chunk schema: docs/reference.md. The pipeline, the Python API, and where the graph guesses (and why): TECHNICAL.md.
📊 See it on real repositories
Not a toy demo — five real, large, public repositories, each indexed at a pinned commit, with the
generated graph committed and the exact reproduction command recorded. Every number is measured,
from benchmarks/results.json, not estimated.
Repository | Language(s) | Scope | Nodes | Edges |
Go | scoped (controllers, scheduler, API server) | 14,451 | 110,246 | |
C++ / Python | scoped (Python/C++ boundary) | 21,380 | 115,984 | |
Python | full repository | 55,810 | 303,339 | |
TypeScript | scoped ( | 113,080 | 656,158 | |
C | scoped (extreme-scale) | 136,219 | 256,413 |
See examples/README.md for the full index and reproduction commands, docs/benchmarks.md for methodology, and docs/limitations.md for what running against five real repositories actually surfaced (parse-error rates on macro-heavy C/C++, call-name ambiguity, cross-language resolution limits).
🧪 Reproducible benchmark & regression suite
Beyond full-scale public codebase indexing, repo2graph includes a reproducible 25-task benchmark across 5 application archetypes (benchmarks/corpus/: TypeScript app, Python backend, modular monolith, React frontend, and dynamic patterns) comparing repo2graph against ripgrep and agent-baseline search:
Metric | repo2graph (GraphRAG) |
| Agent Baseline Search |
Query Correctness | 100% (25/25) | 80.0% (20/25) | 56.0% (14/25) |
Citation Precision | 97.9% (46/47) | 80.9% (38/47) | 55.3% (26/47) |
Mean Query Latency | 1.82 ms | 12.44 ms | 31.84 ms |
Token Budget Compliance | 100% (Clamped) | 0% (Unbounded) | 72% |
Full methodology, ground truth evidence, failure cases, and reproduction scripts: BENCHMARK.md.
Continuous benchmark regression checking is enforced in CI via .github/workflows/benchmark.yml.
📖 CLI & server reference
Command | Does |
| Parse a local repo into a graph + chunks. |
| Fetch, build, and clean up — no local clone needed. |
| Lexical search + one-hop graph expansion. |
| Budget-bounded GraphRAG pack; |
| Compute/verify dense vectors for hybrid search. |
| Regenerate |
| Node/edge/function counts for an existing index; |
| Indexed commit/branch, build time, counts, languages, skipped paths, size, and freshness. |
| Diagnose environment, dependencies, permissions, and index integrity. |
| Privacy-preserving diagnostic bundle for an issue. No source code, no paths by default. |
| Say whether a path would be indexed, and which precedence rule decided. |
`repo2graph explain <edge | node |
| PR / diff architectural blast radius & impact analysis. |
| Print shell tab completion setup script ( |
| stdio MCP server over |
Environment variables (only read by rag --answer, in this precedence order):
GEMINI_API_KEY → OPENAI_API_KEY → ANTHROPIC_API_KEY → OLLAMA_HOST. --model overrides the
provider's best-effort default. No other command makes a network call or reads these. Full flag
tables and budget accounting: docs/cli.md.
🔐 Security
Secure-by-default secret exclusion:
build,github, auto-buildingquery/rag, the GitHub Action and MCP all skip credential files automatically —.env*, private keys, certificates,.ssh,.aws,.gnupg,.kube,credentials/,secrets/and more.--include-secretsopts out; the MCP tools have no equivalent, because an agent returning.envis a different problem from a human choosing to read it.Content-aware secret scanning: Chunks are scanned for high-entropy tokens, cloud API keys (AWS, OpenAI, Google, Slack, GitHub), JWTs, DB URLs and private keys. Matches are redacted line-preservingly (
--secret-policy redact-match|exclude-file|warn-only|off).Sanitized logs and events: Audit logs and structured event sinks enforce cycle detection, container size limits, recursion depth ceilings, and scrub URL basic-auth credentials.
Local by default, and only one path ever sends your code:
build,query,rag,map,statsand the MCP server over stdio open no socket at all — asserted by socket-level tests, not just by reading the code. Four commands can reach the network, and only the first sends anything of yours:rag --answer(uploads the pack; prints provider + hostname first),repo2graph github(clones),repo2graph embed(downloads an embedding model once), andrepo2graph-mcp --auth-oidc-issuer(fetches public keys). No telemetry of any kind, and no setting to turn off.
Where every byte goes and how to delete it: docs/PRIVACY.md. What an attacker could try, and what is out of scope: docs/THREAT_MODEL.md. Hardened configurations to copy: docs/secure-configuration.md. Reporting a vulnerability: .github/SECURITY.md.
🤝 Contributing & community
git clone https://github.com/Srinivasan-78/repo2graph
cd repo2graph
python3 -m venv .venv && .venv/bin/pip install -e ".[dev]"
make lint format-check typecheck test # the four gates CI runsBranch from develop and open the PR against it — main is the release branch. make lint
alone is only ruff check .; ruff format --check . is a separate gate and the one people
miss.
docs/good-first-issues.md — seven starter tasks, each with a file and line to start from, acceptance criteria, and the catch that makes it harder than it looks.
.github/CONTRIBUTING.md — local setup, the CI gates, the branch model, the test house style, and the registry/Glama release process.
docs/ARCHITECTURE.md — the module map: what each module owns, which way dependencies run, and where a change of each kind goes.
docs/BACKLOG.md — deliberately deferred work and why; the closest thing to a roadmap.
LANGUAGE_SUPPORT.md — strategic language roadmap, scorecard generator, and ecosystem relationship extraction (routes, test links, DI, ORMs).
BENCHMARK.md — reproducible 25-task evaluation across 5 archetypes, comparison against ripgrep and agent baseline, and regression gating.
PR_IMPACT.md — flagship PR & diff architectural blast radius analysis, reverse caller exposure, test coverage mapping, and CI workflow.
docs/ROADMAP_LANGUAGE_ISSUES.md — prioritized tracking issues (
LANG-01toLANG-11) for deep language and framework support.AGENTS.md — this codebase's non-obvious conventions (Windows encoding, text slicing, the two budget models) before editing
repo2graph/.POSITIONING.md — what repo2graph claims, what it deliberately does not claim, and the copy for every outward-facing surface. Read it before changing any of them.
CODE_OF_CONDUCT.md — Contributor Covenant v2.1.
Found a bug or have a feature idea? Open an issue.
License
MIT. See LICENSE.
Found repo2graph useful? Star the repo — it's the easiest way to help other people find it.
Available Tools
6 toolsrepo_build_statusARead-onlyIdempotent
Query progress and status of a background index build task under --async-build. Read-only check of in-memory background worker. When to use: use when polling build progress after an async index build was started. When NOT to use: do not use when building synchronously or when queries already succeed. Once completed, use repo_search or repo_map to query code. Output: JSON object with task_id, status, parsed file progress, and error details.
| Name | Required | Description | Default |
|---|---|---|---|
| task_id | Yes | Task ID string returned by a previous tool call when an asynchronous build was initiated. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds useful behavioral context beyond annotations by stating it is a read-only check of an in-memory background worker and by specifying the output fields (task_id, status, parsed file progress, error details). This goes beyond what the annotations alone provide, though it does not disclose edge cases like unknown task IDs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-organized with clear sections for usage, non-usage, follow-up, and output. Every sentence contributes useful information without redundancy, and the most important scoping detail (--async-build) is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-parameter read-only tool with strong annotations, the description is complete: it explains the operation, when to call it, what not to do, what the output contains, and what to use afterward. The absence of an output schema is compensated by listing the expected JSON fields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the task_id parameter is already well documented as the ID returned by a previous async build call. The description does not add significant meaning beyond the schema, so the baseline score of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as a status/progress query for a background index build task under --async-build, with a specific verb and resource. It also distinguishes itself from sibling tools by explicitly mentioning repo_search and repo_map as the post-completion alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit when-to-use guidance (polling after an async build was started) and when-not-to-use guidance (synchronous builds or when queries already succeed). It also names the correct follow-up tools, making the decision boundary unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
repo_cache_statsARead-onlyIdempotent
Retrieve runtime diagnostic counters for the tool result cache (hits, misses, size, max_size, ttl_s). Read-only, in-memory diagnostics, zero side effects. When to use: use when evaluating cache hit rate or debugging server performance. When NOT to use: do not use to search repository contents or inspect code structure; use repo_map or repo_search instead. Output: JSON object with cache metrics.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is covered. The description adds meaningful context beyond those: the data is in-memory, runtime diagnostics, and has zero side effects, which clarifies that the values are ephemeral and the call cannot influence server state.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the action and metric list, and each remaining sentence adds distinct value: side-effect declaration, when to use, when not to use, and expected output. It is concise enough for a zero-parameter tool while providing genuinely useful routing information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given zero parameters, no output schema, and rich annotations, the description covers everything an agent needs: what the tool does, which metrics it returns, the side-effect profile, and explicit alternatives. The output format is also stated ('JSON object with cache metrics'), so no critical information is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, and the schema description coverage is effectively 100% (empty schema), so there are no parameter semantics for the description to add. The baseline for zero-parameter tools is 4, and the description appropriately emphasizes that the tool is a straightforward read of counters without needing inputs.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Retrieve runtime diagnostic counters for the tool result cache' and enumerates the exact metrics returned. It also explicitly contrasts with sibling tools by stating that repo_map and repo_search are for repository contents and structure, making differentiation clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit 'When to use' guidance (evaluating cache hit rate or debugging server performance) and explicit 'When NOT to use' guidance with named alternatives (repo_map or repo_search for repository content/code structure). This is the strongest possible usage guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
repo_impactARead-onlyIdempotent
Analyze PR or git diff impact against a base branch using the code graph. Detects changed symbols, affected public APIs, impacted callers, test coverage, and architectural blast radius with grounded citations. Read-only, deterministic, zero side effects. When to use: use when assessing PR risk, planning test execution, evaluating breaking API changes, or investigating diff blast radius. When NOT to use: do not use for generic lexical code search (use repo_search). Output: structured markdown impact report or PR summary.
| Name | Required | Description | Default |
|---|---|---|---|
| base | No | Base ref or branch to compare against (default 'main'). | |
| diff | No | Optional raw unified diff text. If provided, overrides git diff. | |
| head | No | Head ref or branch to compare (default 'HEAD' or current working tree). | |
| format | No | Report format: 'markdown' (full report), 'pr-comment' (compact PR summary), or 'json'. | |
| max_depth | No | Caller traversal hops around changed symbols (default 2, max 4). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds useful behavioral context beyond annotations: grounded citations, the use of a code graph, deterministic execution, and zero side effects, along with the structured output nature.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured and front-loaded, starting with the core purpose before moving to usage, exclusions, and output. Every sentence contributes value, and the 'When to use'/'When NOT to use' formatting makes it easy for an agent to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only analysis tool with no required parameters and full schema coverage, the description covers purpose, inputs, use cases, exclusions, and output format. Even without an output schema, it tells the agent what kind of report to expect and what content the report includes.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all five parameters and the format enum. The description adds contextual behavior ('diff impact', 'PR summary') but does not need to elaborate on each parameter; this is appropriately at baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: 'Analyze PR or git diff impact against a base branch using the code graph.' It lists concrete detection targets (changed symbols, affected APIs, callers, test coverage, blast radius) and explicitly distinguishes itself from generic lexical search via repo_search.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit 'When to use' guidance with concrete scenarios: assessing PR risk, planning test execution, evaluating breaking API changes, and investigating diff blast radius. It also gives a clear 'When NOT to use' exclusion and names the alternative tool, repo_search.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
repo_mapARead-onlyIdempotent
Retrieve a high-level structural map of the repository: languages, hub files, and top entry points. Read-only, deterministic, zero side effects. When to use: call this first at session start to understand codebase layout and identify entry points before detailed queries. Use when deciding where to investigate. When NOT to use: do not use to search code (use repo_search) or inspect call graphs (use repo_neighbours). Output: markdown summary of languages, hub files, and entry points.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, and the description reinforces this with 'Read-only, deterministic, zero side effects.' It adds useful context beyond annotations by specifying the output format ('markdown summary') and the intended session-start usage, which provide extra behavioral clarity.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with the core purpose, then safety, then usage guidance, and ends with output format. Every sentence adds value; the structure is clear and easy to scan. There is minimal redundancy beyond reinforcing the read-only nature.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless tool with no output schema, the description is complete: it states the artifact produced, the timing of use, the exclusions, and the alternatives. An agent has everything needed to select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema carries no burden and there is nothing to document. The description compensates by clarifying what the tool returns ('markdown summary'), which is the relevant semantic context for a parameterless tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Retrieve a high-level structural map of the repository' and lists the content (languages, hub files, top entry points). It also distinguishes itself from siblings by explicitly naming repo_search and repo_neighbours as alternatives for other tasks.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit guidance on when to use it ('call this first at session start', 'Use when deciding where to investigate') and when not to use it, naming the correct sibling tools for search and call-graph inspection. This fully routes an agent to the right tool with no ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
repo_neighboursARead-onlyIdempotent
Traverse code graph relationships from a known symbol or file node_id (callers, callees, base classes, definitions). Read-only, deterministic traversal, no side effects. When to use: use with a specific node_id (e.g. from repo_search citations) to inspect callers (CALLS in), callees (CALLS out), inheritance, or definitions. When NOT to use: do not use for text search across code (use repo_search) or repo overview (use repo_map). Output: markdown list formatted as - <EDGE_TYPE> <in|out>: <name> (<path:line>) [<node_id>].
| Name | Required | Description | Default |
|---|---|---|---|
| hops | No | Traversal depth from node_id (default 1, max 4). | |
| limit | No | Maximum neighbor rows to return (default 20, max 50). | |
| node_id | Yes | Target graph node identifier to expand from (e.g. 'sym:pkg/mod.py::func', 'file:pkg/mod.py', 'dir:pkg'). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description states 'Read-only, deterministic traversal, no side effects' and describes the output format as a markdown list with edge types. Annotations already convey read-only/idempotent/destructive hints, but the addition of 'deterministic traversal' and the exact output layout adds value beyond annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description front-loads the core purpose, then gives usage guidance, exclusions, and output format in compact sentences. Every sentence earns its place; no redundant filler or restating of schema fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema and moderate complexity, the description covers invocation context, parameter source, exclusions, and output format. It could enumerate all possible edge types, but the format example plus 'callers, callees, base classes, definitions' is sufficient for an agent to use it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with clear descriptions and examples for node_id, hops, and limit. The description adds the hint that node_id typically comes from repo_search citations, but this is contextual rather than necessary to parse the parameters. Baseline 3 is appropriate since the schema already handles semantics.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Traverse') and clearly identifies the resource ('code graph relationships') and scope ('from a known symbol or file node_id'). It also names the sibling tools it is not ('do not use for text search ... use repo_search') which solidifies differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit 'When to use' and 'When NOT to use' sections give concrete conditions and name alternatives (repo_search, repo_map). It also advises pairing with a node_id from repo_search citations, which is actionable guidance agents can follow directly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
repo_searchARead-onlyIdempotent
Search repository code for answers to questions using BM25 lexical ranking expanded with graph neighbours. Read-only, no side effects, secret files (.env) excluded. When to use: use for open-ended queries, locating implementations, or finding error strings. When NOT to use: do not use when you already have a symbol node_id and want callers/callees (use repo_neighbours); do not use for broad repo layout (use repo_map). Output: markdown citation blocks [cite: path:start-end] bounded by budget_tokens.
| Name | Required | Description | Default |
|---|---|---|---|
| k | No | Number of initial seed chunks retrieved via BM25 lexical scoring (default 8, max 50). | |
| hops | No | Graph traversal depth around seed chunks (default 1, max 4; 0 returns seeds only). | |
| query | Yes | Natural language question, search terms, or symbol identifier to search for (e.g. 'pack_context' or 'how does export work'). | |
| budget_tokens | No | Maximum token ceiling for returned markdown pack (default 6000, max 12000). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is covered. The description adds genuinely useful behavioral context beyond annotations: secret files (.env) are excluded, the output format is markdown citation blocks, and results are bounded by budget_tokens. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every sentence earns its place: purpose, safety, when-to-use, when-not-to-use, and output format. It is front-loaded with the core purpose and the exclusions are clearly separated, making it easy for an agent to parse quickly.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with rich annotations and 100% parameter schema coverage, the description covers everything an agent needs to decide when to call it, what it does, what it returns, and how to avoid misuse. The sibling alternatives are named, the output format is specified, and the safety profile is already in annotations. No critical gap remains.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema fully documents all four parameters. The description does not add significant parameter-level meaning beyond what the schema already provides, though it does clarify that output is citation blocks bounded by budget_tokens, which is consistent with the schema's description. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Search repository code for answers to questions.' It also names the search technique (BM25 lexical ranking expanded with graph neighbours) and explicitly contrasts with sibling tools in the usage section, making the tool's distinct role clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides explicit when-to-use guidance ('open-ended queries, locating implementations, or finding error strings') and explicit when-not-to-use guidance with named alternatives (repo_neighbours for callers/callees, repo_map for broad layout). An agent can reliably route to the correct tool without further inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v2.2.0- Added
repo_impact
3 tool updates
v1.5.1- Changed
repo_build_status1 field changed- changed
Input schema / properties / task_id / descriptionPrevious value: -"the id a previous call returned"New value: +"Task ID string returned by a previous tool call when an asynchronous build was initiated."
- Changed
repo_neighbours3 fields changed- changed
Input schema / properties / hops / descriptionPrevious value: -"graph hops (default 1, max 4)"New value: +"Traversal depth from node_id (default 1, max 4)." - changed
Input schema / properties / limit / descriptionPrevious value: -"neighbours (default 20, max 50)"New value: +"Maximum neighbor rows to return (default 20, max 50)." - changed
Input schema / properties / node_id / descriptionPrevious value: -"e.g. sym:pkg/a.py::run"New value: +"Target graph node identifier to expand from (e.g. 'sym:pkg/mod.py::func', 'file:pkg/mod.py', 'dir:pkg')."
- Changed
repo_search4 fields changed- changed
Input schema / properties / budget_tokens / descriptionPrevious value: -"max 12000"New value: +"Maximum token ceiling for returned markdown pack (default 6000, max 12000)." - changed
Input schema / properties / hops / descriptionPrevious value: -"graph hops (default 1, max 4)"New value: +"Graph traversal depth around seed chunks (default 1, max 4; 0 returns seeds only)." - changed
Input schema / properties / k / descriptionPrevious value: -"seed chunks (default 8, max 50)"New value: +"Number of initial seed chunks retrieved via BM25 lexical scoring (default 8, max 50)." - changed
Input schema / properties / query / descriptionPrevious value: -"the question"New value: +"Natural language question, search terms, or symbol identifier to search for (e.g. 'pack_context' or 'how does export work')."
5 tool updates
v0.1.0- First observed
repo_build_status - First observed
repo_cache_stats - First observed
repo_map - First observed
repo_neighbours - First observed
repo_search
TDQS
Scored across 6 tools
Each tool has a clearly distinct purpose: repo_map for overview, repo_search for lexical search, repo_neighbours for graph traversal, repo_impact for diff/PR analysis, and the two status tools for operational diagnostics. The 'When NOT to use' guidance in each description further eliminates boundary confusion.
All tools follow a consistent repo_ prefix with lowercase snake_case names, creating a predictable and recognizable pattern. Minor noun/verb variation (map vs search) is natural and does not undermine consistency.
Six tools is a well-scoped count for a repository code-graph analysis server. Each tool adds a distinct capability without bloating the surface or leaving the set feeling thin.
The tool set covers the main analysis workflow: repo orientation, open-ended code search, symbol graph traversal, PR/diff impact analysis, and operational build/cache diagnostics. Since the server is read-only analysis tooling, missing CRUD operations are not a gap.
Maintenance
Related MCP Connectors
Hosted code graph over MCP: exact callers, dependencies, and cross-repo blast radius for AI agents.
Repository knowledge graph MCP server for codebase understanding and debugging.
Search GitHub, npm, PyPI, StackOverflow, ArXiv from one MCP — built for coding agents.
Codebase graphs, caller impact analysis, and recorded project context for AI coding agents.
Related MCP Servers
- AlicenseBqualityAmaintenanceAn MCP code-intelligence server for AI agents with pre-indexed AST cache, 62 MCP tools, and TOON-compressed output, enabling token-efficient code analysis and project health grading entirely locally.9246 PyPI52MIT
- AlicenseNot gradedqualityBmaintenanceEnables LLM agents to query a codebase's structural knowledge (symbols, imports, call graphs, etc.) via MCP, reducing tokens and improving correctness compared to raw file access.47 npm7MIT
- FlicenseNot gradedqualityAmaintenanceEnables agents to build and query code knowledge graphs for repositories in a folder — finding shortest paths between concepts, explaining concepts with neighbours and community context, and visualizing per-repo graphs through MCP tools.-
- FlicenseNot gradedqualityCmaintenanceEnables coding agents to query a local-first code intelligence graph of Python repositories—covering functions, classes, modules, and their relationships—via MCP, supporting subgraph retrieval, caller lookup, and impact analysis without re-reading the codebase.-