Skip to main content
Glama

repo2graph

Cited, token-bounded answers about a codebase, for coding agents and CI.

PyPI CI License: MIT

Ask a repository a question and get back the source that answers it. Every block is headed [cite: path:start-end], and the whole reply fits under a token ceiling that the tool enforces. One tree-sitter pass records who calls whom, who imports what, and which files keep changing together in git. Retrieval starts from BM25 matches and follows those links.

It needs no model, no API key, no language server and no database. Building, querying and the MCP server make no network calls; the only exceptions are rag --answer (opt-in, sends the pack to an LLM) and repo2graph github (clones a repository). Parsers are part of that promise: tree-sitter-language-pack is pinned below 1.0 so every grammar arrives compiled into the installed wheel, rather than being downloaded on first parse — see SECURITY.md. Use it from the CLI, as an MCP server in Claude Code or Cursor, or as a GitHub Action.

Try it

uvx repo2graph demo                                   # bundled example repo, five questions answered
uvx repo2graph build . -o .r2g                        # index your own code
uvx repo2graph rag "how does routing work" -o .r2g    # a cited, budget-bounded context pack

Give an agent the same thing over MCP:

claude mcp add repo2graph -- uvx --from "repo2graph[mcp]" repo2graph-mcp .

Cursor, Claude Desktop and other clients: docs/mcp.md. Step-by-step with expected output: demo in docs/cli.md.

Related MCP server: Lore MCP Server

Is it better than grep?

On cross-file structural questions, yes. On single-file lexical queries, no — grep is the better tool. Both halves are measured on 40 lexical and 40 structural questions about Flask, requests, FastAPI and Hono that repo2graph was not tuned against, scored against the definitions that answer them, every retriever held to the same token budget (method, full tables and diagnosis; raw rows in results_holdout_final.json).

Where repo2graph wins: cross-file structural questions

grep structurally cannot traverse dependency edges, compute reverse call closures, or follow cross-module delegation. On questions whose evidence provably spans a cross-file graph edge:

Budget

repo2graph

lexical search alone

grep, then read around hits

2,000 tokens

29%

24%

14%

4,000 tokens

60%

48%

21%

8,000 tokens

79%

55%

38%

The graph is what does it, and that is the comparison that matters: expansion adds +5/+12/+24 pp over the same retriever with expansion switched off. Against grep the margin is +15/+39/+41 pp — it widens with budget rather than closing, because grep has no edge to follow however much room it is given. It also gets there for fewer tokens than grep at every budget, by 7, 7 and 398 mean tokens, because a cited definition is a smaller thing to return than a window around every textual hit.

Where grep wins: single-file lexical questions

Budget

repo2graph

grep, then read around hits

2,000 tokens

17%

28%

4,000 tokens

29%

48%

8,000 tokens

45%

59%

On purely lexical questions where the evidence sits in a single file, text search is grep's optimum, and this gap is real: −11/−19/−14 pp. We have not closed it. The remaining cause looks like vocabulary mismatch rather than ranking — a question that says "datetime" does not match code that says "timestamp" — which is why the dense-vector path closes it to −2 pp at 8,000 tokens and more BM25 tuning has not. That path needs a downloaded model, so it is opt-in (pip install "repo2graph[rag]") and the zero-dependency default stays the default.

Where repo2graph wins for agents: fewer turns

In simulated agent workflows (search → read → answer, via scripts/agent_eval.py). No model is in the loop — the "agent" is a deterministic policy over real ripgrep and a real index:

Task set

Method

Success

Mean turns

Mean tokens

Precision per read

Structural (40)

repo2graph

40%

1.9

2,590

12.0%

Structural (40)

ripgrep

10%

2.5

586

29.0%

Lexical (40)

repo2graph

20%

1.4

2,070

21.7%

Lexical (40)

ripgrep

18%

3.0

740

89.1%

Four times the structural success rate, and it reaches an answer in one or two turns where grep needs three. The turn count is the number that moved most — structural went 3.1 to 1.9 — because sharper seed ranking puts the answer in the first pack more often, and a turn saved is worth more to an agent than a token saved. It is not cheap: 4.4× grep's tokens on the structural set and 2.8× on the lexical one, and grep wins precision per read on both. repo2graph buys recall and turns with context; the budget-matched comparison is the single-shot tables above.

For completeness, the 35+10 question set this page used to report — visible since the first version of these tables and therefore a regression set, not evidence — now reads 44/63/80% lexical against grep's 35/61/72%, and 60/80/100% structural against grep's 20/20/70%. Those are the better-looking numbers, which is exactly why the held-out set is the one quoted above.

What it does do that grep doesn't:

  • Relationships in one hop. repo_neighbours returns a symbol's callers, callees, base classes and defining file, each with a line number. Every edge carries a confidence: a name that could mean several definitions is marked ambiguous and priced at 1/n, not guessed.

  • A hard ceiling. The pack is measured, clamped and re-measured before it's returned (12k tokens max over MCP), so an agent can't flood its own context through this tool.

  • Reverse closures from the graph. repo_blast_radius walks what depends on a symbol, bounded by hop count and visit cap, with a citation per edge.

How it compares with Serena, Aider's repo map, CodeGraphContext, code-graph-rag, Sourcegraph, Cursor's index and Claude Code's own search, including when to use those instead: docs/comparison.md.

Five questions to start with

Ask your repo

What the graph adds

Where is authentication enforced?

the guard itself, plus the routes that call it

What calls <function>?

CALLS edges into it, each with a confidence score

What tests cover <module>?

IMPORTS edges from the test module back to the code under test

What would be affected by changing <api>?

the definition, then its direct callers from the CALLS edges into it (explain node)

Trace <a request> from route to persistence.

the handler and its callees one hop at a time, each block cited to file and line

GitHub Action

- uses: actions/checkout@v4
  with: { fetch-depth: 0 }   # full history, so CO_CHANGE edges are meaningful

- uses: Srinivasan-78/repo2graph@v3
  with:
    git-history: "500"
    commit-branch: graph     # optional: publish graph.html to a browsable branch

@v3 follows every 3.x release; pin an exact tag (@v3.0.0) to upgrade by hand. The Action never calls an LLM. Inputs and outputs: docs/cli.md.

Commands

Command

Does

repo2graph build <path> -o .r2g

Parse a repo into a graph and chunks (--incremental, --git-history N)

repo2graph query "<q>" -o .r2g

BM25 search plus one graph hop

repo2graph rag "<q>" -o .r2g

Budget-bounded, cited context pack (--answer sends it to an LLM: opt-in, the only path that sends code anywhere)

repo2graph explain <edge|node|retrieval>

Why an edge exists, or why a block was retrieved

repo2graph github <owner/repo> -o <dir>

Fetch, build and clean up without a local clone

repo2graph demo

Index a bundled example and answer the five questions above

repo2graph map, repo2graph stats, repo2graph index-status, repo2graph embed

Re-render graph.html, report counts and freshness, add optional dense vectors

repo2graph doctor, repo2graph bug-report, repo2graph explain-path, repo2graph completion

Diagnose setup, build a privacy-safe bug bundle, say why a path is (not) indexed, shell completion

repo2graph-mcp <path>

stdio MCP server: repo_map, repo_search, repo_neighbours, repo_blast_radius, and five more

Full flags: docs/cli.md. Python API: docs/python-api.md.

What it can't do

  • Resolve calls by type. Calls are matched by name, scoped by class, file, imports and directory. x.get() on a receiver of unknown type is recorded as a low-confidence guess, not a fact. For exact references, use a language-server tool.

  • See dynamic dispatch, reflection or computed imports. A missing edge doesn't prove that no call exists.

  • Cross language boundaries (Python calling C++ through bindings).

  • Rebuild itself when files change. index-status reports staleness; rebuild with build --incremental.

Symbols, calls and classes are extracted for Python, JS, TS, TSX, Go, Rust, Java, Ruby, C, C++, C#, PHP, Kotlin, Swift, Scala, Bash and Lua. Every other file is still indexed as text. The specific cases that defeat it — reflection dispatch, string-keyed registries, barrel re-exports — are pinned as known failures in the synthetic regression suite, and per-language coverage is scored, unevenly, in docs/architecture.md §4.

Two of those seventeen are benchmarked on real repositories: Python (Flask, requests, FastAPI) and TypeScript (Hono). Treat the rest as parsed-and-unmeasured. That is not a formality — the retrieval benchmark found two parse gaps in the one TypeScript repository as soon as it looked, one of which was hiding Hono's entire public API, and neither would have been visible without a real repository to ask questions about.

Status

The 2.x CLI, MCP tools and output schema follow semver: breaking changes wait for 3.0. Default paths run locally, send no telemetry and exclude secrets from agent replies unconditionally (what never leaves your machine, how credential files are excluded, reporting a vulnerability). A Docker image for read-only, non-root deployments is described in .github/SECURITY.md.

Contributing

git clone https://github.com/Srinivasan-78/repo2graph && cd repo2graph
python3 -m venv .venv && .venv/bin/pip install -e ".[dev]"
make lint format-check typecheck test

Branch from develop. Start with .github/CONTRIBUTING.md and docs/architecture.md. The most useful contribution right now is new questions for the retrieval benchmark, especially on repositories you know well. All docs: architecture, CLI, MCP, Python API, comparison.

MIT licensed.

Available Tools

9 tools
repo_blast_radiusA
Idempotent

Reverse reachability from a node_id: callers of callers, subclasses of subclasses, importers of importers of its file, and historically co-edited files, by hop distance -- the blast radius of changing it. Read-only, deterministic, zero side effects, secret files excluded. When to use: use before editing a symbol to see what depends on it, or to answer 'what is the blast radius of changing X'. When NOT to use: do not use for open-ended search (use repo_search) or to read a symbol's own body (use repo_read). Output: JSON object with callers, subclasses, importers, cochange, and a summary.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMaximum rows per section (default 40, max 50).
node_idYesGraph node id to analyze (e.g. from repo_find_symbol or a repo_search citation).
max_hopsNoReverse-closure depth (default 3, max 6).
include_cochangeNoInclude historically co-edited files (default true).

TDQS

A3.7/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description asserts 'Read-only, deterministic, zero side effects', but the annotations declare readOnlyHint=false, which signals the tool may mutate state. This is a direct contradiction between prose and structured metadata and is the most consequential thing an agent needs to resolve before calling.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core semantic, then a when-to-use/when-not-to-use block and an output summary. Dense but each sentence carries routing or behavioral information; the output line earns its place since no output schema exists.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a graph-analysis tool with no output schema, the description supplies the return shape (callers, subclasses, importers, cochange, summary), usage routing, and semantics. It would be complete were it not for the unresolved read-only claim, which leaves the tool's actual safety profile ambiguous.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so node_id, limit, max_hops, and include_cochange are all documented in the schema with defaults and maxima. The description adds only loose framing (hop distance, co-edited files) and no format or syntax detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a precise verb+resource with scope: 'Reverse reachability from a node_id' enumerating callers, subclasses, importers, and co-edited files by hop distance. It is immediately distinguishable from siblings like repo_search and repo_read, which it names explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use ('before editing a symbol', 'what is the blast radius of changing X') and when-NOT-to-use with named alternatives ('do not use for open-ended search (use repo_search) or to read a symbol's own body (use repo_read)'). This is exactly the routing guidance an agent needs.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

repo_build_statusA
Read-onlyIdempotent

Query progress and status of a background index build task under --async-build. Read-only check of in-memory background worker. When to use: use when polling build progress after an async index build was started. When NOT to use: do not use when building synchronously or when queries already succeed. Once completed, use repo_search or repo_map to query code. Output: JSON object with task_id, status, parsed file progress, and error details.

ParametersJSON Schema
NameRequiredDescriptionDefault
task_idYesTask ID string returned by a previous tool call when an asynchronous build was initiated.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false. The description adds useful behavioral context beyond annotations by stating it is a read-only check of an in-memory background worker and by specifying the output fields (task_id, status, parsed file progress, error details). This goes beyond what the annotations alone provide, though it does not disclose edge cases like unknown task IDs.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-organized with clear sections for usage, non-usage, follow-up, and output. Every sentence contributes useful information without redundancy, and the most important scoping detail (--async-build) is front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a simple one-parameter read-only tool with strong annotations, the description is complete: it explains the operation, when to call it, what not to do, what the output contains, and what to use afterward. The absence of an output schema is compensated by listing the expected JSON fields.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100% and the task_id parameter is already well documented as the ID returned by a previous async build call. The description does not add significant meaning beyond the schema, so the baseline score of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly identifies the tool as a status/progress query for a background index build task under --async-build, with a specific verb and resource. It also distinguishes itself from sibling tools by explicitly mentioning repo_search and repo_map as the post-completion alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit when-to-use guidance (polling after an async build was started) and when-not-to-use guidance (synchronous builds or when queries already succeed). It also names the correct follow-up tools, making the decision boundary unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

repo_cache_statsA
Idempotent

Retrieve runtime diagnostic counters for the tool result cache (hits, misses, size, max_size, ttl_s). Read-only, in-memory diagnostics, zero side effects. When to use: use when evaluating cache hit rate or debugging server performance. When NOT to use: do not use to search repository contents or inspect code structure; use repo_map or repo_search instead. Output: JSON object with cache metrics.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A3.9/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description asserts 'Read-only, in-memory diagnostics, zero side effects,' but the annotations declare readOnlyHint=false and openWorldHint=true, directly conflicting with that claim. An agent relying on the prose would misjudge the safety profile of this tool. This is a genuine annotation contradiction rather than ordinary missing context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Well front-loaded and organized with labeled 'When to use', 'When NOT to use', and 'Output' segments; each sentence is short and purposeful. Slight redundancy in piling 'Read-only, in-memory diagnostics, zero side effects' on top of annotations that already cover mutability, and that stacking is what creates the conflict.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no parameters and no output schema, the description supplies what an agent needs: what the tool reports, the named metrics, and where the sibling tools handle repo searching. The only real gap is that it does not reconcile its read-only framing with the annotation-declared openWorldHint/readOnlyHint values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so per the baseline the schema carries no semantics to explain. The description instead names the metric fields it returns (hits, misses, size, max_size, ttl_s), which is extra useful detail even though it is about output rather than input.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource — 'Retrieve runtime diagnostic counters for the tool result cache' — and enumerates the counters (hits, misses, size, max_size, ttl_s). It is clearly distinguishable from siblings like repo_map and repo_search, which it explicitly names as not-this.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit 'When to use' (evaluating cache hit rate, debugging server performance) and 'When NOT to use' (searching repository contents or inspecting code structure) with named alternatives (repo_map, repo_search). Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

repo_find_symbolA
Idempotent

Look up a symbol or file's node_id by name, for feeding into repo_neighbours, repo_read, repo_path_between or repo_blast_radius. Read-only, deterministic, zero side effects, secret files (.env) excluded. Matches an exact name first, then case-insensitively, then the last segment of a qualname; an ambiguous name returns every candidate with its path so you disambiguate rather than the server guessing. When to use: you already know a name (from a traceback, a grep, a review comment) and want its node_id with no repo_search round trip. When NOT to use: do not use for open-ended text search (use repo_search) or for repo layout (use repo_map). Output: JSON array of {node_id, name, qualname, kind, path, start_line, end_line, lang}.

ParametersJSON Schema
NameRequiredDescriptionDefault
kindNoOptional exact node kind filter (e.g. 'function', 'class', 'method').
nameYesSymbol or file name to look up (e.g. 'validate_token').
limitNoMaximum candidates returned (default 20, min 1, max 50).
path_prefixNoOptional path prefix filter, e.g. 'pkg/'.

TDQS

A3.7/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description asserts 'Read-only, deterministic, zero side effects', but the annotations declare readOnlyHint=false, i.e. the tool is not guaranteed side-effect-free. That is a direct contradiction of structured metadata, so the behavioral claims cannot be trusted as written. Other disclosed traits (matching precedence, secret-file exclusion, ambiguity handling) are genuinely useful but are undercut by the conflict.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action, then matching semantics, then when/when-not, then output shape. Dense but each block earns its place; the matching-precedence sentence is long but load-bearing for disambiguation behavior.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, and the description compensates by enumerating the returned fields ({node_id, name, qualname, kind, path, start_line, end_line, lang}). Ambiguity handling, matching precedence and secret-file exclusion are covered, leaving little an agent needs that is missing — apart from the unresolved read-only conflict.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so kind, name, limit and path_prefix are already documented with types, defaults and bounds. The description adds no per-parameter syntax or format detail beyond the schema, which is the expected baseline when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('look up a symbol or file's node_id by name') and immediately names the consumers (repo_neighbours, repo_read, repo_path_between, repo_blast_radius), so the agent can place it precisely in the tool graph.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit 'When to use' (name already known from a traceback/grep/review comment, want node_id with no repo_search round trip) and 'When NOT to use' (open-ended text search → repo_search; repo layout → repo_map). Alternatives and the selecting condition are both spelled out.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

repo_mapA
Idempotent

Retrieve a high-level structural map of the repository: languages, hub files, and top entry points. Read-only, deterministic, zero side effects. When to use: call this first at session start to understand codebase layout and identify entry points before detailed queries. Use when deciding where to investigate. When NOT to use: do not use to search code (use repo_search) or inspect call graphs (use repo_neighbours). Output: markdown summary of languages, hub files, and entry points, prefixed with a staleness note when the indexed working tree has changed since the index was built -- treat citations as suspect until rebuilt.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description asserts 'Read-only, deterministic, zero side effects,' but the annotations declare readOnlyHint=false. This directly contradicts the structured safety signal an agent relies on, so the behavioral disclosure is untrustworthy rather than merely incomplete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with purpose, then usage, then output, in a logical order with no filler. It is somewhat dense and the staleness caveat is long, but every sentence carries actionable content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema present, the description supplies the return shape ('markdown summary of languages, hub files, and entry points') plus a staleness caveat about changed working trees. That is exactly the context needed to call and interpret the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is nothing for the description to disambiguate; baseline 4 applies. The description correctly implies no inputs are required.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Retrieve a high-level structural map of the repository') and enumerates the content (languages, hub files, entry points). It is clearly distinguishable from siblings like repo_search and repo_neighbours, which it explicitly names.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use ('call this first at session start... before detailed queries'), explicit when-not-to-use with named alternatives ('do not use to search code (use repo_search) or inspect call graphs (use repo_neighbours)'). Routing is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

repo_neighboursA
Idempotent

Traverse code graph relationships from a known symbol or file node_id (callers, callees, base classes, definitions). Read-only, deterministic traversal, no side effects. When to use: use with a specific node_id (e.g. from repo_search citations) to inspect callers (CALLS in), callees (CALLS out), inheritance, or definitions. When NOT to use: do not use for text search across code (use repo_search) or repo overview (use repo_map). Output: markdown list formatted as - <EDGE_TYPE> <in|out>: <name> (<path:line>) [<node_id>].

ParametersJSON Schema
NameRequiredDescriptionDefault
hopsNoTraversal depth from node_id (default 1, max 4).
limitNoMaximum neighbor rows to return (default 20, min 1, max 50).
node_idYesTarget graph node identifier to expand from (e.g. 'sym:pkg/mod.py::func', 'file:pkg/mod.py', 'dir:pkg').

TDQS

A3.8/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description asserts 'Read-only, deterministic traversal, no side effects', but the annotations declare readOnlyHint=false. This is a direct conflict on the tool's mutation profile, so the description cannot be trusted as behavioral disclosure. It also fails to reconcile the openWorldHint=true annotation with the claim of a fully deterministic, side-effect-free traversal.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and behavior are front-loaded, and the labeled When to use / When NOT to use / Output sections are easy to scan with no filler. Slightly over-specified for its length, and one sentence carries the erroneous read-only claim, but overall it is tight and well ordered.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description compensates by specifying the return format ('- <EDGE_TYPE> <in|out>: <name> (<path:line>) [<node_id>]') and the edge direction convention (CALLS in vs CALLS out). Combined with the routing guidance and parameter coverage, an agent has everything needed to call and interpret this tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%: hops, limit, and node_id each carry their own defaults, bounds, and format examples in the schema. The description only adds the provenance hint for node_id (comes from repo_search citations) and does not clarify hops/limit interaction with the traversal, so baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb (traverse) and resource (code graph relationships from a node_id) and enumerates the edge kinds covered (callers, callees, base classes, definitions). It is immediately distinguishable from repo_search (text search) and repo_map (overview), which are named explicitly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives explicit 'When to use' and 'When NOT to use' sections, names the concrete trigger (a node_id obtained from repo_search citations), and routes the agent to the correct alternatives (repo_search for text search, repo_map for overview). Nothing about tool selection is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

repo_path_betweenA
Idempotent

Find a bounded, bidirectional path between two node_ids over CALLS/DEFINES/IMPORTS edges (CO_CHANGE only if named explicitly), reporting the minimum confidence along each path. Read-only, deterministic traversal, no side effects. When to use: use for 'how does X reach Y' questions a text search cannot answer, e.g. the call chain from an HTTP handler to a database write. When NOT to use: do not use for one-hop neighbours (use repo_neighbours) or open-ended search (use repo_search). Output: JSON array of paths, each an ordered array of {node_id, path, start_line, end_line, via_edge, confidence}.

ParametersJSON Schema
NameRequiredDescriptionDefault
to_idYesTarget graph node id.
from_idYesStarting graph node id.
max_hopsNoMaximum path length in hops (default 6, max 8).
max_pathsNoMaximum distinct paths returned (default 3, max 10).
edge_typesNoEdge types to traverse (default ['CALLS', 'DEFINES', 'IMPORTS']). CO_CHANGE is a statistical correlation, not a call/definition path, and is only followed when named here explicitly.

TDQS

A3.9/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description asserts 'Read-only, deterministic traversal, no side effects,' but the annotations declare readOnlyHint=false, which tells an agent the opposite. These two signals directly conflict and an agent cannot tell whether invoking this tool may mutate state. The conflict is flaggable under the contradiction rule even though other behavioral details (bounded traversal, min-confidence reporting) are otherwise informative.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the operation, then clearly delimited 'When to use', 'When NOT to use', and 'Output' segments. Dense but every sentence carries information an agent needs; no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly supplies the return shape (JSON array of ordered {node_id, path, start_line, end_line, via_edge, confidence} objects), plus traversal bounds and default edge types. An agent has everything needed to call and interpret results — the only defect is the safety-signal conflict noted above.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% — every parameter, default, and the CO_CHANGE caveat is documented in the schema itself. The description only restates the CO_CHANGE rule already present in the edge_types field, adding no new syntax or format guidance. Baseline 3 applies when the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource+scope: find a bounded bidirectional path between two node_ids over named edge types, reporting minimum confidence per path. It distinguishes itself from siblings by explicitly naming repo_neighbours and repo_search as the wrong tools for adjacent tasks.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit 'When to use' with a concrete example (HTTP handler to database write) and an explicit 'When NOT to use' that routes to the correct alternatives (repo_neighbours for one hop, repo_search for open-ended search). Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

repo_readA
Idempotent

Read a widened window of source text around a citation, from the indexed chunks rather than the filesystem -- works over HTTP, from a different machine, with no shared filesystem, because chunks have already passed secret-path exclusion and redaction. Read-only, deterministic, zero side effects. When to use: use after repo_search or repo_neighbours to see more lines around a [cite: path:start-end] citation, with optional context lines each side. When NOT to use: do not use for a path never indexed, or an absolute or '..' path (both are refused); use repo_search or repo_find_symbol to find a valid path first. Output: [cite: path:start-end] header plus the text, bounded in size.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathYesRepo-relative path exactly as it appears in a [cite: path:start-end] header. Absolute paths and '..' segments are refused.
contextNoExtra lines of context on each side of the range (default 0, max 500).
end_lineNoLast line to read (default: start_line).
start_lineNoFirst line to read (default 1).

TDQS

A3.8/5.0
Behavior1/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description asserts 'Read-only, deterministic, zero side effects', but the annotations declare readOnlyHint=false, i.e. the tool may modify its environment. That is a same-axis conflict between the prose and the structured hints, so an agent gets contradictory safety signals. The description does add genuinely useful context (indexed chunks have passed secret-path exclusion and redaction, absolute/'..' paths are refused), but the direct contradiction caps this dimension.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core action and scope, then cleanly segmented into when-to-use, when-not-to-use, and output. Slightly dense in the opening clause about HTTP/no-shared-filesystem, but every sentence carries information an agent needs.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description supplies the return shape ('[cite: path:start-end]' header plus bounded text), the error conditions (refused paths), and the workflow placement relative to repo_search/repo_neighbours. Nothing essential for correct invocation is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so path/context/start_line/end_line are already documented with defaults and the 500-line cap. The description adds only the path-provenance rule (must match a [cite: path:start-end] header) and the refusal behavior, which is baseline-level value over an already complete schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource (read a widened window of source text around a citation) and immediately scopes it ('from the indexed chunks rather than the filesystem'). It explicitly distinguishes itself from sibling tools repo_search and repo_neighbours, so an agent can route correctly without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit 'When to use' (after repo_search or repo_neighbours to expand a [cite: path:start-end]) and 'When NOT to use' (never-indexed path, absolute or '..' path), plus the named fallbacks repo_search and repo_find_symbol. Both the positive and negative conditions are spelled out, not implied.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 7 tool updatesv3.0.0
    • Addedrepo_blast_radius
    • Addedrepo_find_symbol
    • Removedrepo_impact
    • Changedrepo_neighbours1 field changed
      • changedInput schema / properties / limit / description
        Previous value: -"Maximum neighbor rows to return (default 20, max 50)."New value: +"Maximum neighbor rows to return (default 20, min 1, max 50)."
    • Addedrepo_path_between
    • Addedrepo_read
    • Changedrepo_search1 field changed
      • changedInput schema / properties / budget_tokens / description
        Previous value: -"Maximum token ceiling for returned markdown pack (default 6000, max 12000)."New value: +"Maximum token ceiling for returned markdown pack (default 6000, max 12000; zero or negative uses the default)."
  2. 1 tool updatev2.2.0
    • Addedrepo_impact
  3. 3 tool updatesv1.5.1
    • Changedrepo_build_status1 field changed
      • changedInput schema / properties / task_id / description
        Previous value: -"the id a previous call returned"New value: +"Task ID string returned by a previous tool call when an asynchronous build was initiated."
    • Changedrepo_neighbours3 fields changed
      • changedInput schema / properties / hops / description
        Previous value: -"graph hops (default 1, max 4)"New value: +"Traversal depth from node_id (default 1, max 4)."
      • changedInput schema / properties / limit / description
        Previous value: -"neighbours (default 20, max 50)"New value: +"Maximum neighbor rows to return (default 20, max 50)."
      • changedInput schema / properties / node_id / description
        Previous value: -"e.g. sym:pkg/a.py::run"New value: +"Target graph node identifier to expand from (e.g. 'sym:pkg/mod.py::func', 'file:pkg/mod.py', 'dir:pkg')."
    • Changedrepo_search4 fields changed
      • changedInput schema / properties / budget_tokens / description
        Previous value: -"max 12000"New value: +"Maximum token ceiling for returned markdown pack (default 6000, max 12000)."
      • changedInput schema / properties / hops / description
        Previous value: -"graph hops (default 1, max 4)"New value: +"Graph traversal depth around seed chunks (default 1, max 4; 0 returns seeds only)."
      • changedInput schema / properties / k / description
        Previous value: -"seed chunks (default 8, max 50)"New value: +"Number of initial seed chunks retrieved via BM25 lexical scoring (default 8, max 50)."
      • changedInput schema / properties / query / description
        Previous value: -"the question"New value: +"Natural language question, search terms, or symbol identifier to search for (e.g. 'pack_context' or 'how does export work')."
  4. 5 tool updatesv0.1.0
    • First observedrepo_build_status
    • First observedrepo_cache_stats
    • First observedrepo_map
    • First observedrepo_neighbours
    • First observedrepo_search

TDQS

A4.1/5.0

Scored across 9 tools

Disambiguation5/5

Each tool targets a distinct resource+action, and the descriptions explicitly cross-reference siblings (repo_search vs repo_neighbours vs repo_map; one-hop neighbours vs multi-hop repo_path_between vs reverse repo_blast_radius). Overlapping graph-traversal tools are cleanly separated by their 'When NOT to use' clauses, leaving no realistic misselection.

Naming Consistency5/5

Every tool uses the same repo_ prefix with a consistent snake_case noun/verb phrase (repo_map, repo_search, repo_find_symbol, repo_blast_radius). No camelCase or verb-style mixing.

Tool Count5/5

Nine tools is well-scoped for a read-only code-graph service: layout, search, symbol lookup, read, and several traversal/analysis operations, plus two diagnostics. Each tool earns its place and none is redundant.

Completeness4/5

Coverage of the read-only query lifecycle is strong: map, search, symbol resolution, read, neighbours, path-finding, and blast radius. Minor gaps remain — no tool to initiate an index build (only repo_build_status to poll) and no direct file/listing operation — but these are workable for the stated purpose.

Maintenance

ActivityActive
ResponsivenessWithin a week

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    A
    maintenance
    An MCP code-intelligence server for AI agents with pre-indexed AST cache, 62 MCP tools, and TOON-compressed output, enabling token-efficient code analysis and project health grading entirely locally.
    9
    54
    MIT
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables LLM agents to query a codebase's structural knowledge (symbols, imports, call graphs, etc.) via MCP, reducing tokens and improving correctness compared to raw file access.
    583 npm
    7
    MIT
  • F
    license
    Not graded
    quality
    A
    maintenance
    Enables agents to build and query code knowledge graphs for repositories in a folder — finding shortest paths between concepts, explaining concepts with neighbours and community context, and visualizing per-repo graphs through MCP tools.
    -
  • F
    license
    Not graded
    quality
    C
    maintenance
    Enables coding agents to query a local-first code intelligence graph of Python repositories—covering functions, classes, modules, and their relationships—via MCP, supporting subgraph retrieval, caller lookup, and impact analysis without re-reading the codebase.
    -