Skip to main content
Glama
nagendra-kon

repodigest-mcp

by nagendra-kon

repodigest-mcp

Token-efficient code context for AI coding assistants: Python signatures, call graphs and budget-packed context, served locally over MCP.

License: MIT Python 3.10+ MCP 2.x Tests: 61 passing

repodigest-mcp is a Model Context Protocol server that wraps RepoDigest's static analysis (AST parsing, call graph, BM25 search, token-budgeted packing) so that Claude Code, Cursor, or any MCP client can ask for exactly the slice of a codebase it needs instead of reading whole files.


The problem

Coding assistants explore a repository by reading files. To learn what one function accepts and returns, the assistant pulls in the entire file: every other function, every body, every import. On a real module that is routinely 90%+ noise for the question being asked, and it compounds over a session:

  • Context bloat. The window fills with code that doesn't matter, leaving less room for the code that does.

  • Cost and latency. You pay for every token in, on every subsequent turn.

  • Worse answers. Relevant details get buried in long contexts.

Most of what an assistant needs is structural, and structure can be extracted deterministically. A function's interface is its signature and docstring. Its blast radius is its callers and callees. The code relevant to a task is a neighbourhood in the call graph around the best-matching symbol. None of that needs a model; it needs an AST.

repodigest-mcp exposes those three operations as MCP tools. It runs as a local stdio subprocess: no daemon, no cloud service, no telemetry, and your source is parsed on your machine. (The only network access is tiktoken fetching its tokenizer vocabulary once on first use, after which it is cached.)

Related MCP server: pyscope-mcp

Tools

All three tools are read-only. Failures come back as MCP tool errors (is_error=True) with a message the model can act on, never as a crash or an empty result.

Tool

Purpose

get_symbol_signature

A function/class/method's signature and docstring, without the body

get_symbol_dependencies

Direct callers and callees, across files

pack_task_context

Best-matching code for a task, packed into a token budget

symbol_name accepts a fully-qualified name (pkg.mod.Class.method) or any dotted suffix (Class.method, method). path is the repository root and defaults to ., the server's working directory.

get_symbol_signature(symbol_name, path)

Returns the definition header and docstring only. Classes come back with every method stubbed to its signature. For a method in a large file this is typically 90%+ fewer tokens than reading the file (measured below).

# repodigest.search.ranker.SymbolRanker.score (repodigest/search/ranker.py)
def score(self, query: str, symbol: str) -> float:
    ...

The first line names the symbol the request resolved to and where it lives, which matters when you passed a short suffix.

get_symbol_dependencies(symbol_name, path)

Direct callers and callees from the repository call graph, across files. Returned as structured content:

{
  "symbol": "repodigest.search.ranker.SymbolRanker.rank",
  "kind": "method",
  "file": "repodigest/search/ranker.py",
  "callers": ["repodigest.search.ranker.SymbolRanker.top"],
  "callees": ["repodigest.search.ranker.SymbolRanker.score"]
}

Edges are matched by simple name, with no import or type resolution. Common names (get, run) can produce false positives, and Foo() links to the class Foo, not to Foo.__init__.

pack_task_context(query, budget, signatures_only, path)

Given a natural-language query, the tool:

  1. ranks every symbol in the repo with BM25 and takes the best match as the root;

  2. expands breadth-first through the call graph, alternating callees and callers, nearest first;

  3. packs symbols until the token budget is spent, skipping (never truncating) anything that no longer fits;

  4. returns compact XML.

Parameter

Default

Meaning

query

required

What you're working on, e.g. "score symbols with BM25"

budget

2000

Max tokens (cl100k_base) of packed source

signatures_only

false

Pack everything except the root as signature + docstring

path

.

Repository root

A real call, query="score symbols with BM25", budget=1500, signatures_only=true (abbreviated with …):

<context>
  <file path="…/repodigest/search/ranker.py">
    <symbol name="repodigest.search.ranker.SymbolRanker" kind="class" tokens="525">
<![CDATA[
class SymbolRanker:
    """BM25 ranking over a corpus of `{symbol: text}` documents."""
    …
]]>
    </symbol>
  </file>
  <file path="…/repodigest/cli.py">
    <symbol name="repodigest.cli.pack_command" kind="function" tokens="247">
      …
    </symbol>
  </file>
  <usage total_tokens="772" budget="1500" symbols_packed="2" symbols_skipped="0" />
</context>

The root is always included in full. Here it is a class, followed by one of its callers.

budget counts the packed source only. The XML tags around it are not counted, so leave roughly 10% headroom.

If the best match cannot fit, the tool says so instead of returning something misleading. This is a real response from the same repository at budget=500:

Error executing tool pack_task_context: Best match 'repodigest.search.ranker.SymbolRanker' needs 525 tokens but the budget is 500; raise budget to at least 525.

Errors the tools handle

Situation

Behaviour

Symbol not found

Error, with "did you mean" suggestions for near misses

Ambiguous suffix (run matches 2 methods)

Error listing the candidates (first 5, then +N more)

path missing, not a directory, or empty

Error naming the path

Directory with no Python symbols

Error, rather than an empty result

Empty symbol_name / query, budget < 1

Error explaining the constraint

Root symbol larger than budget

Error stating the budget needed

Unparsable or non-UTF-8 .py file

Skipped with a warning on stderr; the rest is indexed

Architecture

MCP client (Claude Code, Cursor, ...)
        │  JSON-RPC over stdio
        ▼
server.py    three read-only tools; translates failures into tool errors
        │
        ▼
index.py     cached per-repo index: symbol registry · call graph · BM25 ranker
        │    file discovery, error-tolerant parsing, symbol resolution
        ▼
RepoDigest   py_parser · CallGraph · SymbolRanker · ContextPacker

RepoDigest does the analysis. index.py decides which files it sees and remembers the result.

Optimizations

mtime/size-cached indexer. Each repo root gets one index (registry, call graph, ranker), stamped with a fingerprint of (path, mtime, size) for every Python file. A repeated call against an unchanged tree reuses the index. Editing, adding or deleting a file changes the fingerprint and triggers a rebuild on the next call. A warm lookup on the RepoDigest repo took under 5 ms.

Virtual environments and vendored code are excluded automatically. Discovery prunes:

  • hidden directories (.venv, .git, .tox, .cache, ...);

  • any directory containing a pyvenv.cfg, so virtualenvs are caught whatever they are named (venv, myenv), while an ordinary package that merely happens to be called env is kept;

  • site-packages, node_modules and __pycache__.

This matters because RepoDigest's own directory walker globs every *.py under the root. Pointed at this project's directory, it collected 1,825 files (nearly all of them from the virtualenv) and took about 5 s. repodigest-mcp indexed the same directory in 0.03 s, and none of the virtualenv's symbols leaked into search results. If you deliberately pass a virtualenv as the root, it is indexed.

Fault-tolerant parsing. One file with a syntax error or a stray non-UTF-8 byte does not take down the index. That file is skipped and logged.

Protocol-safe logging. On stdio, stdout belongs to the protocol. All logging goes to stderr; use repodigest-mcp --log-level DEBUG to see more.

Token efficiency

Measured on RepoDigest's own source with cl100k_base. "Signature" is the exact string get_symbol_signature returns, header line included. "Saved" compares it to reading the file the symbol lives in.

Symbol

Signature

Symbol source

Whole file

Saved vs. file

ContextPacker.pack

40

200

1,221

96.7%

SymbolRanker.score

40

146

766

94.8%

CallGraph.from_files

48

346

877

94.5%

parse_source

110

196

1,803

93.9%

to_xml

37

189

450

91.8%

ContextPacker (class)

166

677

1,221

86.4%

Across these six symbols the saving versus reading the whole file is 86% to 97%. The gap narrows against the symbol's own body (44% to 86% here) because the body of a short function is not much larger than its signature: the big win comes from not reading the rest of the file. Your numbers will vary with file size and docstring density.

Installation

Requires Python 3.10+. repodigest is not published to PyPI, so install it from GitHub first, then install this package in editable mode:

git clone https://github.com/nagendra-kon/repodigest-mcp.git
cd repodigest-mcp

python3 -m venv venv
source venv/bin/activate

pip install "repodigest @ git+https://github.com/nagendra-kon/repodigest.git"   # the analysis engine
pip install -e ".[dev]"                                                           # this server + pytest

repodigest-mcp --version

If you already have a local RepoDigest checkout, pip install -e ../repodigest works in place of the GitHub line.

Claude Code

Register the server with the absolute path to the venv's executable, because the client launches it as a subprocess and it must use the interpreter that has the dependencies installed. From the repodigest-mcp directory:

claude mcp add repodigest -- "$(pwd)/venv/bin/repodigest-mcp"

Add --scope user to make it available in every project, or --scope project to share it via .mcp.json. Check it with claude mcp list.

Cursor

Add the server to .cursor/mcp.json in your project (or ~/.cursor/mcp.json for all projects):

{
  "mcpServers": {
    "repodigest": {
      "command": "/absolute/path/to/repodigest-mcp/venv/bin/repodigest-mcp",
      "args": []
    }
  }
}

Which repository does it read?

Tools default to path=".", the server's working directory. If your client starts the server somewhere other than the project you're working on, pass path explicitly (for example, tell the assistant which directory to use) or start the server from the right directory. The server is read-only, but path is not sandboxed: it will read .py files under any directory the client names.

Try it

Once registered, ask your assistant things like:

  • "Show me the signature of AuthService.login."

  • "What calls hash_password, and what does it call?"

  • "Pack context for adding rate limiting to the login handler, in 1,500 tokens."

Testing

pytest tests/ -v          # 61 tests, ~2 s

The suite runs against a synthetic multi-file project built in a temp directory. That project includes a fake virtualenv, a hidden directory, node_modules, a file with a syntax error and a non-UTF-8 file, so the exclusion and fault-tolerance paths are exercised for real.

File

Tests

Covers

tests/test_server.py

38

MCP client integration: every tool called through a real in-process MCP client session (schema, structured output, is_error results), plus a subprocess transport test that launches the installed repodigest-mcp console script over stdio and calls a tool. Also token-budget guarantees, signatures_only, ambiguous and unknown symbols, and invalid paths.

tests/test_index.py

23

Indexing edge cases: virtualenv, hidden-dir and vendored-dir exclusion (venv detected by pyvenv.cfg, not by name), root validation, skipped unparsable and undecodable files, cache reuse, invalidation on edit / add / delete, symbol resolution and BM25 matching.

The budget tests assert that reported total_tokens never exceeds budget, that the per-symbol token counts sum to the total, that smaller budgets pack fewer symbols and report what was skipped, and that a root symbol that cannot fit is an error rather than a silent empty result. Integration tests use the official MCP Python client.

Limitations

  • Python only. Symbols are top-level functions, classes and methods; nested functions are not indexed separately.

  • Name-based call graph. See the note under get_symbol_dependencies.

  • Token counts are a proxy. They use tiktoken's cl100k_base, not Claude's own tokenizer, so treat budgets as close approximations. The budget also excludes the XML markup.

  • mcp>=2.0 only. MCP SDK 2.x renamed FastMCP to MCPServer; this package targets the 2.x API.

License

MIT. Built on RepoDigest.

Available Tools

3 tools
get_symbol_dependenciesA
Read-only

List a symbol's direct callers and callees across the repository.

Edges are matched by simple name (no import or type resolution), so common names such as get or run can produce false positives. Class instantiation Foo() links to the class Foo, not to Foo.__init__.

Args: symbol_name: Qualified name (pkg.mod.Class.method) or any dotted suffix of it. path: Repository root to search (default: the server's working directory).

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNo.
symbol_nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription

No output parameters

TDQS

A4.8/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds valuable behavioral detail beyond the readOnlyHint annotation: edges are matched by simple name with no import resolution, common names can cause false positives, and class instantiation links to the class rather than __init__. This meaningfully prepares the agent for surprising results.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded, with the core action first, followed by essential caveats and parameter documentation. Every sentence adds useful information; there is no filler or redundant restatement.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool is simple (2 params, 1 required) and has an output schema, so the description does not need to explain return values. It covers purpose, matching behavior, limitations, and both parameters, making it complete for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, but the description fully compensates by documenting both parameters: symbol_name accepts a qualified name or dotted suffix, and path specifies the repository root with a default. This is exactly the semantic meaning the schema lacks.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description starts with a specific verb and resource: 'List a symbol's direct callers and callees across the repository.' It clearly differentiates this from siblings like get_symbol_signature (signature lookup) and pack_task_context (task context packaging).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly conveys when to use the tool: when you need direct callers/callees across the repository. It does not explicitly name alternatives or exclusions, but the scope is unambiguous and the caveats about simple-name matching help set expectations.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_symbol_signatureA
Read-only

Return a Python symbol's signature and docstring without its body.

Far cheaper than reading the file. Works for functions, classes (methods listed as signatures) and methods.

Args: symbol_name: Qualified name (pkg.mod.Class.method) or any dotted suffix of it (Class.method, method). Ambiguous suffixes return the candidates. path: Repository root to search (default: the server's working directory).

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNo.
symbol_nameYes

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds meaningful behavioral detail beyond the readOnlyHint: it excludes the function body, handles classes by listing methods as signatures, and returns candidates for ambiguous suffixes. These are concrete, non-obvious behaviors that help an agent predict results without contradicting the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is compact and front-loaded with the core purpose, followed by a cost/benefit note and efficient parameter semantics. Every sentence adds value and there is no redundant or filler content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the output schema exists, the description does not need to explain return values. It covers purpose, behavior, ambiguity handling, and both parameters, leaving no obvious gap for an agent to invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden. It thoroughly explains symbol_name, including qualified names, dotted suffixes, and ambiguity resolution, and documents path's default as the server working directory. This fully compensates for the schema's lack of descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with the specific action and resource: "Return a Python symbol's signature and docstring without its body." It clearly defines what the tool does and distinguishes it from siblings like get_symbol_dependencies by focusing on signatures rather than dependencies and by explicitly excluding the body.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

"Far cheaper than reading the file" gives a clear when-to-use context for retrieving signature-only information. However, it does not explicitly contrast with sibling tools or state when not to use it, so it stops short of full usage routing.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

pack_task_contextA
Read-only

Pack the code most relevant to a task into a token-bounded XML snippet.

Ranks every symbol against query (BM25), takes the best match as the root, then breadth-first expands through its callees and callers, skipping (never truncating) any symbol that no longer fits.

Args: query: Natural-language description of the task, e.g. "validate user login". budget: Max tokens (cl100k_base) of packed source. The XML markup around the source (file/symbol tags, usage summary) is not counted, so allow ~10% headroom. signatures_only: Pack everything except the root symbol as signature + docstring. path: Repository root to search (default: the server's working directory).

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNo.
queryYes
budgetNo
signatures_onlyNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description goes well beyond the readOnlyHint annotation, disclosing the BM25 ranking, root selection, breadth-first expansion through callees and callers, and the skip-never-truncate behavior. It also explains the budget semantics, noting that XML markup is not counted and advising ~10% headroom, plus the signatures_only behavior. This is rich behavioral detail.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well organized: a one-sentence purpose, a compact algorithm summary, and a clearly separated Args block. Every sentence contributes useful information, and the most important behavioral facts are front-loaded before parameter details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex 4-parameter tool with no schema descriptions, the description covers purpose, algorithm, parameter semantics, edge-case behavior (skipping vs truncating), and budget headroom. The output schema exists, so return-value details don't need to be in the description; nothing critical is missing for correct invocation.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Despite 0% schema description coverage, the description fully compensates by explaining every parameter: query with an example, budget with tokenization and markup caveats, signatures_only with its exact packing behavior, and path with its default and scope. This gives an agent all the semantic information needed to set parameters correctly.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb and resource: 'pack the code most relevant to a task into a token-bounded XML snippet.' It also distinguishes the tool from sibling symbol-focused tools by describing a query-based ranking and breadth-first expansion over callees and callers, which is a different scope than simply fetching a signature or dependency list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The intended use is implied through the purpose and the `query` argument ('Natural-language description of the task'), but the description does not explicitly explain when to choose this tool over get_symbol_signature or get_symbol_dependencies, nor does it state when not to use it. There are no exclusions or alternative-routing hints.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv0.1.0
    • First observedget_symbol_dependencies
    • First observedget_symbol_signature
    • First observedpack_task_context

TDQS

A4.6/5.0

Scored across 3 tools

Disambiguation4/5

get_symbol_signature and get_symbol_dependencies share the same argument shape and both target a single symbol, so an agent could briefly confuse them; however, their outputs are clearly different (docstring vs. call graph). pack_task_context is unmistakably distinct in purpose.

Naming Consistency4/5

The get_symbol_* prefix establishes a clear pattern for the two lookup tools, and pack_task_context uses the same snake_case verb_noun style. The one-off pack_ verb is a minor deviation rather than a chaotic mix.

Tool Count5/5

Three tools is a focused, purposeful set for a code-digestion server: signature lookup, dependency lookup, and context packing. Each tool fills a clear niche without redundancy.

Completeness4/5

The core workflow of inspecting a symbol and packing relevant code for a task is well covered. Missing capabilities like repository-wide symbol discovery or batch queries are minor gaps since pack_task_context effectively provides search.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Provides AI assistants with a structured, token-efficient map of a codebase's symbols, dependencies, and relationships via MCP tools like overview, query, and impact analysis.
    8
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    MCP server that exposes Python function- and module-level call graphs for agentic coding clients, enabling tools like callers_of, callees_of, and neighborhood queries.
    MIT
  • A
    license
    A
    quality
    D
    maintenance
    Universal MCP server that analyzes any codebase and provides structured context to AI assistants. Dynamic, accurate, and token-efficient.
    18
    18 npm
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    Serves structured code context via MCP, enabling AI agents to understand codebases with dependency graphs and significantly reduce token usage.
    20 npm
    MIT