citegraph
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@citegraphwho calls place_order and what does it call?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
citegraph
A code map for your AI coding assistant. Ask "what calls this?" or "how does this endpoint reach that code?" and get the exact chain with file and line numbers, even when it crosses services and languages.
citegraph reads your Python and C# repositories once and records which functions call which, including calls that travel through a job queue from one service to another. It serves that map to Claude Code (or any MCP client, such as Claude Desktop or Cursor) as a set of tools. It runs on your machine, never returns source code, and every answer cites its evidence and says how sure it is.
See it in action
A C# API enqueues a job. A Python worker runs it and calls into an analysis package. Three projects, two languages, and nothing in the code a text search can follow from one end to the other:
// components/bff/src/Api/VersionEndpoints.cs (DramatiqTasks.ClassifyStems = "classify_stems")
await queue.EnqueueAsync(DramatiqTasks.ClassifyStems, new object[] { versionId.ToString() });# components/worker/app/tasks.py
@dramatiq.actor(actor_name="classify_stems")
def classify_stems(version_id):
from audio.stems import classify_stems as classify_audio # re-exported by audio/stems/__init__.py
return classify_audio([version_id])You: How does the
ClassifyStemsendpoint end up inclassify_one?
find_path("Site.Api.Endpoints.VersionEndpoints.ClassifyStems", "audio.stems.classify.classify_one")
step project called at rule confidence
Site.Api.Endpoints.VersionEndpoints.ClassifyStems components/bff
→ app.tasks.classify_stems components/worker src/Api/VersionEndpoints.cs:7 queue_match 0.9
→ audio.stems.classify.classify_stems components/analysis app/tasks.py:8 import_scope 0.9
→ audio.stems.classify.classify_one components/analysis src/audio/stems/classify.py:6 same_file 0.95That is real output (the example in tests/test_cross_service.py, shown as a table). citegraph resolved the C# constant to its value, matched it to the Python actor with that name, and followed the package re-export. Your assistant gets the whole chain in one call and opens only the three lines that matter.
It is just as useful inside one repo. Asked who calls flask.json.loads in flask 3.1.3, citegraph returns exactly
the five real callers; a text search for the name finds 21 functions, which the assistant would otherwise open
one by one to tell callers from look-alikes. In C#, calls through interfaces land on the right method because
citegraph follows the declared types of fields, parameters and locals.
Related MCP server: codegraph-mcp
What you can ask
Ask | Tool |
"How does the |
|
"What calls |
|
"What does |
|
"Which endpoint triggers the |
|
"Where is |
|
"Give me an overview of the billing repo." |
|
"Why do you think A calls B?" |
|
Find a symbol, show one, check what is indexed |
|
Every answer carries path:line and the git commit it came from, the rule that linked the two pieces of code, and a
confidence. Name-only guesses are hidden unless asked for, and answers from an index older than the repo are
flagged as stale.
Quick start
You need Python 3.13+, git and uv.
# 1. Install the citegraph command
uv tool install git+https://github.com/rankinbc/citegraph
# 2. Index one repo, or a folder that holds several (seconds; re-run after you pull, only changed files are re-read)
citegraph index ~/src/my-repos
# 3. Connect it to Claude Code
claude mcp add citegraph -- citegraph serve --root ~/src/my-reposThen ask questions as usual; run /mcp in Claude Code to check that citegraph is connected. The same tools work
from the terminal: citegraph query what_calls symbol=place_order --root ~/src/my-repos.
Several services in one repo? Add projects = "auto" to a citegraph.toml in the folder you index, and each
service is indexed as its own project. Job queues (Dramatiq, with C# or Python senders) are linked automatically;
anything else, such as an HTTP call, can be declared in a citegraph.overrides.yaml
(guide).
Setup for other MCP clients, team setup, configuration and tips: docs/guide.md.
How it works
sequenceDiagram
actor You
participant Claude as Claude Code
participant CG as citegraph (local MCP server)
participant Index as Local index (names and locations only)
You->>Claude: "What calls OrderService.place_order?"
Claude->>CG: what_calls(symbol="OrderService.place_order")
CG->>Index: look up resolved call edges
Index-->>CG: callers, each with its rule and confidence
CG-->>Claude: callers with path:line, commit, confidence
Claude->>Claude: opens only those lines to confirm
Claude-->>You: answer that cites each caller's file and lineIndex.
citegraph indexparses every git-tracked Python and C# file with tree-sitter and stores symbols, references, imports and config key names in a local SQLite file. About 3 seconds for 35k lines.Resolve. Each reference becomes an edge through a named rule (same file, import, declared type, queue match, hand-written link, unique name, or a name-only guess), and the rule sets the edge's confidence.
Answer.
citegraph serveexposes nine read-only tools over MCP, in milliseconds.
Safe to leave connected
Stores names and locations only, never source code or config values. The one stored value: name-shaped C#
const stringjob names, redacted like everything else.Redacts anything that looks like a secret, on the way into the index and again on the way out.
Read-only: no write or shell tools. Every call is logged locally (
citegraph audit tail).Runs on your machine. Indexing and serving make no network calls; your assistant sees only the answers it asks for.
How accurate is it?
A deterministic eval asks citegraph and a grep baseline the same 50 questions about two pinned public Python projects (flask and httpx), with answers labeled from a compiler-grade index:
question | citegraph F1 | grep F1 |
what calls X | 0.78 | 0.69 |
what does X call | 0.84 | 0.58 |
where is config key K | 0.60 | 0.86 |
citegraph wins on precision: grep finds every caller but buries them among look-alikes. Config keys are a known weak spot, where plain search still does better. Each rule's confidence is checked against its measured precision, and a subset of the eval runs in CI as a regression gate. C# is not yet benchmarked. Details: guide, eval report.
Limitations
Python and C# only (TypeScript is next).
Static analysis, not a compiler: dynamic dispatch, reflection and dependency-injection wiring are invisible or matched by name, with lower confidence.
Python method calls on untyped variables resolve by name only, the main cause of missed callers.
No generic type inference in C#; extension methods resolve by name.
Links between services: Dramatiq job queues and hand-written links only; HTTP calls between services are not linked automatically yet.
Full list: docs/guide.md#limitations.
Under the hood
An API designed for an agent. All nine tools return one envelope (
data,evidence,confidence,stale,notes), and errors carry a hint and candidates so the agent can recover on its own. mcp/server.pyCalibrated graph resolution. When the eval showed name-only matches were right 17% of the time, not the assumed 50%, the confidence was lowered to match and those edges were hidden by default. resolve/
Security by construction. One sanitizing write path, a leak scan after every run, and a schema with no column that could hold code. store/db.py
Honest evaluation. The eval caught a bug in its own baseline that made citegraph look ten times better on one tool; the corrected numbers are the ones published. analysis
Quality bar. 350+ tests, strict pyright, ruff, CI on Ubuntu and Windows.
More: architecture, design notes, full design.
Roadmap
HTTP links between services (queue links and hand-written links ship today).
A TypeScript extractor.
A C# eval: scip-dotnet labels, a C# golden set and grep baseline, calibrated C# confidence and a CI gate.
An agent-level eval: does an agent answer better and cheaper with citegraph than with grep alone?
License
MIT
Available Tools
9 toolsexplain_edgeB
Why citegraph believes from_symbol calls to_symbol: rule, meaning, confidence, evidence.
| Name | Required | Description | Default |
|---|---|---|---|
| to_symbol | Yes | ||
| from_symbol | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden, but it does disclose that the answer is rule-based inference with an associated confidence and evidence – useful context implying the edge may be heuristic rather than ground truth. It does not explicitly confirm this is a non-mutating read or describe any limits, so it goes beyond the name but stays thin.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence with no filler, and the subject ('Why citegraph believes') comes first. The fragmentary phrasing is dense rather than wasteful, though slightly awkward to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so the description need not enumerate return values, and the parameter direction is covered. However, with zero schema coverage for parameters and no annotations, the definition leaves an agent short on usage context and confirmation of read-only behavior.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% and the parameters only carry bare titles, so the description is the sole source of meaning. It conveys the directionality – from_symbol is the caller, to_symbol the callee – which the titles alone do not make unambiguous, but it adds no format or qualification details.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: it explains why the graph believes a call edge exists from one symbol to another, listing the supporting fields (rule, meaning, confidence, evidence). This clearly distinguishes it from list-style siblings like what_calls and get_symbol, though it never names an alternative explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no when-to-use guidance, no mention of when an agent should reach for explain_edge instead of what_calls or what_does_it_call, and no prerequisites. The diagnostic intent is only implied by the word 'Why'.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_config_keyA
Where config keys are defined (json/yaml/env example) and read in code. * is a wildcard.
A:B, A__B and a.b spellings match each other. Values are never returned.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | ||
| pattern | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so description carries full burden. It discloses important behavioral traits: wildcard support, equivalent spelling forms (A:B, A__B, a.b), and that values are never returned (safety for secrets). However, it doesn't mention permissions, rate limits, or result format details beyond 'values are never returned'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the scope, followed by key semantics. No wasted words. Could be slightly more structured (e.g., bullet points) but is efficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Output schema exists, so return values need not be explained. However, with no annotations and 0% schema coverage, the description should do more to cover behavioral aspects (e.g., what the results contain, ordering, pagination). It adds important semantics but leaves gaps for a search tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so description must compensate. It explains that 'pattern' supports wildcard `*` and that different spellings match each other, which adds crucial syntax/semantic meaning beyond the schema title 'Pattern'. It doesn't explain the 'limit' parameter, but the wildcard guidance is valuable. Baseline for 0% coverage would be low, but the description partially compensates for the main parameter.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (find) and resource (config key), and clarifies the scope: where keys are defined (json/yaml/env) and read in code. This is a clear, distinctive purpose. It doesn't reference sibling tools for differentiation, but 'config key' is unambiguous enough on its own.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No explicit when-to-use or when-not-to-use guidance. The description implies searching for config key definitions and reads, which is clear enough from the purpose, but it doesn't mention alternatives like get_symbol or search_symbols. Implied context only.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_pathB
Shortest call path from one symbol to another, following edges at or above min_confidence.
| Name | Required | Description | Default |
|---|---|---|---|
| max_depth | No | ||
| to_symbol | Yes | ||
| from_symbol | Yes | ||
| min_confidence | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It discloses shortest-path semantics and min_confidence filtering, but omits directionality, no-path behavior, and how max_depth truncates traversal.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single, front-loaded sentence with no filler. It efficiently states the core operation and the confidence-threshold condition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 4-parameter path tool with no annotations and 0% schema descriptions, the one-sentence description leaves max_depth and traversal details undocumented. An output schema exists, so return values need less coverage, but key input semantics remain incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%. The description clarifies from_symbol/to_symbol as endpoints and min_confidence as an edge threshold, but max_depth is never explained despite having a default of 6, so it only partially compensates.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific operation: shortest call path between two symbols, with an edge-confidence threshold. It implies a graph-traversal purpose distinct from direct-call siblings like what_calls, though it does not explicitly name alternatives.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance or alternatives are provided. The agent must infer that this is for connectivity/path queries rather than direct call inspection or symbol lookup.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_symbolA
Look up one symbol (bare, qualified, or repo:qualified). Ambiguous names return candidates.
| Name | Required | Description | Default |
|---|---|---|---|
| name | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses that ambiguous names return candidates rather than a single result, which is real behavioral information, but it says nothing about not-found behavior, errors, or whether the result is authoritative.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no filler, and the primary action is front-loaded ahead of the ambiguity caveat. Every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return structure need not be described, and the candidate-on-ambiguity note covers the main behavioral surprise. The remaining gap is the absence of any tie-breaker against the sibling search_symbols tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% for the single 'name' parameter, so the description must compensate, and it does: it enumerates the accepted name forms (bare, qualified, repo:qualified). It stops short of giving a concrete example or syntax rules beyond the labels.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (look up) and resource (one symbol), and scopes it to a single symbol so it is distinguishable from bulk retrieval. However, it never mentions the sibling search_symbols, which an agent would need to weigh against this tool.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use guidance and no mention of when not to use it. The reader must infer from 'one symbol' that this is for known names rather than discovery, and the obvious alternative (search_symbols) is never named.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
repo_overviewC
Languages, top modules, entry points and fan-in hotspots for one indexed repo.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full behavioral burden. It implies a read-only overview of an indexed repo, but never states safety, permissions, or freshness, and does not clarify prerequisites beyond 'indexed'.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded fragment listing the return contents, with no filler. Every term maps to a distinct output category.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained in the description. For a simple one-parameter read tool, the description covers purpose and scope, but omits usage guidance and read-only/authorization context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the single 'repo' parameter is undocumented except for its type. The description adds that the repo must be indexed, but gives no identifier syntax, expected format, or other meaning beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names the specific resource (repo overview) and enumerates its outputs (languages, top modules, entry points, fan-in hotspots). It distinguishes the high-level summary from siblings like get_symbol or what_calls, though it lacks an explicit verb.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use guidance, no alternatives named, and no exclusions. The content implies a summary tool but does not tell the agent when to prefer it over status or the search/navigation tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_symbolsB
Fuzzy search symbol names. kind: module|class|interface|function|method. Returns path:line evidence.
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | ||
| repo | No | ||
| limit | No | ||
| query | Yes |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations, so the description carries the burden. It usefully discloses the return format ('path:line evidence') and the valid kind vocabulary, but says nothing about limit behavior, repo scoping defaults, or match ranking. Adequate but thin for an annotation-free tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three terse fragments, purpose front-loaded, no filler. Slightly clipped at the cost of completeness, but every element is informative.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Return values are covered by the output schema, so the description needn't explain them. Still missing repo/limit semantics and disambiguation from sibling search/lookup tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, but the description compensates for the important enum-like kind parameter ('module|class|interface|function|method'). repo and limit are left to name inference with no guidance on defaults or scoping.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Fuzzy search symbol names') and hints at return format. An agent can distinguish it from a single-symbol lookup like get_symbol, though no sibling is named explicitly.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No when-to-use or when-not guidance and no alternatives named. The word 'fuzzy' implies exploratory use, but the agent must infer when to pick this over get_symbol or repo_overview.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
statusA
Index status: repos, indexed commit vs current HEAD (stale), counts, last run. Call this first.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full behavioral burden. It usefully discloses that it compares the indexed commit against current HEAD to detect staleness and reports last run, which is meaningful context. However, it never explicitly states that the operation is read-only or side-effect-free, leaving a gap for a tool with no annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the core reporting scope followed by the key usage directive. No filler or redundancy; every phrase contributes.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given zero parameters, full annotation absence, and an output schema that already describes return values, the description is complete enough. It summarizes what the status report contains and tells the agent to call it first, which are the key missing pieces not covered by structured fields.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there are no parameter semantics to document. The schema is fully covered and empty; per the rubric, zero-parameter tools baseline at 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it reports index status, listing specific fields: repos, indexed commit vs current HEAD (stale), counts, last run. It does not explicitly name or differentiate itself from any sibling tool, so it falls short of the 5 threshold for sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The instruction 'Call this first' gives a clear, explicit priority for when to use this tool. It does not name alternatives or when-not conditions, but the workflow guidance is strong enough to route the agent appropriately.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
what_callsA
Callers of a symbol (depth 1-3). Each edge has rule, confidence and the path:line of the call. Edges below min_confidence (default 0.5) are hidden and counted in notes; pass min_confidence=0.1 to see lower-confidence candidates such as name-only (ambiguous) matches.
| Name | Required | Description | Default |
|---|---|---|---|
| depth | No | ||
| limit | No | ||
| symbol | Yes | ||
| min_confidence | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full load and does well: it discloses what each edge contains (rule, confidence, path:line), that edges below the threshold are hidden and tallied in notes, and how to relax the filter. Missing only whether the traversal can be expensive or has any auth/limit caveats.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences with no filler; the core operation leads and the confidence-threshold behavior follows. Every clause contributes either return-shape or filtering information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not have been explained, yet the description usefully characterizes the edge payload anyway. The only real gap is the undocumented `limit` parameter for a traversal whose result count is bounded.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0% across 4 parameters, so the description must compensate. It meaningfully documents depth (1-3), min_confidence (0.5 default, 0.1 for ambiguous matches), but says nothing about `limit` (default 25) or what qualifies as a `symbol`, leaving two parameters undocumented in both schema and prose.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific resource and relationship — 'Callers of a symbol' — and scopes it with the depth range 1-3, immediately distinguishing it from the callee-side sibling what_does_it_call. An agent can identify the operation and its direction without opening any schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives actionable operational guidance: depth is bounded to 1-3 and min_confidence defaults to 0.5, with an explicit instruction to pass 0.1 to surface low-confidence name-only matches. It does not, however, name a sibling alternative or state when not to use this tool, leaving the routing decision to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
what_does_it_callA
Callees of a symbol (depth 1-3). Each edge has rule, confidence and the path:line of the call. Edges below min_confidence (default 0.5) are hidden and counted in notes; pass min_confidence=0.1 to see lower-confidence candidates such as name-only (ambiguous) matches.
| Name | Required | Description | Default |
|---|---|---|---|
| depth | No | ||
| limit | No | ||
| symbol | Yes | ||
| min_confidence | No |
Output Schema
| Name | Required | Description |
|---|---|---|
No output parameters | ||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden and does reasonably well: it discloses the edge payload (rule, confidence, path:line) and, importantly, the non-obvious filtering behavior that low-confidence edges are silently hidden and only counted in notes. It does not describe failure modes or the 'limit' truncation behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two dense sentences, front-loaded with the resource definition and followed by the output/filter contract. Nothing is wasted, though the second sentence crams default, effect, and tuning into one clause.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-shape explanation is not required, and the confidence-threshold semantics are covered. However, an unexplained 'limit' parameter and no mention of what happens at depth boundaries leaves gaps for a 4-parameter query tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate, and it partially does: it documents min_confidence's default (0.5), its suppression effect, and a recommended value (0.1) for ambiguous name-only matches, plus depth's 1-3 range. The 'limit' parameter is entirely unexplained in both schema and description.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific resource (callees of a symbol) and bounds the traversal (depth 1-3), which is a clear verb+resource statement. It does not explicitly name the inverse sibling 'what_calls', so an agent must infer the direction distinction from the name alone.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies the use case (finding what a symbol calls) but never states when to prefer this over 'what_calls' or 'explain_edge', nor any prerequisites. It does give threshold-tuning guidance for min_confidence, which is useful but is parameter advice rather than tool-routing advice.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.1.0- First observed
explain_edge - First observed
find_config_key - First observed
find_path - First observed
get_symbol - First observed
repo_overview - First observed
search_symbols - First observed
status - First observed
what_calls - First observed
what_does_it_call
TDQS
Scored across 9 tools
Each tool has a largely distinct purpose: get_symbol vs search_symbols (exact vs fuzzy) and what_calls vs what_does_it_call (inverse directions) are clearly separated by their descriptions. Minor overlap between status and repo_overview (both report index/repo state) is the only real friction, but descriptions disambiguate them.
Three conventions coexist: verb_noun (get_symbol, search_symbols, find_path, find_config_key, explain_edge), question-style (what_calls, what_does_it_call), and bare nouns (status, repo_overview). It remains readable but there is no single predictable pattern.
Nine tools is well-scoped for a call-graph/symbol-index server, with each tool covering a distinct query direction or diagnostic. No redundant or filler tools.
The core call-graph lifecycle is covered: status/lookup/search, bidirectional traversal, path finding, config tracing, repo overview, and edge explanation. Minor gaps exist (e.g. listing indexed repos explicitly or triggering a reindex), but agents can work around them.
Maintenance
Related MCP Connectors
Ask a codebase what calls what: search, blast radius, paths between symbols, and diffs.
Codebase graphs, caller impact analysis, and recorded project context for AI coding agents.
Search indexed code, trace dependencies, assess change impact, and recall repository memory.
Deterministic context layer for your codebase: change impact, blast radius, answers with receipts.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceEnables LLM agents to efficiently understand and navigate a codebase by providing semantic search over symbols and a reference graph, replacing expensive grep/glob calls with structured tools like definition lookup, caller/callee queries, and change-impact analysis.3MIT
- FlicenseNot gradedqualityAmaintenanceProvides efficient code navigation and graph-based analysis for AI agents, enabling symbol resolution, callers, implementations, and type schemas with minimal token usage.-
- AlicenseAqualityAmaintenanceEnables AI coding agents to query a semantic cross-repository code graph for symbols, references, callers, dependencies, and change impact across registered repositories.1221Apache 2.0
- FlicenseNot gradedqualityAmaintenanceEnables developer agents to perform semantic codebase search, dependency and impact analysis, cross-file refactoring, and full-stack API tracing through a unified query DSL over a high-performance graph engine.-