aspark-graph
This MCP server lets you build and query a deterministic local knowledge graph linking a repository's code to its aSPARK delivery artifacts, enabling story tracing and change impact analysis.
Build or rebuild the graph from a repo and its
.spark/artifacts, writing.aspark-graph/.Trace a user story through acceptance criteria, QA verdicts, plan tasks, and code links.
Compute the blast radius of file changes or a git diff range, tagging affected stories/ACs with confidence tiers.
Check gate health for a feature: orphan tasks, unverified acceptance criteria, and open findings.
Report whether the built graph is stale relative to the repo on disk.
Navigate the graph: look up nodes by id, find nodes by substring/type, get neighbors, and find shortest paths.
All query tools are read-only and return JSON;
build_graphis the only writing tool.
Integrates with local Git repositories to infer code-to-artifact links from commit history and to compute the blast radius of changes from a Git diff range (e.g., impact --diff HEAD~1..HEAD).
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@aspark-graphwhat code implements user story US-2?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
πΈοΈ aspark-graph
A lean, local knowledge graph that joins a repo's code to its delivery artifacts β so agents and humans can trace a user story to the code that implements it, and see the story-level blast radius of a change.
aspark-graph reads one repository β its source code and its
aSPARK .spark/ delivery trail (specs,
plans, reviews, QA reports) β and builds a single queryable graph, served over a
CLI and an MCP server. It is deterministic (tree-sitter + declared
artifact links; no LLM, no network) and disposable (the graph is a
rebuildable read model, never a source of truth).
βΆ Watch the 60-second explainer β real output only: the graph on screen is this repository's own graph at v0.7.1.
Try it in 30 seconds
No aSPARK project of your own? This repository has a real delivery trail, so it works as the demo. You need uv and a full clone:
git clone https://github.com/a-lottes/aSPARK-graph && cd aSPARK-graph
uvx aspark-graph build .
uvx aspark-graph query impact src/aspark_graph/queries.pyThat prints the user stories and acceptance criteria a change to queries.py touches, each
tagged declared, extracted or inferred. Use a full clone: inferred links come from
the git history, and a shallow clone (--depth 1) builds without them.
Related MCP server: kg-memory-mcp
The two questions it exists to answer
Everything else is in service of these:
Question | Tool | Plain meaning |
"Which code implements this user story, and did its acceptance criteria pass QA?" |
| Follow story β ACs β plan tasks β code β QA results, with zero grepping. |
"If I change these files, which stories and acceptance criteria are in the blast radius β what must QA re-verify?" |
| Walk code β tasks β stories/ACs, tagging each hit with how trustworthy the link is. |
Why this exists
When an AI agent (or the developer supervising it) works on an aSPARK-managed
repo too large to hold in your head, those two questions are exactly the ones
aSPARK's own review and QA gates depend on β and today the only way to answer
them is Grep/Glob plus reading .spark/ files by hand.
The spec β plan β review β QA trail is machine-parseable, and it is linked to code by intent. But nothing joins the two, so an agent re-derives the link every time it greps: slowly, incompletely, and non-reproducibly. aspark-graph computes that join once, deterministically, and lets you query it.
It does this by deliberately doing less on the code side than a general code graph, and adding the one thing general graphs don't have: the delivery artifacts. That artifact layer is what makes story tracing, gate-aware impact, and orphan detection possible at all.
When to use it β and when not
β Use it on an aSPARK repo big enough that "read every relevant file" isn't viable, when you need a fast, reproducible answer before opening files.
β Skip it on a repo small enough to hold in your head (just read the files), or a repo with no
.spark/artifacts (the artifact layer is the whole point; to see what it does, run the demo above on this repository).π€ Want a broad semantic code graph too? Run Graphify alongside it β different scope, no conflict. aspark-graph is an accelerant for aSPARK, not a replacement for a code-search tool.
Using aspark-graph in aSPARK gates? See
docs/aspark-integration.md
for drop-in CLAUDE.md blocks that wire the /peer-review and /demo-day
gates to the query tools.
Trust boundary, non-guarantees, and how to report a vulnerability: see
SECURITY.md
β read it before treating this server's output as anything other than data.
The graph model (read this to interpret any result)
The graph is a typed, directed multigraph. Every node id is stable and
deterministic, derived only from content and location, so two builds of an
unchanged repo produce byte-identical ids and a byte-identical graph.json.
Node types
Layer | Types | Source |
Code |
| tree-sitter extraction |
Artifact |
|
|
Edge types
Edge | Direction | Meaning |
| File β Class/Function | code structure |
| File β File | resolved import |
| Function β Function | best-effort, may be absent |
| Feature β Story | artifact structure |
| Story β AcceptanceCriterion | " |
| Feature β Task | " |
| Task β Story | plan links a task to the story it serves |
| Task β File/Function | the codeβstory bridge (best-effort; see below) |
| QACheck β AcceptanceCriterion | QA result for an AC |
| Finding β File | a review finding's location |
Confidence tiers β every artifact/code link carries a tier, and impact
reports the weakest link on the strongest path so you can trust a result
appropriately:
Tier | Rank | Where it comes from |
| strongest | an explicit |
| middle | tree-sitter ( |
| weakest | self-derived from git history β treat as a hint, confirm before acting |
Reading a result: an
impacthit taggedinferredreached the story only through a git-history guess; adeclaredhit rests on an author-written link. The tier never raises confidence β it reports the weakest step, so an inferred edge can only ever lower a path's trust, never mask a real one.
Node id schemes (useful when constructing get_node/shortest_path queries):
file:<relpath> e.g. file:src/aspark_graph/queries.py
def:<relpath>::<qualname> e.g. def:src/foo.py::Widget.render
feature:<name> e.g. feature:aspark-graph
story:<feature>:<id> e.g. story:aspark-graph:US-1
ac:<feature>:<id> e.g. ac:aspark-graph:AC-1.2
task:<feature>:<id> e.g. task:aspark-graph:T3
finding:<feature>:<id> e.g. finding:aspark-graph:F1
qa:<feature>:<ac>#<index> e.g. qa:aspark-graph:AC-1.1#0Install
Requires Python β₯ 3.11. aspark-graph is published on PyPI:
pip install aspark-graph
# or, with uv:
uvx aspark-graph build . # build the graph for the current repo, no install stepAdd it to Claude Code as an MCP server:
claude mcp add aspark-graph -- uvx aspark-graph serveBuilding from source (for contributors) is documented under Development below.
Update
pip install --upgrade aspark-graph
# or, with uvx, the latest published version always runs β no separate update stepThe graph is not forwards-compatible across versions: always rebuild after
updating (aspark-graph build .). Incremental builds (v0.4.0+) make this fast β
only changed files are re-parsed, so a routine update rebuild takes seconds on
most repos.
Build the graph
aspark-graph build [path] # scans code + .spark/, writes .aspark-graph/graph.jsonThe graph is written to .aspark-graph/graph.json at the repo root (gitignore
it β it's rebuildable). Re-running build on an unchanged repo produces a
byte-identical graph. Parsing fails loudly on .spark/ template drift
(it names the file and the mismatch) rather than silently guessing.
Query
Every query is available on both the CLI and MCP, and they return identical answers by construction (all query logic lives in one shared module; the CLI and server are thin adapters over it, and a parity test enforces it). Output is JSON.
CLI
# The two headline queries
aspark-graph query story_trace US-2 --feature my-feature
aspark-graph query impact src/foo.py src/bar.py
aspark-graph query impact --diff HEAD~1..HEAD # blast radius of a change range
# Gate & freshness
aspark-graph query gate_health my-feature # are this feature's ACs covered / passing?
aspark-graph query staleness # does the graph still match the repo on disk?
# Graph navigation
aspark-graph query get_node "file:src/foo.py"
aspark-graph query find_nodes Widget --type Class
aspark-graph query get_neighbors "story:my-feature:US-1" --edge-type has_ac
aspark-graph query shortest_path "task:my-feature:T1" "ac:my-feature:AC-1.1"MCP
The same operations are exposed as MCP tools: story_trace, impact,
gate_health, staleness, get_node, find_nodes, get_neighbors,
shortest_path β all eight read-only, plus build_graph, the one tool
that writes (<target>/.aspark-graph/graph.json and parse-cache.json).
Querying before a build (or any domain error) returns a clean
{"found": false, ...}-shaped result β never a raw traceback. See
SECURITY.md
for the full trust boundary.
Linking code to stories
impact and story_trace are only as useful as the implements (taskβcode)
links they can find. aspark-graph establishes those links three ways, strongest
to weakest confidence:
Confidence | Source | How to opt in |
| An explicit | In |
| Git commit history | Reference the task id and its story id in the commit message β subject |
| tree-sitter ( | Automatic β no action needed. |
Recommendation for aSPARK repos: make one commit per task whose message names
the task and story ids (the convention aSPARK's own workflow already encourages).
That alone lets impact answer on a repo that was never hand-annotated.
Inference is deterministic (it reads only committed state β file paths and
message ids, never timestamps) and offline; if git is unavailable it is
simply skipped. When multiple .spark/ features reuse the same T<n>/US-<n>
numbering, a commit is resolved to a single feature before linking (by the
.spark/<feature>/ tree it touched, or by a unique taskβstory pairing in its
message); a genuinely ambiguous commit contributes no edge β an honest
absence over a wrong cross-feature link.
Shallow clones and ZIP downloads
inferred links need commit history. A shallow clone (git clone --depth 1, the
default in many CI setups) has only part of it, and a ZIP download has none. When
the graph has at least one plan task, build says so on stderr:
Shallow git history: inferred links may be missing (run 'git fetch --unshallow', then rebuild).
No git history: inferred links are missing (build from a full git clone to get them).At most one of the two lines appears. MCP build_graph reports the same as two
booleans, shallow_history and no_git_history. In a shallow clone, git shows
the oldest kept (boundary) commits as adding every file, so the build skips their
inferred links: impact can show fewer links than on a full clone, never extra
ones. The exit code and the stdout format are unchanged, and on a full clone
graph.json is byte-identical to earlier versions. To get the missing links, run
git fetch --unshallow (or build from a full clone) and rebuild.
Supported languages
Code extraction covers Python, TypeScript/JavaScript, Java, Go and Rust (tree-sitter).
Files in other languages are recorded as unparsed File nodes β the build never
fails on an unknown language.
Artifacts read
Per feature folder in .spark/, the graph parses spec.md and plan.md, plus
review, QA and release files under either name aSPARK has used:
Artifact | Current name | Legacy name |
Review |
|
|
QA |
|
|
Release |
|
|
When both names exist in one folder, the current name wins. The build names the ignored
file (Ignored β¦ (qa.md takes precedence) on stderr, ignored_legacy_files over MCP), and
the ignored file leaves no trace in the graph. A QA result cell is read by its marker
first β β fail, β unverified, β
pass β and by its words only when it has none.
Design guarantees (why you can trust the output)
Deterministic. Byte-identical rebuild on an unchanged repo; parse-affecting dependencies are pinned exactly; a double-build test enforces it. The five tree-sitter grammars (Python, TypeScript, Java, Go, Rust) and tree-sitter core are pinned with
==, and the committeduv.lockis part of the determinism contract for the rest β a grammar version can change the node types it extracts, so an unpinned grammar would silently break byte-identity. The guarantee's boundary: byte-identical rebuild holds for an unchanged repo on a fixed grammar set. A grammar bump is a deliberate, changelog-documented, version-bumped event β never a silent change to what a past build promised.Offline & LLM-free. No network, no model calls β just AST parsing and artifact-template parsing.
Fails loudly, never silently. Template drift raises a named error; it never skips or guesses.
Clean errors. Domain errors (drift, graph-not-built) surface as one-line messages with a non-zero exit / a structured dict β never a stack trace.
Disposable. The graph is a read model. Delete
.aspark-graph/and rebuild; the source of truth is always the code and the.spark/files.
Out of scope
Languages beyond the five currently supported, an LLM/natural-language layer, precise call-graph resolution, a visualization UI, exports (Neo4j/GraphML/Obsidian), HTTP/team mode, and authenticated or remote MCP transport are out of scope. The current language support is Python, TypeScript/JavaScript, Java, Go, and Rust.
Development
Requires uv. Build and run from a checkout:
git clone https://github.com/a-lottes/aSPARK-graph.git aspark-graph
cd aspark-graph
uv sync --extra dev
uv run aspark-graph build . # build the graph for the current repo
uv run pytestAdd a checkout to Claude Code as an MCP server:
claude mcp add aspark-graph -- uv run --directory /path/to/aspark-graph aspark-graph serveTo pick up upstream changes: git pull && uv sync --extra dev.
The project dogfoods itself: its own .spark/aspark-graph/ trail is the
primary test fixture, so touching the parser or a query is checked against a real
aSPARK trail.
License
MIT Β© Andreas Lottes. Part of the aSPARK product family. Code-graph prior art: Graphify β different scope.
Available Tools
9 toolsbuild_graphA
(Re)scan a repository and its .spark/ artifacts and persist the graph. Writes /.aspark-graph/ β this is the one MCP tool that writes.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | . |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden and does disclose the key side effect: it writes to <path>/.aspark-graph/ and is the only write-capable MCP tool. It does not specify overwrite or idempotency behavior, but for a scan-and-persist tool the main behavioral risk is clearly flagged.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences with no filler. The action is front-loaded, and the second sentence delivers the crucial mutating side effect and sibling differentiation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with one optional parameter and no output schema, the description provides purpose, side effects, output location, and distinction from siblings. It could add explicit guidance about when to rebuild the graph, but the core operational context is sufficiently documented.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has no descriptions for the single `path` parameter, but the description compensates by using `path` in the output location pattern, indicating it is the repository root and determines where .aspark-graph/ is written. This gives the agent enough semantic meaning to use the parameter correctly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific verb ('(Re)scan'), a specific resource (the repository and .spark/ artifacts), and a concrete outcome (persist the graph). It also explicitly differentiates this tool from the read-oriented siblings by stating it is the only MCP tool that writes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly situates this tool as the write/persist operation among read-style siblings, so an agent can infer when it is needed. However, it does not explicitly name alternatives or state when not to use it, so it stops short of a full usage-guide.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_nodesC
Find nodes whose id or name contains a substring, optionally by type.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | . | |
| type | No | ||
| query | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description must disclose behavioral traits. It fails to mention case sensitivity, whether the substring match is exact or fuzzy, pagination, or potential performance impact. For a read-only search, it does not confirm immutability.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single concise sentence with no wasted words, but it omits essential details that would benefit the agent. It is not verbose but is under-specified.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema and sparse description, the tool lacks critical context such as return format, error handling, or result limits. For a search operation, this is insufficient for confident invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, and the tool description does not explain any parameters. The meaning of 'repo' (likely repository path) and 'type' (scope of nodes) is left ambiguous, forcing the agent to guess or assume defaults.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (find) and resource (nodes), and specifies matching on id or name substring with optional type filter. This distinguishes it from sibling tools like get_node (single node) or get_neighbors (adjacent nodes).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives such as get_node or search. The description only states what it does without context or exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
gate_healthC
The aSPARK gate invariants as data: orphan tasks, unverified acceptance criteria, and open findings for a feature.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | . | |
| feature | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description bears full responsibility for disclosing behavioral traits. It does not state if the tool is read-only, destructive, requires authentication, or has side effects. 'As data' implies read-only, but this is not explicit.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence, which is concise, but it front-loads jargon ('aSPARK gate invariants') without explanation. While not verbose, the structure could be improved by clarifying the tool's action earlier.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of output schema, annotations, and parameter explanations, the description is incomplete. It does not explain the return format, how to interpret the data (orphan tasks, etc.), or how the tool fits into the broader set of sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description must explain parameters, but it only mentions 'feature' generically. It does not clarify what 'feature' means, what 'repo' (default '.') refers to, or how they affect the output. The parameter semantics are almost entirely absent.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description vaguely indicates that the tool retrieves data about gate invariants (orphan tasks, unverified acceptance criteria, open findings) for a feature, but lacks a clear verb like 'get' or 'list', making its purpose ambiguous. Compared to sibling tools, it's not immediately obvious what specific action it performs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus alternatives like 'impact', 'staleness', or 'story_trace'. There is no mention of prerequisites, context, or scenarios where gate_health is appropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_neighborsC
Nodes within depth hops of a node (both directions); 'what touches this?'.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| repo | No | . | |
| depth | No | ||
| edge_types | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description is minimal and does not disclose important behavioral details such as performance characteristics, ordering of results, pagination, or treatment of edge types. Annotations are absent, so the description carries full burden but fails to provide sufficient transparency.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very concise (one line), which is efficient but comes at the cost of completeness. It front-loads the purpose but omits essential details, making it borderline under-specified.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, four parameters with zero descriptions, and sibling tools that suggest complex graph operations, the description is severely incomplete. It fails to explain the return format, the meaning of 'edge_types' and 'repo', or how depth works, leaving the agent with significant ambiguity.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description does not explain any of the four parameters beyond the mention of 'depth'. With 0% schema description coverage, the description should compensate, but it offers no meaning for 'id', 'repo', or 'edge_types'. The agent cannot infer parameter semantics from this description alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it returns nodes within 'depth' hops in both directions, using the intuitive phrase 'what touches this?'. This effectively conveys the core functionality and distinguishes it from siblings like 'shortest_path' (which finds paths) or 'build_graph' (which constructs full graph).
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives like 'shortest_path' or 'find_nodes'. There is no mention of when not to use it, prerequisites, or context where it might be inappropriate.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_nodeB
Look up a single node by its id (e.g. 'file:src/foo.py').
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | ||
| repo | No | . |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, description carries full burden. Only states 'look up a single node' without disclosing return format, side effects, authorization, or any constraints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence with no wasted words. Front-loaded with core purpose.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema and no annotations; description fails to explain return value structure or error handling. Too minimal for a tool likely returning a complex node object.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage 0%; description only mentions 'id' with an example but ignores the 'repo' parameter entirely. Does not explain parameter meanings beyond a format hint for id.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the verb 'look up' and resource 'single node by its id', with a concrete example of the id format. Differentiates from sibling tools like find_nodes by specifying singleton lookup.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives like find_nodes or get_neighbors. Does not mention prerequisites or contexts.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
impactA
Blast radius of a change: the stories and acceptance criteria that depend
on the given files (or the files in a git diff range), each tagged with its
weakest-edge confidence. Pass either files or diff, not both.
| Name | Required | Description | Default |
|---|---|---|---|
| diff | No | ||
| repo | No | . | |
| files | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It discloses the mutual exclusivity of input parameters and the output format (stories with confidence), but does not cover error handling, performance, or other behavioral traits.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences: first defines purpose and output, second gives usage constraint. No unnecessary words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no output schema or annotations, the description adequately defines input, output, and a key usage rule. It lacks details on error conditions or return structure confidence interpretation, but is likely sufficient for an AI agent.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, but the description adds meaning for 'files' (list of files) and 'diff' (git diff range) and the 'not both' constraint. However, the 'repo' parameter is not explained, which is a minor gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it computes the blast radius of a change, listing stories and acceptance criteria depending on files or a diff range, each tagged with confidence. It is specific and distinguishes from sibling graph tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description specifies the input constraint (pass either files or diff, not both) but does not explicitly compare to alternative tools like build_graph or find_nodes, leaving the agent to infer usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
shortest_pathC
An ordered path connecting two nodes, or an explicit 'no path' result.
| Name | Required | Description | Default |
|---|---|---|---|
| a | Yes | ||
| b | Yes | ||
| repo | No | . |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations provided. The description only describes the output format but does not disclose behavior like read-only nature, algorithm, or requirements (e.g., graph must already be built).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The one-sentence description is too short and vague, lacking substantive content. It does not balance brevity with informativeness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity of a shortest path tool, with no output schema and no annotations, the description is insufficient. It fails to explain return format, error cases, or integration with sibling tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 0% schema description coverage, the description should add meaning to parameters. It hints at 'two nodes' but does not explain a, b, or repo, leaving their semantics unclear.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description mentions 'ordered path' and 'no path' result, vaguely indicating graph path finding, but does not specify 'shortest' or the context (e.g., code repository graph). It fails to differentiate from sibling tools like get_neighbors or story_trace.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives such as build_graph or get_neighbors. No prerequisites or exclusions are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
stalenessC
Report whether the built graph still matches the repo on disk (US-4).
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | . |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, description carries full burden. It only states the function but does not disclose side effects, auth needs, or any behavioral traits beyond the name.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Description is very short but includes irrelevant 'US-4' in parentheses. It is efficient but could be cleaner without the extra noise.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given one optional parameter and no output schema, the description should at least hint at the return type (e.g., boolean). It does not, leaving the agent guessing about output format.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0% for the 'repo' parameter. Description adds no explanation about the parameter's meaning or usage, leaving the agent without clarity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description clearly states the tool reports whether the built graph matches the repo on disk. Verb 'report' and resource 'staleness' are specific, and it distinguishes from sibling tools like build_graph and find_nodes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool vs alternatives. No exclusions or context provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
story_traceB
Full thread of a user story: acceptance criteria (with their latest QA verdict), mapped plan tasks, and any best-effort code links.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | . | |
| story | Yes | ||
| feature | No |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations present, so description must convey behavior. It discloses the output structure (acceptance criteria, tasks, code links) but does not mention read-only nature, side effects, or performance characteristics. Adequate but not comprehensive.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Single sentence, front-loaded with the main action. Efficient but could benefit from breaking out the components for clarity.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Without an output schema, the description partially describes the return value (QA verdict, tasks, code links). However, it omits how parameters like 'repo' and 'feature' influence results, and whether multiple stories are returned. Adequate but has gaps.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, but the description fails to clarify the roles of 'repo', 'story', and 'feature'. It only mentions 'user story' implicitly. This leaves the agent without understanding how each parameter affects the result.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states it retrieves a 'full thread of a user story' including acceptance criteria with QA verdict, plan tasks, and code links. It provides a specific verb ('trace' implied) and resource, and is distinct from sibling tools like build_graph or find_nodes.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool or when to prefer alternatives. The description only lists output contents, not context or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.7.1- Added
build_graph - Added
get_node
2 tool updates
v0.4.0- Removed
build_graph - Removed
get_node
9 tool updates
v0.3.0- First observed
build_graph - First observed
find_nodes - First observed
gate_health - First observed
get_neighbors - First observed
get_node - First observed
impact - First observed
shortest_path - First observed
staleness - First observed
story_trace
TDQS
Scored across 9 tools
Most tools are clearly distinct: get_node is exact-id lookup while find_nodes is substring search, and get_neighbors/impact/shortest_path are separated by adjacency vs. dependency impact vs. pathfinding. There is mild conceptual overlap between story_trace and gate_health, and between impact and get_neighbors, but the descriptions provide enough boundary clarity.
All names use snake_case, and several follow verb_noun (build_graph, get_node, find_nodes, get_neighbors), but others are concept-noun phrases (story_trace, gate_health, impact, shortest_path, staleness). The mix is readable but not a consistent naming convention.
Nine tools is a well-scoped set for a repository graph server: one write/build operation, multiple query modes, traversal, impact analysis, health checks, and staleness reporting. Each tool contributes a distinct capability without obvious bloat.
The surface covers the core graph lifecycle well: build/rescan, node lookup, substring search, neighbor traversal, shortest path, story tracing, gate health, impact, and repo freshness. Minor gaps exist, like no explicit raw edge listing or graph-level metadata query, but agents can complete the intended workflows.
Maintenance
Related MCP Connectors
Repository knowledge graph MCP server for codebase understanding and debugging.
Knowledge coverage map and health score. Ingest docs into a governed knowledge graph via MCP.
Hosted code graph over MCP: exact callers, dependencies, and cross-repo blast radius for AI agents.
Codebase graphs, caller impact analysis, and recorded project context for AI coding agents.
Related MCP Servers
- AlicenseAqualityCmaintenanceA local knowledge graph MCP server that provides AI agents with permanent, structured memory about codebases, enabling semantic search, blast radius analysis, and convention enforcement.82MIT
- FlicenseNot gradedqualityDmaintenanceAn in-memory knowledge graph MCP server that gives coding agents structural and semantic recall over codebases by indexing Python source, ADR documents, and project configuration, exposing 7 tools for search, traversal, context retrieval, and natural-language Q&A.-
- AlicenseNot gradedqualityBmaintenanceExposes code graphs across multi-program repositories via MCP, enabling humans and agents to query the fleet with evidence.MIT
- AlicenseNot gradedqualityAmaintenanceLocal repository intelligence MCP server that builds a reusable graph of code structure for AI coding agents, providing 34 network-free tools for understanding, searching, and analyzing repositories without data leaving the machine.60 npmMIT