runecho
It is a read-only MCP oracle server for deterministic codebase facts, letting AI agents verify symbols and structural state before acting.
structure: inspect an enrolled repo's files and symbols, filtered by path globs, at different detail levels (tree, symbols, hashes, full).
diff: compare two snapshots, or a labeled snapshot against live code, to see structural drift.
hash: get a deterministic root hash and file count for the current code.
status: check per-repo health — last indexed, staleness, parse errors, coverage, snapshot count, latest hash, file cap.
health: get store-wide health — schema version, integrity check, enrolled repo count, and database path.
locate: find where a symbol is defined (name → file:line), searchable by exact name, prefix, or last dotted segment, with kind filters and pagination — useful to confirm a symbol exists before using it.
RunEcho
Your AI agent calls a function that doesn't exist. RunEcho stops the edit before it's written —
not after the build fails. It runs as a PreToolUse hook, checks every
Edit/Write against the symbols your code actually declares, and asks you
first when a reference has nothing behind it:
[runecho-guard] 1 symbol reference(s) not found in the indexed code — possible hallucination:
snippet line 2: validateSnapshotChecksum
Approve if these are legitimate (new/local/dynamic, or an intended removal).Deterministic. A parse and a lookup, not a model: the same edit gets the same verdict on every run, machine and agent. No LLM, no API keys, no network.
Free in context. A clean check writes nothing; only an edit it stops costs anything (~100 tokens). ~12 ms, no build, no language server.
Honest about its reach. It reads bare calls, constants and type references, which caught 4 of 9 real hallucinations in our benchmark; the other 5 need receiver types and are out of scope by design. One layer, like a type checker, not the whole answer.
How Well It Works
The same code produces the same answer. Every check is a parse and a lookup, so the verdict is identical on every run, every machine, and every agent — there is no model to sample from and nothing to re-roll. That is what makes it a gate rather than an opinion: it is model-free and vendor-neutral — no LLM, no API keys, no network, no build, no language server.
The guard costs zero context tokens — measured, not asserted: a clean check
writes nothing at all, and only an edit it actually stops costs anything (~100
tokens). It is a PreToolUse hook, so the agent never spends context deciding
whether to call it. The oracle MCP server is a separate surface and is not
free — its tool schemas cost ~968 tokens at session start, and structure
unscoped is expensive enough to be worth scoping. Every number, including the
unflattering ones, is in bench/TOKEN-COST.md.
How often it's wrong, measured against git history, not approvals. A
user's approve-anyway rate turned out to have no variance to measure — 308 of
308 ask-gated edits approved, zero denied, across 30 days of dogfood traffic.
So fpaudit judges each flagged symbol against dated git history instead:
was it defined at ask time, and is it defined now. Live reading across this
project's own dogfood corpus: 15.3% false-positive, 33.7% premature (the
guard was correct, just fired before the symbol it flagged existed — an
agent writing a caller before its callee), 51.0% stands (a real unbacked
reference). The full method, and what NOT to conclude from the 51%, is in
bench/FPAUDIT.md.
The scope, stated up front. RunEcho reads unqualified references — bare
calls, constant references, and type annotations. Measured against its own
corpus of real model hallucinations, that catches 4 of 9 (N=15 hand-verified
cases mined from live session transcripts, each backed by a compiler or runtime
error as independent ground truth). The other 5 are qualified positions —
df.groupby(…), tree.Root() — which need receiver-type resolution and are out
of scope by design. The numbers, including the misses, are in
bench/FINDINGS.md.
That is the honest shape of the thing: one cheap layer against AI coding mistakes — not the whole answer. A narrow, fast, certain check that runs before the write, not a system that makes your agent correct. Run it the way you run a type checker — one layer that removes one class of mistake completely, alongside the tests and review that catch the rest.
Related MCP server: codebase-rag
Why RunEcho Exists
Coding agents are useful, but they routinely make three kinds of mistakes:
they refer to functions or types that do not exist
they describe structural changes inaccurately
they keep reasoning from stale repo state after the code has moved on
RunEcho exists to give those agents a local source of truth they can query before they speak, edit, or commit.
Use it when you want:
a deterministic answer to "does this symbol actually exist?"
a structural diff instead of a vague summary of what changed
a guard that catches invented helper calls before they land in your repo
If your main problem is broad semantic search or general codebase exploration, RunEcho is not trying to be that. Its job is narrower: verify repo facts and reduce hallucinated code changes.
How It Works
RunEcho parses your source into a compact Intermediate Representation (IR): per file: its content hash plus the functions, classes, exports, and imports it declares. The IR has a deterministic root hash, so "did the structure change?" becomes a cheap hash comparison, and "what changed?" becomes a structural diff.
Snapshots of that IR are stored in a single central history database. Each enrolled repo has a stable identity, so the oracle can answer questions about any of your repos and compute drift between any two snapshots.
Three binaries make up the surface area:
runecho-ir— a CLI to enrol repos, index them, take snapshots, and inspect diffs and churn from the terminal.runecho-mcp— a stdio MCP server that exposes read-only oracle tools (structure,diff,hash,status,health,locate) to an AI agent.locateanswers "where is symbol X" deterministically (name → file:line), so an agent finds definitions without grepping or guessing.runecho-guard— a guard that checks new code against the indexed IR and flags references to symbols that don't exist (likely hallucinations). Runs as a git pre-commit hook, as a Claude CodePreToolUsehook that vets everyEdit/Write/MultiEditbefore it lands, or as--protocol— an edit on stdin, a versioned verdict document on stdout, for anything that is not Claude Code.
source ──▶ parser ──▶ IR (hashed) ──▶ snapshot ──▶ ~/.runecho/history.db
│
AI agent ──(MCP)──▶ runecho-mcp ──▶ structure / diff / hash / ...
│
git commit / agent edit ──▶ runecho-guard ──▶ "symbol X doesn't exist — block/ask"Prerequisites
Nothing to run a tagged release — the prebuilt binaries are self-contained (no runtime, no API keys).
Go 1.26+ only if you build from source (
bash install.sh).A POSIX or Windows shell. Storage lives under
~/.runecho/by default.No external services, no API keys.
Languages parsed today: Go, JavaScript, TypeScript, JSX, TSX, Google Apps
Script (.gs), Python, shell (.sh/.bash), Rust (.rs), and Ruby (.rb).
Extraction is intentionally shallow and deterministic: top-level structure, not
full semantic analysis.
Quick Start
Get the binaries. Either download a prebuilt release (no Go needed) — pick your OS/arch from the latest release:
# example: macOS arm64 — adjust the asset name for your platform. # NOTE: the tag in the URL path is v-prefixed; the asset filename is not. TAG=v0.17.1; NUM=0.17.1 curl -sSL "https://github.com/inth3shadows/runecho/releases/download/${TAG}/runecho_${NUM}_darwin_arm64.tar.gz" | tar -xz install -m755 runecho-ir runecho-mcp runecho-guard ~/.local/bin/…or build from source (needs Go 1.26+), which also installs the guard hooks:
bash install.sh runecho-ir install --periodic # optional, run inside the checkout: hourly reindex that also keeps the binaries at the newest releaseEnrol a repo and capture its current structure:
runecho-ir repo add /path/to/your/repo runecho-ir repo reindex <name> # name is shown by `repo add`If the directory you want to enrol is not the directory you want parsed, set a separate source root:
runecho-ir repo add /path/to/worktree --source-root=/path/to/sourceSee what's enrolled and ask for drift since the last snapshot:
runecho-ir repo list runecho-ir diff --since=reindex /path/to/your/repoRegister the oracle with your AI agent so it can query directly:
claude mcp add runecho -- ~/.local/bin/runecho-mcpFor Codex, add this to
~/.codex/config.toml:[mcp_servers.runecho] command = "/home/YOUR_USER/.local/bin/runecho-mcp" # absolute path; TOML does not expand ~Install the edit-time guard in Claude Code — the primary integration if you want RunEcho to vet assistant edits before they are written:
/plugin marketplace add inth3shadows/runecho /plugin install runecho-guard@runechoThe plugin wires both hooks —
PreToolUse(the guard) andPostToolUse(records the outcome and refreshes the index). It does not ship the binary, so step 1 still has to have happened. If the binary is missing the hook defers silently rather than erroring on every edit. Uninstall with/plugin uninstall runecho-guard@runecho.Without plugin support, print the equivalent
~/.claude/settings.jsonsnippet and merge it by hand:bash install.sh --print-hook-config(Optional) Install the commit-time guard in a repo you've enrolled:
bash install.sh --hook # run from the target repo's rootIt blocks commits that call functions which exist nowhere in the indexed code (with a "did you mean …?" suggestion when there's a close match). Bypass any single commit with
RUNECHO_GUARD_SKIP=1 git commit ….(Maintainers/forks only) If you cut release tags from this repo, install the tag-monotonicity safety net:
bash install.sh --hook-pre-pushRejects a
vX.Y.Ztag push that isn't semver-greater than the highest existing tag — see issue #51.
Current Boundaries
RunEcho is strongest when you want deterministic structure and guardrails, not general-purpose code intelligence.
It tracks top-level symbols and imports/exports, not full type information.
Parsers are AST-based but intentionally shallow — they extract definitions (functions, classes, methods), not semantics: no type inference, call graph, or cross-file binding. Go uses the stdlib
go/ast; Python, JS/TS, Rust, and Ruby use a pure-Go tree-sitter runtime; shell uses a masking scan. Imports and exports for the tree-sitter languages are still regex. Each language has known gaps — see the Parser Capability Matrix for the per-language honest accounting.Indexing covers more languages than the guard checks. Shell, Rust, and Ruby feed the index (
structure,locate,diff) but are not validated at edit time — the guard's reference checks exist for Go, JS/TS, and Python only.The guard validates unqualified references: bare calls (
foo(...)), bare type annotations (x: SomeType), and SCREAMING_SNAKE constant references. It does not flag qualified references (obj.method(...),pkg.Thing,x.attr) — those would need receiver-type resolution, semantic analysis RunEcho deliberately avoids. This boundary is measured, not asserted: seebench/FINDINGS.md, where a corpus of real transcript-observed hallucinations places the guard's catch-rate by reference position (the qualified positions are the deliberate gap).Snapshots, diffs, and hash queries are local and deterministic. There is no semantic search, embedding index, or hosted control plane here.
The guard runs unattended on every commit/edit with no sandboxing — see SECURITY.md for the threat model, what's stored, and how to report a vulnerability.
Project Structure
Path | Purpose |
| The CLI: snapshot, diff, map, log, churn, verify, truth-trail, validate-claims, contract, guard-stats, fpreport, fpaudit, repo, backup, install — plus indexing, which is the no-subcommand default ( |
| The stdio MCP oracle server |
| The guard: pre-commit mode, Claude Code hook mode, and |
| Per-language structure extraction (Go/JS/TS/JSX/TSX/.gs/Python/shell/Rust/Ruby) |
| IR build, deterministic hashing, JSON storage |
| Central store: migrations, registry, diff, churn, contracts, backup |
| Minimal MCP plumbing + the oracle tools |
| Diff parsing, symbol extraction, validation, did-you-mean |
| Edit-scope contract format and parsing |
| Memoized export sets for Go dependencies (qualified-call checks) |
|
|
| Symbol-reference extraction from prose ( |
| Canonical git-common-dir resolution (worktree identity) |
| Builds all three binaries; |
Related Documentation
Technical Reference — architecture, storage schema, the IR, the MCP tools, maintenance
Usage Guide — day-to-day operations: enrolling repos, integrations, reading drift, troubleshooting
Token cost — measured context cost of every surface, including where RunEcho is expensive
False-positive audit — the guard's fp/premature/stands rate against git history, and why approval rate isn't a false-positive proxy
Changelog — notable changes per release; versioning policy
License
MIT — see LICENSE.
Available Tools
6 toolsdiffARead-onlyIdempotent
Structural drift for an enrolled repo: files and symbols added, removed, or whose body hash changed. With a+b (snapshot ids) diffs those snapshots; with since=label diffs that snapshot vs live code; default diffs the latest snapshot vs live code. Errors if the repo is not enrolled.
| Name | Required | Description | Default |
|---|---|---|---|
| a | No | snapshot id A (with b) | |
| b | No | snapshot id B (with a) | |
| repo | Yes | name of an enrolled repo | |
| since | No | diff latest snapshot with this label vs live code | |
| session | No | with `since`: pin the reference snapshot to this session id |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and non-open-world, so the safety profile is covered. The description adds real context beyond that: the failure condition (errors if the repo is not enrolled) and a definition of what counts as drift.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences, front-loaded with the output definition and then the three invocation modes. No filler or redundant restatement of the tool name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description compensates by describing the returned content (added/removed/body-hash-changed files and symbols) and the error case. It stops short of describing output shape or ordering, but is adequate for a 5-param read tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so baseline is 3, but the description goes further by explaining the pairwise relationship of a+b, the standalone meaning of since, the default fallback with no parameters, and the session pinning modifier. That interaction logic is not fully captured by the per-property schema text.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific computation (structural drift) and enumerates the exact output domain: files and symbols added, removed, or whose body hash changed. An agent can tell this apart from hash/structure siblings, though the description does not explicitly name when those siblings are preferable.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Spells out three distinct calling modes: a+b diffs two snapshots, since=label diffs a labeled snapshot vs live code, and the default diffs latest snapshot vs live code. This is genuine mode-selection guidance, but it does not address when to reach for this tool over a sibling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
hashARead-onlyIdempotent
Deterministic root hash + file count of an enrolled repo's current code. Same code → identical hash across machines. Use to confirm two machines/checkouts hold identical code, or that code changed since a stored hash (compare with status's latest stored hash). Errors if the repo is not enrolled.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | Yes | name of an enrolled repo |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover the safety profile (readOnly, idempotent, closed-world), so the description's job is to add beyond that. It does: determinism guarantees identical hashes across machines for identical code, and it discloses the failure mode (errors when the repo is not enrolled). It does not cover latency or hashing scope limits, but the added context is substantial.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with what is returned before the usage guidance and the error condition. No sentence is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
There is no output schema, so the description carries the return-value burden and does so ('root hash + file count'), plus the determinism contract and the not-enrolled error. For a single-param, read-only tool, nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Only one parameter, fully documented in the schema (100% coverage), so the schema does the heavy lifting. The description only implicitly clarifies that `repo` must be enrolled ('an enrolled repo', 'Errors if the repo is not enrolled'), adding marginal value over the schema's 'name of an enrolled repo'.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: computes a deterministic root hash plus file count for an enrolled repo's current code. It also names the sibling it relates to (`status`'s stored hash), so an agent can distinguish it from `diff`, `structure`, `health`, or `locate`.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use conditions: to confirm two machines/checkouts hold identical code, or that code changed since a stored hash. It also names the alternative source of comparison (`status`'s latest stored hash), which is exactly the routing information an agent needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
healthARead-onlyIdempotent
Store-wide health: schema version, integrity check, enrolled repo names and count, db path. Use once to check the store itself is sound (integrity, schema) or to list what is enrolled; for one repo's freshness use status. Takes no arguments.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, and openWorldHint=false, so the safety profile is covered. The description adds real value by enumerating the returned fields (schema version, integrity, enrolled repos, db path) even though there is no output schema, and confirms it takes no arguments.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, front-loaded with the resource scope, then the usage routing. Every clause carries information; no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a parameterless read-only diagnostic tool with no output schema, the description is complete: it says what it checks, what it returns, when to use it, and which sibling to use instead.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so there is nothing for the description to disambiguate beyond confirming 'Takes no arguments,' which matches the empty schema. Baseline 4 applies for a no-parameter tool.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Store-wide health') and enumerates exactly what it reports: schema version, integrity check, enrolled repo names and count, db path. It also explicitly distinguishes itself from the sibling `status`.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ('check the store itself is sound... or to list what is enrolled') and names the alternative with its condition ('for one repo's freshness use `status`'). Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
locateARead-onlyIdempotent
Deterministically locate symbols in an enrolled repo: name → file:line (+ short body hash). Pass symbol to find a specific definition without grepping (a named lookup searches every kind); omit it to list all (capped, paginate with offset) — the unfiltered list defaults to functions+classes. Use this to verify a symbol exists before claiming it does: a zero-match result is definitive (parsed from the live AST), unlike grep, which can miss real symbols (formatting/whitespace variance, multi-line signatures) or hit false positives (comments, strings).
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | restrict to func|class|export|import (default: func+class) | |
| repo | Yes | name of an enrolled repo | |
| offset | No | skip this many matches before returning a page (default 0). Page again with the response's next_offset until it's absent. | |
| symbol | No | symbol to locate: matches by exact name, name prefix, or last dotted segment (e.g. "fetch" finds "Reader.fetch"). Omit to list all (capped). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and closed-world, so safety is covered. The description still adds real behavioral context beyond them: determinism from the live AST, that a zero-match result is definitive, that unfiltered results are capped and paginated, and that the default kind set is func+class. It stops short of detailing result size or latency, but that is a minor gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded: the core operation (name → file:line) comes first, then parameter behavior, then the verification rationale. It is dense but nearly every clause carries information; the grep contrast runs slightly long, keeping it just under a 5.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the return-value burden and does so: it specifies the file:line plus short body hash output and the next_offset pagination contract. For a read-only, four-parameter lookup tool, nothing needed to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds meaning beyond the schema: a named lookup ignores the kind narrowing and searches every kind, and the unfiltered list defaults to functions+classes. It also reinforces the capped/paginated behavior of `offset`, which is genuinely useful for interpreting results.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a precise verb+resource: deterministically locate symbols in an enrolled repo, returning name → file:line plus a body hash. It distinguishes itself from grep explicitly and from siblings like hash/structure/diff by emphasizing deterministic AST-parsed lookup. An agent can identify exactly what this tool produces without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit conditional usage: pass `symbol` for a named lookup that searches every kind, omit it to list all with pagination. It also states the purpose-level trigger (verify a symbol exists before claiming it does) and contrasts with the grep alternative, which is exactly the routing guidance an agent needs.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
statusARead-onlyIdempotent
Per-repo index health: last indexed, staleness, parse errors, coverage %, snapshot count, latest stored hash, file cap. Use to judge the stored index that diff's default baseline and the guard read from (a stale index means a stale baseline), and to check coverage before trusting a structure/locate miss: those parse live code on every call, so staleness does not affect them, but the file cap does, and coverage below 100% means files they never see. For the store as a whole use health. Errors if the repo is not enrolled.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | Yes | name of an enrolled repo |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only declare readOnly/idempotent/openWorld=false, so the burden is on the text, and it delivers: it discloses the error condition (repo not enrolled), explains that coverage below 100% means files other tools never see, and clarifies that the file cap affects live-parse tools while staleness does not. That is real behavioral context beyond the annotations rather than a restatement of them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the payload of returned fields, then the decision guidance. It is information-dense and every clause is load-bearing, though the middle sentence is long enough that a reader must parse it carefully rather than skimming.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description enumerates the returned fields so the agent knows what to expect, and it covers the error case plus cross-tool implications for diff, structure, locate and health. Nothing needed to call this single-parameter, read-only tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single parameter with 100% schema description coverage, so the schema already explains that `repo` names an enrolled repo. The description reinforces this indirectly via the 'not enrolled' error but adds no format or naming detail beyond the schema, making the baseline 3 appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource combination (per-repo index health) and enumerates exactly what it reports: last indexed, staleness, parse errors, coverage %, snapshot count, hash, file cap. It explicitly distinguishes the repo-scoped tool from its sibling `health`, which covers the store as a whole, so an agent can pick the right one without opening a schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use guidance tied to concrete decisions: judge the stored index that `diff`'s default baseline and the guard read from, check coverage before trusting a `structure`/`locate` miss. It also states when NOT to worry (staleness doesn't affect live parse tools, but the file cap does) and names the alternative (`health`) for store-wide questions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
structureARead-onlyIdempotent
Deterministic structure (files + symbols) of an enrolled repo's current code. Use to ground claims about what functions/types/exports exist. Scope with paths globs and pick a detail level to keep responses small.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | Yes | name of an enrolled repo (see `health`/registry) | |
| paths | No | optional glob filters; return only matching files (e.g. "internal/mcp/**" or "*.go"). `**` matches across directories. Omit for the whole repo. | |
| detail | No | tree = file paths + symbol counts only (cheapest); symbols (default) = per-file symbols[] (name/kind/line, plus `doc` — the verbatim first line of the symbol's doc comment where it has one, absent otherwise; it is the one field NOT verified against the code, so treat it as what the author wrote, not as checked intent) + refs; hashes = symbols plus each symbol's content hash (~2.5x the tokens; only needed to detect body-level drift, which `hash`/`diff`/`status` answer far more cheaply); full = also the legacy imports/functions/classes/exports arrays + symbol_hashes (redundant with symbols[], for back-compat) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (read-only, idempotent, closed-world), so the bar is lower. The description adds real value beyond them: determinism, the ~2.5x token cost of `hashes`, and the important caveat that the `doc` field is the one field NOT verified against the code.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the purpose and followed by usage and scoping. Dense but every clause carries information; no filler or restated title.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description sketches the return shape (files + per-file symbols with name/kind/line, refs, hashes) and the cost implications of each level. That is nearly enough for correct invocation, though exact response nesting is left implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description reinforces the cost/verbosity tradeoff across `detail` levels and the glob semantics of `paths` ('`**` matches across directories'), which is meaning an agent can act on when choosing arguments.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource: a deterministic files+symbols structure of an enrolled repo's current code, which an agent can distinguish from mutation or diff tools. It does not explicitly name siblings (hash/diff/status) as things this is *not*, but the scope is concrete and unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives a clear when-to-use: 'ground claims about what functions/types/exports exist', plus scoping advice (paths globs) and guidance on selecting a detail level. It even routes a sub-case away ('body-level drift, which hash/diff/status answer far more cheaply'), but stops short of stating when not to use this tool at all.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v0.1.0- First observed
diff - First observed
hash - First observed
health - First observed
locate - First observed
status - First observed
structure
TDQS
Scored across 6 tools
Most tools are cleanly separated and the descriptions explicitly cross-reference each other (status vs health, diff vs structure). The only real overlap is between structure and locate, both of which surface files/symbols, but the descriptions differentiate full-structure listing from targeted symbol lookup well enough.
All six names are single lowercase words with a uniform style, which is easy to predict. The only slight inconsistency is that some are nouns (structure, hash, status, health) and others are verbs (diff, locate), so the semantic pattern shifts even though the form is consistent.
Six focused tools is well-scoped for a deterministic code-structure/index server, and none appears redundant. Each tool earns its place in the read/inspection workflow.
The surface is read-only and coherent, but every tool repeatedly says it 'errors if the repo is not enrolled,' yet there is no enroll tool, and diff references snapshot ids/labels with no snapshot-creation tool. These are notable missing lifecycle operations that create dead ends for a fresh agent.
Maintenance
Related MCP Connectors
Deterministic context layer for your codebase: change impact, blast radius, answers with receipts.
Code intelligence for coding agents: semantic, AST, graph, and full-text search. 279+ languages.
Search indexed code, trace dependencies, assess change impact, and recall repository memory.
Ask a codebase what calls what: search, blast radius, paths between symbols, and diffs.
Related MCP Servers
- FlicenseAqualityAmaintenanceDeterministic code intelligence engine — indexes 27 languages into a queryable symbol graph for real-time blast-radius analysis, no embeddings or LLM calls.526-
- AlicenseNot gradedqualityCmaintenanceEnables LLM agents to efficiently understand and navigate a codebase by providing semantic search over symbols and a reference graph, replacing expensive grep/glob calls with structured tools like definition lookup, caller/callee queries, and change-impact analysis.3MIT
- AlicenseNot gradedqualityAmaintenanceExposes a codebase's symbol graph and symbol-aware editing tools to AI agents, enabling targeted symbol lookup, impact analysis, and atomic multi-file edits with reduced context tokens.7 npmMIT
- AlicenseNot gradedqualityAmaintenanceIndexes a codebase into a symbol-level graph and exposes tools for finding symbols, querying relationships, and assessing impact, letting AI coding agents answer structural questions in a single call within a token budget.70 npm4Apache 2.0