trace-mcp
trace-mcp is an MCP server that indexes a codebase (and markdown vaults) into a framework-aware dependency graph and serves it to AI agents so they can answer code questions with far fewer tokens than raw file reading.
Project orientation:
get_project_map,suggest_queries,get_index_health,get_coverage_report— see frameworks, languages, file counts, and index status.Symbol & text search:
search(symbols by name/kind/FQN, fuzzy/semantic/fusion modes),search_text(regex/full-text across files),suggest_queriesfor onboarding.Code reading without Read:
get_symbol(one symbol's source),get_outline(file's symbols/signatures),get_context_bundle(symbol + imports within a token budget),get_feature_context(ranked snippets by topic).Graph navigation:
find_usages(references),get_call_graph(bidirectional callers/callees),get_change_impact(blast radius, risk score, breaking changes, affected tests),get_task_context/plan_turn(optimal starting context for a dev task).Framework-aware edges: routes→controllers→views, ORM relations, migrations→schema, DI, events, cross-language component links (88 framework integrations).
Decision memory:
remember_decision,mine_sessions,query_decisions,invalidate_decision— capture, mine, search, and link architectural decisions to code symbols.Session analytics & savings:
get_session_analytics,get_optimization_report,get_real_savings,get_usage_trends,get_session_stats— measure token waste and trace-mcp savings.Diagnostics & maintenance:
get_diagnostics(run tsc/mypy/pyright mapped to symbols),register_edit(incremental reindex after edits),batch(run up to 10 tools in one request),load_tools/get_preset_info(manage deferred tools).
Provides search capabilities across past session content related to GraphQL discussions, with verbatim conversation fragments and file references.
Provides framework-aware understanding of connections between Laravel controllers and Vue components, enabling tracing of request flows from PHP to rendered Vue pages with prop mapping.
Supports Laravel framework features including routes, controllers, Eloquent models, migrations, and Inertia integration for cross-language dependency analysis.
Utilizes bundled ONNX embeddings for semantic code search capabilities that work offline without requiring API keys.
Includes security scanning capabilities for OWASP Top-10 vulnerabilities and taint analysis as part of the security scanning features.
Supports decision memory linking to PostgreSQL usage decisions (e.g., JSONB support) and schema reconstruction from migrations for dependency analysis.
npm install -g trace-mcp # MCP server, no app
trace init # wire it into your agent, once per machine
trace add # index the repo you are in72.7% fewer input tokens to review a pull request — median over 60 merged PRs in six repos that are not ours, 13,595 → 3,291 per pull request. Method and reproduction →
Measured at trace-mcp 3.23.2 (cb8ab30c) on 7 September 2026 — a result from that build, not a claim about the current one. What it set out to measure, the bar it had to clear and the verdict: preregistration.
Cheaper is not the same as better, so the same 60 pull requests were reviewed twice and scored blind. The trace-mcp arm understood the change in 67% of them against 65% for naive file loading, at 0.80 false positives per PR against 0.58. Quality half of the benchmark →
The problem
AI agents pay repeatedly for work they have already done. Every turn, the agent re-reads the same files, re-traverses the same dependencies, and re-inflates the context window with structure it discovered five steps ago. That repeated work is most of what a long session costs in tokens and latency.
trace-mcp builds a framework-aware graph of your codebase once, then serves it through MCP so the agent reasons from a precomputed structure instead of brute-reading the repo. Ask "what breaks if I change this model?" — instead of 80 Grep calls and 190 file reads, the agent calls get_change_impact once and gets the blast radius across PHP, Vue, migrations, and DI. 88 framework integrations across 81 languages, 182 tools.
The binding constraint is recomputation, not model capability: token bills, latency, and hallucinations all grow with project size instead of with task complexity. trace-mcp closes the recomputation leak. The graph is built once, kept incrementally fresh, and served to every agent that asks — so the same work isn't paid for over and over.
Lower cost — fewer tokens per successful answer, on average and at peak
Lower latency — fewer sequential tool calls, fewer round-trips to the model
Higher accuracy — less noise in context means fewer hallucinations and stronger first-response correctness
Production stability — context growth tracks task complexity rather than repository size
We started with code intelligence, where the repetition is most expensive, and the same engine now indexes markdown knowledge vaults (Obsidian, Logseq, plain MD) as a peer domain. Wikilinks, tags, frontmatter, and embeds become graph edges and symbol metadata; search, find_usages, get_change_impact, and apply_rename work identically over both.
Related MCP server: LiLBrain
What trace-mcp does for you
You ask | trace-mcp answers | How |
"What breaks if I change this model?" | Blast radius across languages + risk score + linked architectural decisions |
|
"Why was auth implemented this way?" | The actual decision record with reasoning and tradeoffs |
|
"I'm starting a new task" | Optimal code subgraph + relevant past decisions + dead-end warnings |
|
"What did we discuss about GraphQL last month?" | Verbatim conversation fragments with file references |
|
"Show me the request flow from URL to rendered page" | Route → Middleware → Controller → Service → View with prop mapping |
|
"Find all untested code in this module" | Symbols classified as "unreached" or "imported but never called in tests" |
|
"What's the impact of this API change on other services?" | Cross-subproject client calls with confidence scores |
|
"What notes link to this concept?" | Backlinks across the vault, with section + alias context |
|
"What breaks if I rename this note?" | Every |
|
Four capabilities that are rare among adjacent tools:
Framework-aware edges — trace-mcp understands that
Inertia::render('Users/Show')connects PHP to Vue, that@Injectable()creates a DI dependency, that$user->posts()means apoststable from migrations. 88 framework integrations.Code-linked decision memory — when you record "chose PostgreSQL for JSONB support", it's linked to
src/db/connection.ts::Pool#class. When someone runsget_change_impacton that symbol, they see the decision. MemPalace stores decisions as text; trace-mcp ties them to the dependency graph.Cross-session intelligence — past sessions are mined for decisions and indexed for search. When you start a new session,
get_wake_upgives you orientation in ~300 tokens;plan_turnshows relevant past decisions for your task;get_wake_up { scope: "resume" }carries over structural context from previous sessions.Code and knowledge in one graph — point trace-mcp at a markdown vault (Obsidian, Logseq, plain MD) and the same engine indexes it: each note becomes a
note:<basename>symbol, headings become nested sections,[[wikilinks]]and![[embeds]]become graph edges, frontmatter and#tagsride on metadata. PageRank, Signal Fusion ranking, embeddings, and rename refactoring all apply unchanged. The agent does not learn a second tool: it is the same graph, holding both the codebase and the notes.
Why agents keep re-reading
AI coding agents recompute the same work every turn — and they're framework-blind while doing it.
They re-read UserController.php, then re-read it again next turn. They don't know that Inertia::render('Users/Show', $data) connects a Laravel controller to resources/js/Pages/Users/Show.vue. They don't know that $user->posts() means the posts table defined three migrations ago. They can't trace a request from URL to rendered pixel — so they trace it again, and again, every session.
The result: 5–15× repeated reads of hot files in a single task, context windows used as scratch databases, and agents that get more expensive the bigger the project gets — instead of more capable.
The solution
trace-mcp builds a cross-language dependency graph from your source code and exposes it through the Model Context Protocol — the plugin format Claude Code, Cursor, Windsurf and other AI coding agents speak. Any MCP-compatible agent gets framework-level understanding out of the box.
Without trace-mcp | With trace-mcp |
Agent reads 15 files to understand a feature |
|
Agent doesn't know which Vue page a controller renders |
|
"What breaks if I change this model?" — agent guesses |
|
Schema? Agent needs a running database | Migrations parsed — schema reconstructed from code |
Prop mismatch between PHP and Vue? Discovered in production | Detected at index time — PHP data vs. |
Desktop app
trace-mcp ships with an optional Electron desktop app (packages/app) that gives you a visual surface over the same index the MCP server uses. It manages multiple projects, wires up MCP clients, and provides a GPU-accelerated graph explorer — all without opening a terminal.
Projects & clients. The menu window lists indexed projects with live status (Ready / indexing / error) and re-index / remove controls. The MCP Clients tab detects installed clients (Claude Code, Claw Code, Claude Desktop, Cursor, Windsurf, Continue, Junie, JetBrains AI, Codex, AMP, Warp, Factory Droid) and wires trace-mcp into them with one click, including enforcement level (Base / Standard / Max — CLAUDE.md only, + hooks, + tweakcc & agent-behavior rules; Max-tier features are Claude Code–specific). Warp and JetBrains AI require manual paste in the IDE because their config storage is GUI-only.
Per-project overview. Each project opens in its own tabbed window: Overview (files, symbols, edges, coverage, linked services, re-index), Ask (natural-language query over the index), and Graph. Overview also surfaces Most Symbols files, last-indexed timestamp, and the dependency coverage meter.
GPU graph explorer. The Graph tab renders the full dependency graph on the GPU via cosmos.gl — tens of thousands of nodes/edges at interactive frame rates. Filter by Files / Symbols, overlay detected communities, highlight groups, toggle labels/FPS, and step through graph depth. Good for getting a feel for coupling, hotspots, and how a codebase is actually shaped before you dive into tools.
Install on macOS: Download the .dmg — open it and drag trace-mcp to Applications. The button on the site picks Apple Silicon or Intel for you; if you would rather choose yourself, both builds are on the Releases page. The app is signed with a Developer ID and notarized by Apple, so it opens without a warning — if macOS ever does warn you about a trace-mcp build, that warning is real and the download should not be trusted.
Install on Windows: run trace-mcp.Setup.<version>.exe from Releases.
In-app updater stuck on an old build? App versions 3.10.0 and earlier on macOS/Windows can't update themselves — "Check for updates…" shows Cannot set properties of undefined (setting 'autoDownload') and does nothing, a bug fixed in 3.11.0 that the affected builds can't fetch their own way out of. Reinstall by hand: download the .dmg (macOS) or grab the latest trace-mcp.Setup.<version>.exe from Releases (Windows) — or, if you have the CLI installed, run trace-mcp install-app.
The app talks to the same trace-mcp daemon (http://127.0.0.1:3741) that MCP clients use, so anything you index from the app is immediately available to Claude Code / Cursor / etc. If you only want the MCP server and the CLI, you do not need the app at all — npm install -g trace-mcp is the whole install.
How trace-mcp compares
trace-mcp combines code graph navigation, cross-session memory, and real-time code understanding in a single tool. Most adjacent projects solve one of these — trace-mcp unifies all three and is the only one with framework-aware cross-language edges (88 framework integrations) and code-linked decision memory.
vs. token-efficient exploration (Repomix, jCodeMunch, cymbal) — trace-mcp adds framework edges, refactoring, security, and subprojects on top of symbol lookup.
vs. session-memory tools (MemPalace, claude-mem, ConPort) — trace-mcp links decisions to specific symbols/files, so they surface automatically in impact analysis.
vs. RAG / doc-gen (DeepContext, smart-coding-mcp) — trace-mcp answers "show me the execution path, deps, and tests," not "find code similar to this query."
vs. code-graph MCP servers (Serena, Roam-Code) — trace-mcp has the broadest language coverage (81 languages) and is the only one with cross-language framework edges.
Full side-by-side tables with GitHub stars, languages, and per-capability coverage: trace-mcp vs. other code intelligence MCP servers.
Head-to-head: vs Repomix · vs Serena · vs codegraph · vs codebase-memory-mcp · vs Claude Code context mode · vs code-review-graph · Repomix vs codegraph.
Token reduction — what we measured
AI agents burn tokens recomputing what they already discovered last turn — re-reading files, re-traversing dependencies, re-inflating context. trace-mcp replaces that with precision context: only the symbols, edges, and signatures relevant to the query, served from a graph that was computed once.
Start with the measurement that isn't ours. Everything else in this section is trace-mcp measured on trace-mcp's own repository — the first table row against real responses, everything below it by trace-mcp's own synthetic estimators. The PR review context benchmark is the exception: assembling review context for 60 merged pull requests across six open-source repositories — hono, axios, express, requests, flask, got — cost a median 3,291 input tokens against 13,595 for loading the diff plus every file it touches, 72.7% less, counted with gpt-tokenizer rather than estimated. The base and head SHAs are pinned in benchmarks/pr-context/dataset.json, npx tsx scripts/bench-pr-context.ts re-runs it, and the 56 pull requests where the index did not pay off are published alongside the wins — 13 that cost more than reading the files outright, and 42 where the bundle's token budget did not deliver every changed symbol's body, a shortfall the benchmark could not see until this run made it score delivery rather than listing.
What to expect — by workload:
Workload | Typical reduction |
Mixed real-world production (measured tool responses vs. the file reads they replace) | 67.4% fewer tokens |
Structured code-navigation tasks (symbol lookup, impact analysis, type hierarchy, call graph) | up to 99% less redundant processing — synthetic estimate |
Targeted research / planning queries (composite tasks that replace ~10 sequential operations) | up to ~40× on individual calls — synthetic estimate |
Non-code workloads (raw text, unstructured data) | Out of scope today |
The 67.4% is the honest number to plan against, and the reason it moved is not that the product got faster. We used to print "~40–50% on average". That figure descended from a counter that scored each call before the tool ran — RAW_COST_ESTIMATES[tool] × 0.15, a constant with no variance — which we found and fixed ourselves in #915. The honest replacements read 29.3%, then 21.1%, then 21.0% as coverage grew to 97.2% of recorded calls. Then we found that three tools carrying 76% of the weight were each priced from one sample, and that four defensible ways of choosing those samples price the same build at 21.0%, 30.7%, 56.0% and 67.4%. So the sampling frame is now generated, frozen and committed before the run that uses it, with every claim about how it was built re-checked against the repository in CI (preregistration). 67.4% is what the registered frame measures; read the jump from 21% as a change of frame, not a change of product. It weights real o200k_base counts of real responses by 18,329 recorded calls from one machine (per-tool table, generated into docs/_data/response_tokens.json). Three caveats travel with it: the baseline half — what a Read/Grep would have cost instead — is still a hand-written estimate; four of the twenty-three tools measured return more tokens than they replace, and the old counter booked a saving for them anyway; and two more tools (register_edit, reindex) replace no file read at all, so they are now credited zero and counted as overhead — with them on the spend side the all-in figure is 66.5%. The peaks below (up to 99% on individual structured calls) are a synthetic estimate, per-call, not per-session.
Measured at trace-mcp 3.31.0 (76996eb9) on 21 September 2026. Its preregistration publishes it as the first pass of the 25% bar we declared before measuring — and says in the same breath that the pass came from registering a sampling frame, not from shipping a faster product. The prediction written before that run named an interval the result landed above; it was wrong and it stays on the page.
Benchmark Lab — the same question, asked by the app. The desktop app's Benchmark Lab tab runs a pinned battery of 8 recall fixtures over three arms and saves every run to ~/.trace/benchmark-runs. Measured at trace-mcp 3.33.0 (4e1ac4fd) on 26 September 2026: the file-reading control spent 82,412 tokens over 13 calls (8/8 answered); the minimal arm answered 7 of 8 for 512 tokens (−99.4%); the standard arm answered 8 of 8 for 13,657 tokens (−83.4%). The minimal miss is a real measurement — raw search_text does not rank src/indexer/pipeline.ts in its top 10 — published rather than re-run until it passes. Method, battery hash and re-run →
Benchmark: trace-mcp's own codebase (694 files, 3,831 symbols → 929 files, 5,197 symbols in v1.30):
Task Without trace-mcp With trace-mcp Reduction
───────────────────────────────────────────────────────────────────────────
Symbol lookup 42,518 tokens 1,162 tokens 97.3%
File exploration 27,486 tokens 855 tokens 96.9%
Search 22,860 tokens 8,000 tokens 65.0%
Find usages 11,430 tokens 1,720 tokens 85.0%
Context bundle 12,847 tokens 3,485 tokens 72.9%
Batch overhead 16,831 tokens 8,299 tokens 50.7%
Impact analysis 49,141 tokens 1,856 tokens 96.2%
Call graph 178,345 tokens 9,285 tokens 94.8%
Type hierarchy 94,762 tokens 855 tokens 99.1%
Tests for 22,590 tokens 1,150 tokens 94.9%
Composite task 223,721 tokens 14,245 tokens 93.6%
───────────────────────────────────────────────────────────────────────────
Total 702,532 tokens 50,812 tokens 92.8%Across 11 structured task categories, recomputation drops by up to ~99% per call when the agent reuses the graph instead of re-reading files. Read that as a peak structured-task result on a well-supported TS/Vue codebase, not a number you should expect on every project. In production, on mixed workloads, expect 67.4% — the measured figure above, not this synthetic one. Less noise in context also means fewer hallucinations and better first-response accuracy — a quality benefit you don't see in token counts.
Savings scale with project size — argued, not measured. Without trace-mcp the agent reads more wrong files before finding the right one, while graph traversal stays O(relevant edges) rather than O(total files). We have no per-project-size measurement to put behind that, so this README no longer quotes one; the per-session token figure that used to sit here came from the same pre-#915 estimator as the "40–50%".
Composite tasks deliver the biggest wins. A single get_task_context call replaces a chain of ~10 sequential operations (search → get_symbol × 5 → Read × 3 → Grep × 2). That's one round-trip instead of ten, which is where most of the latency saving comes from.
Run it yourself
npx trace-mcp benchmark .Per-category token savings against your actual repo in ~5 minutes — no install, no signup, all local. It reads an existing index, so run trace-mcp index . first if the project isn't registered yet. Numbers above are from trace-mcp's own TypeScript/Vue codebase (929 files, 5,197 symbols) under structured benchmarks; production reduction on mixed workloads is lower (67.4% measured, see above), but the per-task patterns hold for any well-supported stack.
This is a synthetic estimate, not measured savings: the "without trace-mcp" side is computed from file sizes in the index, and the "with trace-mcp" side from per-scenario multipliers — not from actual tool calls. It shows the theoretical ceiling. To measure real savings from your own usage, run trace-mcp for a while, then:
trace-mcp analytics savings # real sessions: reads vs. what trace-mcp would have cost
trace-mcp analytics optimize # recommendations based on your actual usageSee Session analytics & token savings tracking for details.
Estimated using benchmark_project — it walks eleven task categories (symbol lookup, file exploration, text search, find usages, context bundle, batch overhead, impact analysis, call graph traversal, type hierarchy, tests-for, composite task context) over the indexed project. No trace-mcp tool is invoked. Every figure on both sides is a scenario-specific synthetic heuristic, and the heuristics differ per scenario. They draw on three kinds of input, mixed differently in each one:
Real values from the index — file
byte_length, symbol source and signature sizes. These carry the baseline for symbol lookup, file exploration, impact analysis and call graph traversal.Assumed result shapes for operations with no indexed equivalent — e.g. text search and find-usages baselines assume a fixed grep yield (matches × context lines × 80 chars),
get_tests_foris assumed to answer in ~400 characters, and the batch-overhead scenario adds fixed per-call MCP framing / hint / metadata token constants.A fixed fraction of the baseline, between 0.05 and 0.45, where neither of the above applies.
Character counts are converted to tokens by an estimator calibrated against cl100k_base when gpt-tokenizer is installed, and by a fixed chars-per-token ratio of 4.0 otherwise. The result is an upper bound on the reduction, not a measurement of it — the same caveats are printed in the tool output and documented at the top of src/analytics/benchmark.ts.
Reproduce it yourself:
# Via CLI (no install)
npx trace-mcp benchmark /path/to/project
# Or via MCP tool
benchmark_project # runs against the current projectKey capabilities
Request flow tracing — URL → Route → Middleware → Controller → Service, across backend frameworks
Component trees — render hierarchy with props / emits / slots (Vue, React, Blade)
Schema from migrations — no DB connection needed
Event chains — Event → Listener → Job fan-out (Laravel, Django, NestJS, Celery, Socket.io)
Change impact analysis — reverse dependency traversal across languages, enriched with linked architectural decisions
Graph-aware task context — describe a dev task → get the optimal code subgraph (execution paths, tests, types) + relevant past decisions, adapted to bugfix/feature/refactor intent
Call graph & DI tree — bidirectional call graphs with 4-tier resolution confidence, optional LSP enrichment for compiler-grade accuracy, NestJS dependency injection
ORM model context — relationships, schema, metadata for 7 ORMs
Dead code & test gap detection — find untested exports/symbols (with "unreached" vs "imported_not_called" classification), dead code, per-symbol test reach in impact analysis
Security scanning — OWASP Top-10 pattern scanning and taint analysis (source→sink data flow). Exportable MCP-server security context for skill-scan
Semantic search, offline by default — bundled ONNX embeddings work out of the box, no API keys; switch to Ollama/OpenAI for LLM-powered summarisation
Decision memory — mine sessions for decisions, link them to symbols/files, auto-surface in impact analysis
Multi-service subprojects — link graphs across services via API contracts; cross-service impact + service-scoped decisions
CI/PR change impact reports — automated blast radius, risk scoring, test-gap detection, architecture violations on every PR
Supported stack
Languages: PHP, TypeScript, JavaScript, Python, Go, Java, Kotlin, Ruby, Rust, C, C++, C#, Swift, Objective-C, Objective-C++, Dart, Scala, Groovy, Elixir, Erlang, Haskell, Gleam, Bash, Lua, Perl, GDScript, R, Julia, Nix, SQL, PL/SQL, HCL/Terraform, Protocol Buffers, GraphQL, Prisma, Vue SFC, HTML, CSS/SCSS/SASS/LESS, XML/XUL/XSD, YAML, JSON, TOML, Assembly, Fortran, AutoHotkey, Verse, AL, Blade, EJS, Zig, OCaml, Clojure, F#, Elm, CUDA, COBOL, Verilog/SystemVerilog, GLSL, Meson, Vim Script, Common Lisp, Emacs Lisp, Dockerfile, Makefile, CMake, INI, Svelte, Astro, Markdown, MATLAB, Lean 4, FORM, Magma, Wolfram/Mathematica, Ada, Apex, D, Nim, Pascal, PowerShell, Solidity, Tcl
Frameworks: Laravel (+ Livewire, Nova, Filament, Pennant), Django (+ DRF), FastAPI, Flask, Express, NestJS, Fastify, Hono, Next.js, Nuxt, Rails, Spring, tRPC
ORMs: Eloquent, Prisma, TypeORM, Drizzle, Sequelize, Mongoose, SQLAlchemy
Frontend: Vue, React, React Native, Blade, Inertia, shadcn/ui, Nuxt UI, MUI, Ant Design, Headless UI
Other: GraphQL, Socket.io, Celery, Zustand, Pydantic, Zod, n8n, React Query/SWR, Playwright/Cypress/Jest/Vitest/Mocha
Knowledge vaults: Obsidian, Logseq, plain markdown — [[wikilinks]], ![[embeds]], [text](path.md), frontmatter (YAML), #tags, ATX headings. Each note becomes a note:<basename> symbol with sections nested inside; wikilinks resolve to references / embeds edges between notes. Mix vault and code in one project — point root at a directory that contains both and run a single find_usages across them.
Full details: Supported frameworks · All tools
Quick start
See your waste first — 5 minutes, no setup, no signup:
npx trace-mcp benchmark .Indexes the project, runs 11 structured task benchmarks (symbol lookup, impact analysis, call graph, type hierarchy, …), and prints estimated per-task token cost — without trace vs. with. You'll see exactly where your agent recomputes work it could reuse. It is a synthetic estimate computed from your index, not a record of real tool calls (see the Methodology block under “Token reduction” above); for measured savings from your own sessions use trace-mcp analytics savings.
Then wire it into your AI agent:
npm install -g trace-mcp
trace init # one-time global setup (MCP clients, hooks, CLAUDE.md)
trace add # register current project for indexinginit— configures your MCP client (Claude Code, Cursor, Windsurf, Claude Desktop, …), installs the guard hook, adds routing rules to~/.claude/CLAUDE.md.add— detects frameworks, creates the per-project index, registers the project. Re-run in every project you want trace to understand.
(The npm package is still called trace-mcp — only the command it installs is shortened. trace-mcp init, trace-mcp add, and every other trace-mcp … invocation keep working.)
All state lives in ~/.trace/ (with automatic fallback from ~/.trace-mcp/) — your project directory stays clean unless you opt into .traceignore or .trace/.config.json.
Using Claude Code or Codex CLI? After npm install -g trace-mcp, skip trace init's client-wiring step and install the plugin directly instead — no git clone needed either way:
# Claude Code
claude plugin install @nikolai-vysotskyi/trace-mcp
# Codex CLI
codex plugin marketplace add nikolai-vysotskyi/trace-mcp
codex plugin install trace-mcp@nikolai-vysotskyi-trace-mcpBoth register the trace-mcp MCP server plus the Bash guard hook in one step. Details: .claude-plugin/README.md · .codex-plugin/README.md.
Then in your MCP client:
> get_project_map to see what frameworks are detected
> get_task_context("fix the login bug") to get full execution context for a task
> get_change_impact on app/Models/User.php to see what depends on itIndexing a markdown vault (Obsidian / Logseq / plain MD). Point trace add at the vault root — .md/.mdx/.markdown are picked up by default. Each note becomes a note:<basename> symbol, headings nest as sections, [[wikilinks]] and ![[embeds]] resolve to graph edges, frontmatter aliases: make alternate names resolvable, and #tags aggregate so every note carrying #sgr is one find_usages away.
> find_usages on note:my-concept // backlinks across the vault
> find_usages on tag:sgr // every note tagged #sgr
> get_change_impact on note:legacy // what breaks if I rename or delete it
> search "schema-guided reasoning" // PageRank + embeddings over the vaultPrefer a GUI? The desktop app handles install, indexing, MCP-client wiring, and re-indexing without touching a terminal.
Going further: adding more projects / upgrading / manual setup · stdio vs HTTP setup (per-repo or team) · semantic search (local ONNX) · indexing & file watcher · .traceignore.
Migration from trace-mcp to trace
The project is still trace-mcp. The command is now trace. The rename lives at exactly that boundary and nowhere else — the npm package, this repo, the domain, and the registry entry all keep the trace-mcp name. The reason is ergonomics, the same shape as rg for ripgrep or kubectl for kubernetes — not token savings: the measured saving from the shorter MCP tool prefix is real but small, 66–366 tokens per turn depending on tokenizer and preset, 0.74–1.23% of a tool list that already costs 8k–45k tokens.
The npm package name does not change. It is still trace-mcp, and it always will be — trace on npm is an unrelated package by another author. Install with npm install -g trace-mcp or npx -y trace-mcp@latest.
What does change, and what stays:
Command name —
trace <cmd>is the new spelling.trace-mcp <cmd>stays as an alias permanently — on macOS,/usr/bin/traceis Apple's owntrace(1), so keep usingtrace-mcpin scripts, CI, or any PATH you don't control yourself.MCP client entries —
trace initandtrace upgraderename an existingmcpServers["trace-mcp"]entry tomcpServers["trace"]and point it at the new command. Nothing is deleted; an entry left astrace-mcpkeeps working, it just costs more tokens.State directory —
~/.trace/, falling back to~/.trace-mcp/when the old one exists and the new one does not. Indexes are not rebuilt.Project config —
.trace.jsonis read first,.trace-mcp.jsonafter it. Existing files keep working where they are.Plugin and registry identifiers — unchanged:
@nikolai-vysotskyi/trace-mcpfor the Claude Code plugin,io.github.nikolai-vysotskyi/trace-mcpin the MCP registry.
One thing init can't do for you. The MCP tool prefix moves too — mcp__trace-mcp__search becomes mcp__trace__search. init migrates the mcpServers entry it owns, but not text you wrote yourself: Claude Code permission allowlists, hook matchers, or your own mcp__trace-mcp__* mentions in CLAUDE.md/AGENTS.md prose. If a hook stops matching or an allowlisted tool starts re-prompting after upgrading, grep your own config for mcp__trace-mcp__ and replace it with mcp__trace__. Everything else above happens automatically the next time you run trace init or trace upgrade — details: Configuration.
Local-first by design
trace-mcp runs entirely on your machine. Nothing about your source code is uploaded, and there is no account to create.
Indexing happens locally. The MCP server is a Node process you run yourself — stdio or
http://127.0.0.1:3741.Index lives in
~/.trace/(falling back to~/.trace-mcp/if that's what you already have), never inside your project and never uploaded. Your repo directory stays clean unless you opt into.traceignoreor.trace/.config.json.Semantic search is offline by default — bundled ONNX embeddings, no API keys, no outbound calls. Switch to Ollama (local) or OpenAI (opt-in) via config.
No telemetry about your code, queries, or usage. The only thing that ever leaves your machine is described below and on the privacy page — nothing else is phoned home.
What your AI client sees is governed by your AI client. trace-mcp returns graph results over MCP; how Claude Code / Cursor / Codex / Windsurf forward them to a model is up to that client's privacy model.
The daemon trusts loopback and nothing else.
serve-httpis unauthenticated by design: a caller on127.0.0.1is already you. A non-loopback--hostis therefore refused unless you pass--allow-remoteand front the port with your own auth — see Configuration.To wipe everything, delete
~/.trace/(or~/.trace-mcp/on an install that hasn't migrated yet) — that directory is the whole footprint.
Usage telemetry
trace-mcp sends at most one anonymous ping per day, per install, so we can count active installs: version, OS, MCP client, and aggregate counts. No code, no paths, no IP address, and no per-install identifier beyond a UUID generated locally on your machine. It is suppressed in CI, and its GA4 credentials ship as plaintext in the published bundle so you can verify where the ping goes.
Turn it off with TRACE_MCP_TELEMETRY=off, or with "telemetry": { "usage_ping": false } in ~/.trace/.config.json.
The complete field list, both opt-outs and how to delete local state are on the privacy page. Source: src/telemetry/usage-ping.ts.
For security-sensitive environments, review SECURITY.md before use.
Getting the most out of trace-mcp
trace-mcp works on three levels to make AI agents use its tools instead of raw file reading:
Level 1: Automatic (works out of the box)
The MCP server provides instructions and tool descriptions with routing hints that tell AI agents when to prefer trace-mcp over native Read/Grep/Glob. This works with any MCP-compatible client — no configuration needed.
Level 2: CLAUDE.md (recommended)
trace-mcp init adds a Code Navigation Policy block to ~/.claude/CLAUDE.md (or your project's CLAUDE.md) that tells the agent which trace-mcp tool to prefer over Read/Grep/Glob for each kind of task. If you skipped init, see System prompt routing for the full block and how to tune enforcement.
Level 3: Hook enforcement (Claude Code only)
For hard enforcement, trace-mcp init installs a PreToolUse guard hook that blocks Read/Grep/Glob on source files and redirects the agent to trace-mcp tools (non-code files, Read-before-Edit, and safe Bash commands pass through). Manage manually with trace-mcp setup-hooks --global / --uninstall. Details: System prompt routing.
Level 4: Max tier — system prompt rewrites + agent behavior rules
Picking Max during trace-mcp init (the default) layers on two more amplifiers:
tweakcc system-prompt rewrites patch Claude Code's core tool descriptions so the model internalizes "use trace-mcp search" instead of "use Grep" from the start. Claude Code only.
agent_behavior: "strict"ships a compact set of discipline rules via MCP instructions — no flattery, disagree on wrong premises, never fabricate, goal-driven execution, 2-strike session hygiene, no drive-by refactors. Cross-client (Claude Code, Cursor, Codex, Windsurf) and auto-updates onnpm upgrade trace-mcpwithout re-runninginit.
This is the setup for making the same discipline rules apply to every teammate's agent without asking anyone to configure it. Tune or disable via tools.agent_behavior in ~/.trace/.config.json — see Tool exposure & agent behavior.
Decision memory
Decisions, tradeoffs, and discoveries from AI-agent conversations usually vanish when the session ends. trace-mcp captures them and links each decision to the code it's about — so when someone later runs get_change_impact on src/db/connection.ts::Pool#class, the "we chose PostgreSQL for JSONB" decision surfaces automatically.
Mine —
mine_sessionsscans Claude Code / Claw Code JSONL logs and extracts decisions via pattern matching (0 LLM calls). Types: architecture, tech choice, bug root cause, tradeoff, convention.Link — each decision attaches to a symbol or file; supports service-scoped decisions for subprojects.
Surface — decisions auto-enrich
get_change_impact,plan_turn, andget_wake_up. Temporal validity (valid_from/valid_until) makes "what was true on 2025-01-15?" queries possible.Search —
query_decisions(FTS5 + filters) for decisions;search_sessionsfor raw conversation content across all past sessions.
trace memory mine # extract decisions from sessions
trace memory search "GraphQL migration" # search past conversations
trace memory timeline --file src/auth.ts # decision history for a fileFull tool list, CLI, temporal validity, service scoping: Decision memory.
Subprojects
A subproject is any repo in your project's ecosystem — microservice, frontend, shared lib, CLI tool. trace links dependency graphs across subprojects: if service A calls an endpoint in service B, changing the endpoint in B shows up as a breaking change for A.
Discovery is automatic. On each index, trace detects subprojects (Docker Compose, flat/grouped workspaces, monolith fallback), parses API contracts (OpenAPI, GraphQL SDL, Protobuf/gRPC), scans code for HTTP client calls (fetch, axios, Http::, requests, http.Get, gRPC stubs, GraphQL ops), and links the calls to known endpoints.
cd ~/projects/my-app && trace add
# → auto-detects user-service (openapi.yaml) and order-service
# → links order-service → user-service via /api/users/{id}
trace subproject impact --endpoint=/api/users
# → [order-service] src/services/user-client.ts:42 (axios, confidence: 85%)External subprojects can be added manually with trace subproject add --repo=... --project=.... MCP tools: get_subproject_graph, get_subproject_impact, get_subproject_clients, subproject_add_repo, subproject_sync.
Full CLI, detection modes, MCP-tool reference, topology config: Configuration — topology & subprojects.
CI/PR change impact reports
trace ci-report --base main --head HEAD produces a markdown or JSON report per pull request: summary, blast radius (depth-2 reverse dep traversal), test coverage gaps (per-symbol hasTestReach), risk analysis (30% complexity + 25% churn + 25% coupling + 20% blast radius), architecture violations (auto-detects clean / hexagonal presets), and new dead exports.
Use --fail-on high to block merges on high-risk changes. See .github/workflows/ci.yml for a ready-to-use GitHub Action that runs build → test → impact-report and posts a sticky PR comment on every push.
Pilot program — for teams running LLM in production
If you're shipping AI features in production — internal copilots, customer-facing assistants, RAG over a code or knowledge base — and you're hitting cost, latency, or quality ceilings, we'll run a focused pilot with you.
Format: 2–4 weeks. Minimal integration. One or two real production use cases — not a demo.
What we measure (before / after):
Tokens per successful answer
First-response accuracy (% of queries resolved without retry)
Retries and fallback calls
End-to-end latency
User success rate on a fixed evaluation set
What you get: a clear, before/after report on whether context optimization moves the metrics that matter for your stack — and a path to scale usage with confidence instead of throttling it on cost.
The target is a system that stays predictable as usage grows, not a one-off cost cut: teams usually want to reach reliable production first and expand their LLM footprint after.
Get in touch: open an issue at github.com/nikolai-vysotskyi/trace-mcp/issues tagged pilot, or reach out to @nikolai-vysotskyi.
How it works
Source files (PHP, TS, Vue, Python, Go, Java, Kotlin, Ruby, HTML, CSS, Blade)
│
▼
┌──────────────────────────────────────────┐
│ Pass 1 — Per-file extraction │
│ tree-sitter → symbols │
│ integration plugins → routes, │
│ components, migrations, events, │
│ models, schemas, variants, tests │
└────────────────────┬─────────────────────┘
│
▼
┌──────────────────────────────────────────┐
│ Pass 2 — Cross-file resolution │
│ PSR-4 · ES modules · Python modules │
│ Vue components · Inertia bridge │
│ Blade inheritance · ORM relations │
│ → unified directed edge graph │
└────────────────────┬─────────────────────┘
│
▼
┌──────────────────────────────────────────┐
│ Pass 3 — LSP enrichment (opt-in) │
│ tsserver · pyright · gopls · │
│ rust-analyzer → compiler-grade │
│ call resolution, 4-tier confidence │
└────────────────────┬─────────────────────┘
│
▼
┌──────────────────────────────────────────┐
│ SQLite (WAL mode) + FTS5 │
│ nodes · edges · symbols · routes │
│ + embeddings (local ONNX by default) │
│ + optional: LLM summaries │
└────────────────────┬─────────────────────┘
│
▼
┌──────────────────────────────────────────┐
│ Decision Memory (decisions.db) │
│ decisions · session chunks · FTS5 │
│ temporal validity · code linkage │
│ auto-mined from session logs │
└────────────────────┬─────────────────────┘
│
▼
MCP server (stdio or HTTP/SSE)
182 tools · 10 resourcesIncremental by default — files are content-hashed; unchanged files are skipped on re-index.
Plugin architecture — language plugins (symbol extraction) and integration plugins (semantic edges) are loaded based on project detection, organized into categories: framework, ORM, view, API, validation, state, realtime, testing, tooling.
Documentation
Full docs live at trace-mcp.com (same content as docs/ in this repo).
Document | Description |
Complete list of languages, frameworks, ORMs, UI libraries, and what each extracts | |
All 182 MCP tools with descriptions and usage examples | |
The seven tools retired in 2.0 ( | |
Config options, AI setup, environment variables, security settings | |
How indexing works, plugin system, project structure, tech stack | |
Decision knowledge graph, session mining, cross-session search, wake-up context | |
Session analytics, token savings tracking, optimization reports, benchmarks | |
Complexity, security and coupling thresholds, and how | |
Measured token savings of the TOON output format on real tool calls | |
OpenTelemetry-compatible spans for every AI provider call and MCP tool call | |
Optional tweakcc integration for maximum tool routing enforcement | |
Full side-by-side tables vs. other code intelligence / memory / RAG tools | |
Building, testing, contributing, adding new plugins | |
The desktop app's macOS 26 design system — tokens, type, geometry, materials, primitives, accessibility floors |
Star History
Project health
License
Built by Nikolai Vysotskyi
Available Tools
29 toolsbatchARead-onlyIdempotent
Execute multiple trace-mcp tools in a single MCP request. Dispatches any registered tool by name, including tools this session's preset defers — so a deferred tool is callable here without a load_tools round-trip (tools.exclude stays a hard restriction). Use to reduce round-trips when you need several independent queries (e.g., get_outline for 3 files, or search + get_symbol together). Read-only (delegates to other tools). Returns JSON: { batch_results: [{ tool, result }], total }.
| Name | Required | Description | Default |
|---|---|---|---|
| calls | Yes | Array of tool calls to execute (max 10) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnly/idempotent annotations, the description explains that the tool delegates to other tools, can invoke deferred tools without load_tools, enforces tools.exclude as a hard restriction, and documents the JSON return shape. This adds meaningful behavioral context not present in annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences cover purpose, use case, restrictions, read-only behavior, and return shape. Every sentence earns its place and key information is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema, the description provides the return format. The schema covers parameter validation, annotations cover safety, and the description covers usage, restrictions, and batching intent. Nothing essential for an agent to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already fully documents the 'calls' parameter with item schemas and min/max constraints, so the baseline is 3. The description adds value with concrete usage examples and clarifies that deferred tool names are acceptable, going slightly beyond schema-only information.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states a specific verb ('Execute multiple trace-mcp tools in a single MCP request') and distinguishes itself from siblings by enabling batched dispatch, including deferred tools. It is immediately clear what this tool does and how it differs from individual tools like get_outline or search.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Use to reduce round-trips when you need several independent queries,' gives concrete examples, and clarifies when a tool is still not callable ('tools.exclude stays a hard restriction'). This gives clear when-to-use and when-not-to-use guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
find_usagesARead-onlyIdempotent
Find all references to a symbol or file (imports, calls, renders, dispatches). Use instead of Grep for symbol usages — semantic, not text matches. For raw text use search_text; for a bidirectional call graph use get_call_graph. Weakly-grounded text_matched edges into a name-colliding target are dropped by default (phantom god-node filter); include_ambiguous_text_matched: true keeps them. Read-only. Returns JSON: { references: [{ edge_type, resolution_tier, file, symbol }], total, truncated?, ambiguous_filtered? } — page caps at 50, total counts all.
| Name | Required | Description | Default |
|---|---|---|---|
| fqn | No | Fully qualified name to find references for | |
| limit | No | Max references returned (default 50). | |
| file_path | No | File path to find references for | |
| symbol_id | No | Symbol ID to find references for | |
| detail_level | No | Output verbosity. "minimal" saves ~40-60% tokens (drops scores, fqn, signatures, summaries). Use to pick a candidate before get_symbol. Default: "default". | |
| include_ambiguous_text_matched | No | Keep text_matched edges whose target name collides with >=3 other symbols (default false — they produce phantom god-nodes). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the annotations (readOnlyHint, idempotentHint, destructiveHint), the description discloses important behaviors: read-only status, the JSON return shape, the 50-item page cap, total counting, truncation, and the phantom-god-node filtering semantics. This is much richer than merely repeating annotation flags.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded, with every clause earning its place. It covers purpose, alternatives, edge-case filtering, read-only status, and return shape without fluff or redundant schema repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only query tool with no output schema, the description compensates fully by specifying the response structure, pagination behavior, and filtering caveat. It also names the relevant sibling tools, making it complete for an agent deciding whether and how to call it.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real value: it explains the ambiguous-filter behavior, the token savings of detail_level='minimal', and the default/limit semantics. It does not clarify how fqn, file_path, and symbol_id interact (e.g., mutual exclusivity), but overall it meaningfully extends the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific, unambiguous definition: 'Find all references to a symbol or file (imports, calls, renders, dispatches).' It also distinguishes itself from Grep and sibling tools by emphasizing semantic vs text matches, so an agent can clearly tell this tool apart from search_text and get_call_graph.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit routing guidance: 'Use instead of Grep for symbol usages — semantic, not text matches. For raw text use search_text; for a bidirectional call graph use get_call_graph.' This tells the agent both when to use this tool and which alternatives to choose in different situations.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_call_graphARead-onlyIdempotent
Build a bidirectional call graph centered on a symbol (who calls it + what it calls). Each branch keeps its direction: depth 2 = callers of callers, callees of callees. Use to understand control flow through a function. For flat list of all references use find_usages instead. Read-only. Returns JSON: { root: { symbol_id, name, calls: [...], called_by: [...] } }.
| Name | Required | Description | Default |
|---|---|---|---|
| fqn | No | Fully qualified name to center the graph on | |
| depth | No | Traversal depth on each side (default 2) | |
| symbol_id | No | Symbol ID to center the graph on |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, and the description reinforces 'Read-only'. More importantly, it adds non-obvious behavior: branches keep their direction, depth 2 means callers of callers and callees of callees, and the JSON response shape is provided. A small gap is unspecified behavior when neither fqn nor symbol_id is provided.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: it opens with the core behavior, immediately explains depth semantics, provides an explicit alternative, and closes with output format. Every sentence earns its place with no redundancy or filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool is well covered: purpose, usage alternative, depth behavior, read-only safety, and return structure are all present, and annotations cover the safety profile. The main omission is that no parameter is required in the schema while the description assumes a symbol is centered, leaving the fqn-or-symbol_id contract implicit.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds useful semantics beyond the schema by explaining what depth means concretely ('callers of callers, callees of callees') and reinforcing the bidirectional traversal. It does not clarify how fqn and symbol_id relate or take precedence, but the main parameter semantics are enriched.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description is specific: it names the exact operation ('Build a bidirectional call graph centered on a symbol'), clarifies both call directions, and explains depth semantics. It also distinguishes itself from find_usages, making the tool's unique purpose immediately clear.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use the tool ('Use to understand control flow through a function') and when not to, directing the agent to the correct alternative ('For flat list of all references use find_usages instead'). This is strong routing guidance among sibling tools.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_change_impactARead-onlyIdempotent
Full change impact report: risk score + mitigations, breaking change detection, enriched dependents (complexity, coverage, exports), module groups, affected tests, co-change hidden couplings. Pass symbol_ids to scope to changed symbols. Use before modifying code. For a quick risk score alone use assess_change_risk; for who-calls-what use get_call_graph. compact pages results; bundle recalls. Read-only. Returns JSON: { risk, dependents, affectedTests, breakingChanges, totalAffected }.
| Name | Required | Description | Default |
|---|---|---|---|
| fqn | No | Fully qualified name to analyze (alternative to symbol_id) | |
| depth | No | Max traversal depth (default 3) | |
| bundle | No | Handle; @N = page N. | |
| compact | No | Paged recall (opt-in). | |
| file_path | No | Relative file path to analyze | |
| symbol_id | No | Symbol ID to analyze | |
| symbol_ids | No | Diff-aware: only analyze impact of these specific symbols (e.g. from get_changed_symbols) | |
| max_dependents | No | Cap on returned dependents (default 200) | |
| decorator_filter | No | Filter dependents to only those with this decorator/annotation/attribute (e.g. "Route", "Transactional", "csrf_protect") |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark the tool read-only and idempotent, and the description explicitly states 'Read-only.' It adds useful behavioral detail beyond annotations: paging via compact, bundle recall semantics, and the returned JSON shape. It doesn't overstate side effects or contradict the safety hints.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense and front-loaded; the first sentence states the core value, and later sentences cover usage, alternatives, paging, and output. Minor terseness like 'compact pages results; bundle recalls' could be clearer, but there is no filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex 9-parameter tool with no output schema, the description does enough by summarizing the report contents and the main return fields. It does not fully explain how the multiple locator params (fqn, symbol_id, file_path, symbol_ids) interact or what happens with no params, leaving a small gap in an otherwise complete picture.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the structured parameters already define fqn, depth, bundle, symbol_ids, max_dependents, decorator_filter, etc. The description adds guidance for symbol_ids as a diff-aware scope and mentions compact/bundle, but it does not materially enrich the parameters beyond the schema, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description defines a concrete deliverable ('Full change impact report') and enumerates its contents: risk score, mitigations, breaking-change detection, enriched dependents, module groups, affected tests, and co-change hidden couplings. It also contrasts itself with get_call_graph, so an agent can distinguish it from the closest sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit trigger ('Use before modifying code'), scoping advice ('Pass symbol_ids to scope to changed symbols'), and routes quick risk checks to assess_change_risk and call-graph queries to get_call_graph. However, assess_change_risk is not present in the provided sibling list, which slightly weakens the referral; still, the conditions of use are clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_context_bundleARead-onlyIdempotent
Get a symbol's source code + its import dependencies + optional callers, packed within a token budget. Supports batch queries with shared-import deduplication. Use instead of chaining get_symbol calls. For a single symbol without imports, use get_symbol — lighter. Read-only. Returns JSON: { primary: [{ symbol_id, file, source }], imports: [{ file, source }], token_usage }.
| Name | Required | Description | Default |
|---|---|---|---|
| fqn | No | Alternative: look up by FQN | |
| symbol_id | No | Single symbol ID | |
| symbol_ids | No | Batch: multiple symbol IDs | |
| token_budget | No | Max tokens (default 8000) | |
| output_format | No | Output format (default json). | |
| include_callers | No | Include who calls these symbols (default false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already carry readOnlyHint=true and idempotentHint=true, and the description's 'Read-only' agrees with them rather than contradicting them. The description adds genuine behavior beyond annotations: token-budget packing, shared-import deduplication for batch queries, and the exact return shape (primary, imports, token_usage). It loses a point because the 'Returns JSON' claim ignores the markdown output_format option, and truncation/error behavior when the budget is exceeded is undisclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, each earning its place: function, batch/dedup semantics, routing to the sibling tool, and return shape. The core capability is front-loaded in the first clause with zero filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 6-parameter read-only tool with no output schema, the description covers the essentials: return shape, token budget behavior, batch semantics, and the lighter alternative. Gaps remain at the edges — the interplay between fqn/symbol_id/symbol_ids and behavior when the token budget is exceeded — but the per-parameter schema descriptions and rich annotations carry much of that load.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3; the schema already documents all six parameters. The description earns an extra point by adding intent behind key parameters: 'packed within a token budget' clarifies the purpose of token_budget, and 'batch queries with shared-import deduplication' gives symbol_ids behavioral meaning beyond its bare listing.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb-object pair — 'Get a symbol's source code + its import dependencies + optional callers' — making the tool's function unmistakable. It also explicitly differentiates itself from get_symbol, its closest sibling, by describing what this tool bundles that get_symbol does not.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit routing rules: 'Use instead of chaining get_symbol calls' for import-heavy or batch needs, and 'For a single symbol without imports, use get_symbol — lighter' as the exclusion condition with the alternative named. This is exactly the when/when-not guidance this dimension requires.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_coverage_reportARead-onlyIdempotent
Technology profile of the project: detected frameworks/ORMs/UI libs from manifests (package.json, composer.json, etc.), which are covered by trace-mcp plugins, and coverage gaps. Read-only. Returns JSON: { detected, covered, gaps }.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and non-destructive behavior; the description adds value by specifying the data source (manifests), the read-only nature, and the exact return shape { detected, covered, gaps }. This goes beyond the structured annotations without contradicting them.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the core purpose, followed by the read-only trait and return shape. Every sentence earns its place, with no redundant elaboration.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter, read-only report with no output schema, the description is complete: it explains what data is gathered, from where, and what the JSON response contains. An agent has enough information to invoke and interpret the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the baseline is 4. The description provides enough context about the returned fields to make the no-argument call understandable, and there is no parameter documentation gap.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly identifies the tool as producing a technology profile with detected frameworks/ORMs/UI libs and coverage gaps, which is specific and not a tautology. It distinguishes the report's focus on plugin coverage from siblings like get_optimization_report or get_usage_trends, though it does not explicitly name any sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is implied rather than stated: an agent can infer this tool is for inspecting project technology coverage and gaps, but there is no explicit 'use when' guidance or mention of alternatives. No exclusion criteria are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_diagnosticsARead-onlyIdempotent
Execute type-checker (tsc, mypy, pyright) and map errors to enclosing AST symbols. Read-only.
| Name | Required | Description | Default |
|---|---|---|---|
| checker | No | Checker (tsc, mypy, pyright) | |
| file_path | No | Filter by file path | |
| max_files | No | Max files reported | |
| timeout_ms | No | Timeout in ms | |
| max_per_file | No | Max errors per file | |
| reduce_output | No | Shrink long output to a receipt |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, idempotent, and non-destructive behavior. The description adds value by disclosing that external type-checkers are executed and that results are mapped to enclosing AST symbols. It also explicitly repeats the read-only trait, consistent with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single sentence with no filler. Every clause carries meaning: the action, the specific checkers, the result transformation, and the safety characteristic.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with no output schema, the description gives a useful sense of what happens and what the result is about (errors mapped to AST symbols). Parameter details are fully covered by the schema. It stops short of describing the exact return shape or the effect of reduce_output, but it is adequate for selection and invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all six parameters are already documented. The description adds minimal parameter meaning beyond listing the same type-checker names found in the enum, so the baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource: execute type-checkers (tsc, mypy, pyright) and map errors to enclosing AST symbols. It clearly identifies what the tool does, even though it does not explicitly contrast itself with sibling tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage context is implied: an agent can infer to use this tool when type-checking diagnostics are needed. However, there is no explicit guidance about when to choose this over alternatives like get_index_health, get_coverage_report, or get_optimization_report, and no exclusions are stated.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_feature_contextARead-onlyIdempotent
Search code by keyword/topic → returns ranked source snippets within a token budget. Use when you need to READ actual code for a concept or feature. For structured task context with tests and entry points use get_task_context instead; for symbol metadata without source use search. Read-only. Returns JSON (default) or Markdown: { items: [{ symbol_id, name, file, source, score }], token_usage } | { content: "...markdown..." }. Supports output_format: "toon". Capped by memory.recall.timeoutMs (default 5000ms); on timeout returns { items: [], token_usage, degraded: true }.
| Name | Required | Description | Default |
|---|---|---|---|
| description | Yes | Natural language description of the feature to find context for | |
| detail_level | No | Output verbosity. "minimal" saves ~40-60% tokens (drops scores, fqn, signatures, summaries). Use to pick a candidate before get_symbol. Default: "default". | |
| token_budget | No | Max tokens for assembled context (default 4000) | |
| output_format | No | "json" (default, structured items), "markdown" (fenced code blocks, ~15-20% cheaper), or "toon" (lossless, 30-60% fewer tokens). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so that baseline is covered. The description adds valuable non-obvious behavior: timeout behavior ('Capped by memory.recall.timeoutMs (default 5000ms); on timeout returns { items: [], token_usage, degraded: true }'), which is not present in annotations or schema. It also states read-only and return format, adding context beyond structured fields.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is efficient: first sentence states purpose, second gives usage and alternatives, third states read-only and return format, fourth adds output_format nuance, fifth covers timeout behavior. Each sentence earns its place, though the timeout detail adds length. It is appropriately sized for the tool's complexity and well front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description must explain return values, and it does: 'Returns JSON (default) or Markdown: { items: [{ symbol_id, name, file, source, score }], token_usage } | { content: "...markdown..." }.' It also covers the timeout degraded response. It doesn't mention rate limits or permissions, but those are not critical given read-only annotations. The description is complete enough for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% (all 4 parameters have descriptions), so the baseline is 3. The description does mention token_budget ('token budget') and output_format ('toon'), but these are already fully documented in the schema with comparable detail. It adds no new semantic meaning for parameters beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource ('Search code by keyword/topic → returns ranked source snippets within a token budget'), making the core function immediately clear. It also differentiates from siblings by naming get_task_context and search with their distinct use cases, so an agent can tell them apart without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly states when to use this tool ('when you need to READ actual code for a concept or feature') and gives alternatives with guidance: 'For structured task context with tests and entry points use get_task_context instead; for symbol metadata without source use search.' This satisfies the explicit when/when-not/alternative criterion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_index_healthARead-onlyIdempotent
Get index status, statistics, health, and pipeline progress (indexing, summarization, embedding). Includes the session projectRoot; when the index is empty, next_steps names the empty root and points at list_projects + call_project_tool for other registered projects. Read-only, no side effects. Use to verify the index is ready before running queries. Returns JSON: { status, stats, projectRoot, next_steps?, config, warnings, pipelineProgress, embedding }.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already carry readOnly/idempotent/destructive safety, and the description reinforces this and adds genuinely new behavioral context: the empty-index next_steps behavior, the projectRoot inclusion, and the exact JSON envelope. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is two sentences: a front-loaded purpose statement and a state-sensitive behavior note, plus a compact return-shape hint. Every sentence contributes information relevant to correct invocation.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description carries the full burden of explaining return values, and it enumerates the JSON fields. It also explains empty-index behavior and provides follow-up tool names, making it complete for a zero-parameter read-only health check.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so the schema leaves nothing to document and the baseline is 4. The description correctly adds no parameter syntax and instead uses the space to describe the return payload.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific action-resource pair ('Get index status, statistics, health, and pipeline progress') and spells out the payload, so an agent knows exactly what the tool reports. It is distinct in subject matter from siblings like get_session_stats, but it does not explicitly name or contrast those alternatives, so it stops just short of top differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives a clear invocation condition: 'Use to verify the index is ready before running queries.' It also hints at routing to list_projects and call_project_tool when the index is empty. No explicit when-not-to-use or alternative health tools are named, so it misses the top criterion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_optimization_reportARead-onlyIdempotent
Detect token waste patterns in AI agent sessions: repeated file reads, Bash grep instead of search, large file reads, unused trace-mcp tools. Provides savings estimates. Read-only. For usage/cost overview use get_session_analytics; for A/B savings comparison use get_real_savings. Returns JSON: { patterns: [{ type, description, savings_estimate }], total_waste }.
| Name | Required | Description | Default |
|---|---|---|---|
| period | No | Time period (default: week) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, so the safety type is covered. The description adds value beyond annotations by describing the JSON return format ({ patterns: [{ type, description, savings_estimate }], total_waste }) and stating it 'Provides savings estimates.' No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: the first names specific waste patterns, the second covers read-only safety and savings estimates, the third gives the return shape and sibling routing. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple one-optional-parameter tool with no output schema, the description provides the return JSON structure, the tool's scope, and explicit sibling alternatives. Nothing needed to invoke it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema covers 100% of the single parameter, including enum values and default. The description does not add any parameter-level detail, so the schema carries the full burden. Baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Detect token waste patterns in AI agent sessions,' then lists concrete pattern types (repeated file reads, Bash grep instead of search, large file reads, unused trace-mcp tools). It also provides the return shape and distinguishes itself from siblings by naming their use cases.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly routes to alternatives with conditions: 'For usage/cost overview use get_session_analytics; for A/B savings comparison use get_real_savings.' This tells the agent when not to use this tool and which sibling to pick instead.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_outlineARead-onlyIdempotent
Get all symbols for a file (signatures only, no bodies) — cheaper than Read for understanding a file before editing. Follow up with get_symbol to read one symbol's source. nested: true expands large top-level symbols (default ≥100 LOC) into inner declarations, each carrying parentId + depth (max 3). Read-only. Returns JSON: { path, language, symbols: [{ symbolId, name, kind, signature, lineStart, lineEnd, parentId?, depth? }] }. Supports output_format: "toon".
| Name | Required | Description | Default |
|---|---|---|---|
| path | Yes | Relative file path | |
| nested | No | Walk the body of each top-level symbol past min_loc_for_nesting and emit inner declarations as extra rows carrying `parentId` + `depth`. Default false. | |
| detail_level | No | Output verbosity. "minimal" saves ~40-60% tokens (drops scores, fqn, signatures, summaries). Use to pick a candidate before get_symbol. Default: "default". | |
| output_format | No | "json" (default) or "toon" (lossless, 30-60% fewer tokens). "markdown" is unsupported here and behaves as json. | |
| min_loc_for_nesting | No | Minimum (line_end - line_start) for a top-level symbol to be expanded when nested=true. Default 100. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, destructiveHint=false, so the safety profile is covered. The description adds behavioral value beyond annotations: the nesting expansion semantics (parentId + depth, max 3, default ≥100 LOC), the 'Read-only' confirmation, and the explicit return JSON shape. No contradiction with annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Purpose is front-loaded in the first sentence. The description is dense but every clause earns its place: cost comparison, follow-up workflow, nesting behavior, return format, and toon support. Slightly long but structured well with a clear return-shape example that helps an agent parse the output.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only outline tool with 5 parameters and 2 enums, the description covers the core purpose, cost benefit, follow-up routing, nesting semantics, return structure, and output formats. Minor gaps: no mention of error behavior (e.g., file-not-found or unsupported language), but for a non-destructive tool with strong annotation coverage this is reasonably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3 and the schema carries the parameter documentation. The description adds some marginal value — clarifying the toon output format support, nesting depth cap, and default LOC threshold — but these mostly restate or lightly extend what the schema already covers. Does not fully compensate beyond the baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ('Get all symbols for a file') with a precise scope qualifier ('signatures only, no bodies'), and distinguishes itself from the sibling get_symbol by naming the follow-up workflow. An agent can immediately tell what this does and how it differs from its sibling.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly positions the tool: 'cheaper than Read for understanding a file before editing' tells when to prefer it, and 'Follow up with get_symbol to read one symbol's source' defines the recommended workflow. The detail_level schema description also guides when to use 'minimal' ('Use to pick a candidate before get_symbol'). Clear usage context with no ambiguity.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_preset_infoARead-onlyIdempotent
Show active tool preset, available presets, which tools are registered in this session, and which are deferred (loadable via load_tools). Read-only. Returns JSON: { active_preset, registered_tools, tool_names, available_presets, deferred_tools }.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only, idempotent, and non-destructive behavior. The description reinforces this with 'Read-only' and adds the return JSON shape (active_preset, registered_tools, etc.), disclosing what the agent will receive. This extra context goes beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two concise sentences, front-loaded with the core purpose and followed by a compact JSON key listing. Every sentence adds value; no filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read-only introspection tool, the description is complete: it states what the tool reports, the read-only nature, and the exact response fields. There is no output schema, so the description carries the burden of return-value disclosure and does so adequately.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so schema coverage is trivially 100%. Baseline for zero-parameter tools is 4; the description need not document parameter semantics because there are none.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb ('Show') and enumerates the exact resources: active tool preset, available presets, registered tools, and deferred tools. It clearly differentiates this introspection tool from the many sibling get_* tools and explicitly connects deferred tools to load_tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description makes the tool's context clear: use it to inspect tool registration and preset state. It mentions that deferred tools are loadable via load_tools, providing adjacent guidance, though it does not explicitly state when not to use it or name a directly competing alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_project_mapARead-onlyIdempotent
Get project overview: detected frameworks, languages, file counts, structure. Read-only, no side effects. Call with summary_only=true at session start to orient yourself before diving into code. Use instead of manual ls/find. Returns JSON: { frameworks, languages, fileCount, symbolCount, structure }.
| Name | Required | Description | Default |
|---|---|---|---|
| summary_only | No | Return only framework list + counts (default false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint, so the safety profile is covered. The description adds useful behavioral context beyond that: it confirms 'no side effects', suggests a cheap orientation call pattern, and specifies the return shape. This exceeds the annotation baseline without contradicting it.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four concise sentences front-load the core behavior and return shape. The 'Read-only, no side effects' clause is slightly redundant with annotations, but the rest earns its place by adding usage guidance and output fields.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is complete for a simple, optional-parameter read-only tool. It names the return JSON fields, gives a recommended invocation, and explains the tool's role. It could mention symbolCount semantics or potential costs, but nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents the only parameter. The description adds the session-start use case for summary_only=true, which is helpful, but it does not add meaning beyond what the parameter description already conveys.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Get') and resource ('project map') and enumerates contents (frameworks, languages, file counts, structure). It does not explicitly distinguish itself from sibling tools like get_outline or get_context_bundle, but the term 'project map' combined with the described fields is clear enough for an agent to understand what it returns.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear usage context: call at session start with summary_only=true to orient before diving into code, and use instead of manual ls/find. It does not discuss when not to use it or name alternative tools, but the intended scenario is explicit and actionable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_real_savingsARead-onlyIdempotent
A/B comparison: how many tokens could be saved by using trace-mcp instead of raw Read/Bash file reads. Per-file breakdown. Read-only. For pattern-based waste detection use get_optimization_report instead. Returns JSON: { files: [{ file, raw_tokens, compact_tokens, savings }], total_savings }.
| Name | Required | Description | Default |
|---|---|---|---|
| period | No | Time period (default: week) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and non-destructive hints, so the description does not need to repeat safety traits. It adds useful behavioral context by describing the exact return shape: files with raw_tokens, compact_tokens, savings, and total_savings. This is valuable because there is no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the core purpose, then clearly differentiates the tool from a sibling, and finishes with the return JSON shape. Every sentence earns its place with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a simple read-only tool with one optional parameter, the description is complete: it explains what the tool does, when to use it, what it returns, and how it differs from the closest sibling. No important context is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the only parameter, period, is fully documented in the schema with an enum and default of 'week'. The description does not add parameter-level meaning beyond the schema, which is expected given high coverage.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose with a specific verb and resource: 'A/B comparison: how many tokens could be saved by using trace-mcp instead of raw Read/Bash file reads.' It also specifies the per-file breakdown, making it easy to distinguish from other reporting tools like get_optimization_report.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description explicitly directs when to use this tool versus an alternative: 'For pattern-based waste detection use get_optimization_report instead.' This provides a clear exclusion and alternative, so an agent can route correctly without opening the schema.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_session_analyticsARead-onlyIdempotent
Analyze AI agent session logs: token usage, cost breakdown by tool/server, top files, models used. Parses Claude Code JSONL logs automatically. Read-only. For waste detection use get_optimization_report; for cost trends use get_usage_trends. Returns JSON: { sessions, tokens, cost_usd, tools, models, topFiles }.
| Name | Required | Description | Default |
|---|---|---|---|
| period | No | Time period (default: week) | |
| session_id | No | Specific session ID to analyze |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark it read-only and idempotent; the description adds beyond that by noting it 'Parses Claude Code JSONL logs automatically' and specifies the exact return shape. It does not over-disclose or contradict annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short, information-dense sentences. The main action is first, then a behavioral note, then alternative routing, then the return format. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a read-only analytics tool with no output schema, the description compensates by stating the return JSON structure and parsing behavior. It lacks explicit prerequisites like 'log files must exist,' but the parameters and routing are sufficiently covered.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, with both 'period' and 'session_id' described in the schema. The description does not add parameter-specific semantics, but the schema fully covers it, so a baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Analyze') with a clear resource ('AI agent session logs') and enumerates the concrete outputs (token usage, cost breakdown, top files, models). It also distinguishes itself from sibling tools by naming what it is not for.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use alternatives: 'For waste detection use get_optimization_report; for cost trends use get_usage_trends.' This gives clear routing guidance and prevents mis-selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_session_statsARead-onlyIdempotent
Token savings stats for this session: per-tool call counts, estimated token savings, reduction percentage, dedup savings, and per-tool latency (p50/p95/max/error_rate). Read-only. Returns JSON: { session: { ..., latency_per_tool }, cumulative, dedup_saved_tokens, report }.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover read-only, idempotent, and non-destructive behavior, so the description is not burdened with those disclosures. It adds valuable behavioral detail by specifying the returned JSON structure—session with latency_per_tool, cumulative, dedup_saved_tokens, and report—which is especially useful given there is no output schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that leads with the purpose and follows with a compact return layout. It contains no filler, though the inline JSON example is somewhat dense and could be slightly clearer as a structured list.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a no-parameter, read-only metrics tool, the description covers purpose, scope, and output shape adequately. It does not clarify how 'session' is defined or when a sibling analytics tool would be more appropriate, leaving a small but meaningful gap for an agent selecting among similar tools.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There are zero parameters and the schema coverage is trivially complete, so the baseline is 4. The description reinforces that the tool is parameterless by scoping everything to 'this session,' but it does not need to add more parameter-level meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description identifies the resource ('this session') and the concrete purpose ('Token savings stats'), and enumerates the metrics returned: per-tool call counts, token savings, reduction percentage, dedup savings, and latency percentiles/error rate. It is clear and specific, but it does not explicitly distinguish itself from overlapping siblings like get_session_analytics.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The phrase 'for this session' implies that the tool should be used when an agent wants token-savings metrics for the current session, which provides some usage context. However, the description gives no explicit guidance about when not to use it or how it compares to alternatives such as get_session_analytics, get_optimization_report, or get_real_savings.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_symbolARead-onlyIdempotent
Look up a symbol by symbol_id or FQN and return its source code. Use instead of Read when you need one specific function/class/method — returns only the symbol, not the whole file. For multiple symbols at once, prefer get_context_bundle. Read-only. Returns JSON: { symbol_id, name, kind, fqn, signature, file, line_start, line_end, source }.
| Name | Required | Description | Default |
|---|---|---|---|
| fqn | No | The fully qualified name to look up | |
| max_lines | No | Truncate source to this many lines (omit for full source) | |
| symbol_id | No | The symbol_id to look up | |
| verify_against_git | No | Compare the indexed source against the current git HEAD slice; mismatches set `git_mismatch: true` in the response (index may be stale). Read-only. Silently skipped when git is unavailable or the file is untracked. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish readOnly, idempotent, and non-destructive behavior. The description adds useful context beyond annotations: it returns only the symbol rather than the whole file and provides the response shape. It does not mention the verify_against_git comparison behavior or git_mismatch in the return payload, though that is documented in the parameter schema, so this is a minor omission rather than a contradiction.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded with the core purpose. It packs usage guidance, behavioral scope, read-only information, and return shape into three sentences with no wasted words.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description provides a return shape despite the lack of an output schema and clearly routes between siblings. However, it does not explicitly state that exactly one of symbol_id or fqn is required, and the verify_against_git behavior only lives in the schema. These are real but minor gaps given the otherwise rich annotation and schema context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already explains each parameter. The description reinforces that either symbol_id or fqn can be used, but it does not add significant meaning beyond the schema for max_lines or verify_against_git. This matches the baseline for fully covered schema parameters.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'Look up a symbol by symbol_id or FQN and return its source code.' It clearly distinguishes the tool from Read and get_context_bundle, so an agent can immediately understand what it does and which sibling it is not.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit usage direction: 'Use instead of Read when you need one specific function/class/method' and 'For multiple symbols at once, prefer get_context_bundle.' This tells the agent when to select this tool and when to select a sibling.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_task_contextARead-onlyIdempotent
All-in-one context for starting a dev task: execution paths, tests, entry points, adapted by task type. Use as your FIRST call when beginning any new task — replaces manual chaining of search → get_symbol → Read. For narrower feature-code lookup use get_feature_context instead. Read-only. Returns JSON (default) or Markdown.
| Name | Required | Description | Default |
|---|---|---|---|
| task | Yes | Natural language description of the task | |
| focus | No | Context strategy: minimal (fast, essential only), broad (default, wide net), deep (follow full execution chains) | |
| detail_level | No | Output verbosity. "minimal" saves ~40-60% tokens (drops scores, fqn, signatures, summaries). Use to pick a candidate before get_symbol. Default: "default". | |
| token_budget | No | Max tokens (default 8000) | |
| include_tests | No | Include relevant test files (default true) | |
| output_format | No | "json" (default, structured fields) or "markdown" (single LLM-optimized document with code fences, ~15-20% cheaper). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true and idempotentHint=true, so the safety profile is covered. The description adds useful behavioral context beyond annotations: outputs are adapted by task type, it returns JSON or Markdown, and it consolidates what would otherwise require multiple tools. It doesn't go into response structure or cost details, but it meaningfully exceeds what annotations alone provide.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tightly packed sentences with zero filler. The most actionable instruction ('Use as your FIRST call') is front-loaded, and both the sibling differentiation and return format are stated in a single breath.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given six parameters, a rich annotation set, and no output schema, the description covers nearly everything an agent needs: purpose, timing, alternatives, content delivered, and output format. It would be complete with a hint about response fields beyond 'execution paths, tests, entry points', but the schema covers the parameters and the description covers usage context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all six parameters thoroughly in terms of enums, defaults, and token impact. The description only lightly references output_format ('Returns JSON (default) or Markdown') and the task type adaptation, adding marginal value beyond the schema's own param descriptions.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a concrete, specific purpose: 'All-in-one context for starting a dev task' and enumerates the included content (execution paths, tests, entry points). It explicitly differentiates from the sibling tool get_feature_context, making the intent unmistakable against the surrounding toolset.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides explicit when-to-use guidance: 'Use as your FIRST call when beginning any new task.' It even names the workflow it replaces ('search → get_symbol → Read') and points to the correct alternative for narrower lookups (get_feature_context). This is model behavior for routing an agent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_usage_trendsARead-onlyIdempotent
Daily token usage time-series: sessions, tokens, estimated cost, tool calls per day. For spotting cost spikes. Read-only. For detailed session breakdown use get_session_analytics instead. Returns JSON: { days, daily: [{ date, sessions, tokens, cost_usd, tool_calls }], totals }.
| Name | Required | Description | Default |
|---|---|---|---|
| days | No | Number of days to show (default: 30) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint, idempotentHint, and destructiveHint=false, so the read-only nature disclosed in the description is consistent but not new. The description adds value by revealing the exact return shape (JSON with days, daily array, totals) and the daily time-series granularity, which the annotations do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences with no fluff. The main purpose is front-loaded, the use case and alternative follow immediately, and the return format is compactly summarized. Every sentence earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The tool has only one optional parameter, no output schema, and low complexity. The description covers the purpose, the use case, the alternative tool, the safety profile is in annotations, and the return JSON shape is fully specified. There is nothing an agent needs to call this correctly that is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% with the parameter description already stating it is the number of days and the default of 30. The tool description adds no additional semantics for the 'days' parameter, but it does corroborate that the output contains a 'days' field, which is minimal extra context. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: produces a daily token usage time-series with sessions, tokens, estimated cost, and tool calls. The phrase 'For spotting cost spikes' clarifies the intended use, and it explicitly differentiates itself from get_session_analytics, so an agent can distinguish it from siblings without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit use case ('For spotting cost spikes') and names the alternative when a different need exists ('For detailed session breakdown use get_session_analytics instead'). This gives clear selection criteria with no inference required.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
invalidate_decisionAIdempotent
Mark a decision as no longer valid. The decision remains in the knowledge graph for historical queries but is excluded from active queries. Use when a decision is superseded or reversed. Mutates the decision store; idempotent. Returns JSON: { invalidated: { id, title, valid_until } }.
| Name | Required | Description | Default |
|---|---|---|---|
| id | Yes | Decision ID to invalidate | |
| valid_until | No | ISO timestamp when decision became invalid (default: now) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses key behavioral traits beyond annotations: the decision remains for historical queries, is excluded from active queries, mutates the decision store, and is idempotent. It also specifies the return shape. This adds meaningful context beyond the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four short sentences, each carrying useful information: the action, the historical/active distinction, the usage condition, and the return format. No filler or redundancy; the most important purpose is front-loaded.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the absence of an output schema, the returned JSON is explicitly stated. The mutation, idempotency, and retention behavior are all covered. The tool's effect on the knowledge graph and query behavior is clear, making it complete for the agent to invoke correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so both id and valid_until are documented there. The description itself does not add much parameter-level meaning beyond the schema, but it correctly implies valid_until is part of the return object. Baseline of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Mark as no longer valid'), names the resource ('decision'), and clearly defines the outcome: the decision stays in the knowledge graph for historical queries but is excluded from active queries. This differentiates it from siblings like remember_decision and query_decisions.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly states when to use the tool: 'Use when a decision is superseded or reversed.' It does not name alternative tools or give when-not-to-use guidance, but the usage context is clear enough for an agent to select it correctly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
load_toolsARead-onlyIdempotent
Load tools this session's preset deferred, by preset name and/or explicit tool names. Call with no arguments to list what is deferred. Emits notifications/tools/list_changed and returns the loaded tools' schemas, so they are usable even if your client ignores that notification (call them through batch). Returns JSON: { loaded, already_loaded, unknown, blocked, tools, hint }.
| Name | Required | Description | Default |
|---|---|---|---|
| tools | No | Explicit tool names to load. Unions with `preset` when both are given. | |
| preset | No | Preset whose members to load (minimal, standard, review, architecture, full). "full" loads everything deferred. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description discloses behavior beyond the annotations: it emits notifications/tools/list_changed, returns loaded tools' schemas, and remains usable even if the client ignores the notification (via batch). It also lists the exact JSON response fields, giving the agent a clear model of side effects and results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but efficient: it front-loads the core purpose, then covers invocation modes, side effects, and return shape in a few sentences. No sentence is filler, and the structure makes the most important information immediately visible.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite having no output schema, the description names all return fields and explains the notification behavior and batch fallback. Combined with the annotations signaling read-only/idempotent behavior, the description gives the agent enough context to invoke the tool correctly and interpret its response.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already describes both parameters with 100% coverage, including the union behavior between `tools` and `preset`. The description adds value by specifying the no-arguments listing behavior, which is a parameter-level semantic not captured by the schema alone.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('load') with a clear resource: deferred tools for the session's preset, by preset name and/or explicit names. It also distinguishes the no-argument listing mode, making the tool's purpose unambiguous and distinct from siblings like get_preset_info.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives clear usage context: call with no arguments to list deferred tools, or pass preset/tool names to load them. It does not explicitly mention alternatives or when-not-to-use conditions, but the provided context is sufficient for typical invocation.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mine_sessionsAIdempotent
Mine Claude Code / Claw Code session logs for architectural decisions, tech choices, bug root causes, and preferences. Strategies: "regex" (default, free, ~20-40% recall), "llm" (higher recall, costs tokens), "hybrid" (regex + LLM safety net). Skips already-mined sessions unless force=true. Mutates the decision store; idempotent. Returns JSON: { mined, decisions_extracted, sessions_processed, strategy?, llm_sessions?, llm_decisions_extracted? }.
| Name | Required | Description | Default |
|---|---|---|---|
| force | No | Re-mine already processed sessions (default: false) | |
| strategy | No | Extraction strategy: regex (default, free/fast/low recall), llm (AI provider, costs tokens, higher recall), hybrid (regex + LLM safety net). Falls back to regex with a warning if no AI provider is configured. | |
| project_root | No | Only mine sessions for this project path (default: all projects) | |
| min_confidence | No | Legacy reject floor — drops decisions below this. Superseded by reject_threshold; kept for back-compat. | |
| reject_threshold | No | Reject floor (default: config decisions.reject_threshold, fallback 0.45). Decisions in [reject_threshold, review_threshold) queue for review; below it, dropped. | |
| review_threshold | No | Auto-approve cutoff (default: config decisions.review_threshold, fallback 0.75). Decisions ≥ this enter the active graph immediately. | |
| incremental_cursor | No | Per-call override for `memory.mining.incrementalCursor`. true (default) reuses byte-offset cursors for appended turns; false falls back to legacy mined/unmined semantics. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description explicitly states "Mutates the decision store; idempotent," and "Skips already-mined sessions unless force=true." This adds meaningful behavioral context beyond the annotations, specifying exactly what side effect occurs, the idempotency guarantee, and the skip behavior. It aligns with idempotentHint=true and readOnlyHint=false, with no contradictions.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and dense: two sentences carrying purpose, strategy tradeoffs, behavioral notes, and return shape. It front-loads the core purpose, uses structured lists for strategies, and contains no filler. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a tool with 7 optional parameters, no output schema, and moderate complexity, the description supplies the return JSON structure, strategy cost/recall tradeoffs, mutation behavior, and skip logic. Combined with 100% schema coverage, an agent has everything needed to select and invoke the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, and the schema already documents all 7 parameters with detailed descriptions, including enum choices, thresholds, and the incremental_cursor override. The description merely summarizes strategy and force, adding no new semantic information beyond what the schema already provides. Baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: "Mine Claude Code / Claw Code session logs for architectural decisions, tech choices, bug root causes, and preferences." This clearly distinguishes it from sibling read/query tools like search and query_decisions by stating it processes session logs and mutates the decision store. The purpose is unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives concrete guidance on strategy selection (regex vs llm vs hybrid) and explains the skip-already-mined behavior with force=true. However, it never explicitly names alternatives or states when NOT to use this tool (e.g., "use query_decisions instead to read stored decisions"). Context is clear but exclusions are absent.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
plan_turnARead-onlyIdempotent
Opening-move router for new tasks. Combines BM25/PageRank search + session journal (negative evidence + focus signals) + framework-aware insertion-point suggestions + change-risk + turn-budget advisor into ONE call. Returns verdict (exists/partial/missing/ambiguous), confidence, ranked targets with provenance, scaffold hints when missing, and recommended next tool calls. Call this FIRST on a new task to break the empty-result hallucination chain. Read-only. For broader task context with source code use get_task_context instead. Returns JSON: { verdict, confidence, targets, scaffoldHints, nextSteps }.
| Name | Required | Description | Default |
|---|---|---|---|
| task | Yes | Natural-language task description (e.g. "add a webhook endpoint for stripe payments") | |
| intent | No | Optional intent hint; auto-classified from task if omitted | |
| skip_risk | No | Skip change-risk assessment for the top target (default false) | |
| max_targets | No | Cap on returned targets (default 5) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already mark this read-only, idempotent, and non-destructive, and the description reinforces this with 'Read-only.' It adds behavioral context by explaining the tool's combined search/journal/risk mechanism and its role in preventing empty-result hallucinations, going beyond what annotations alone convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is information-dense and mostly front-loaded, starting with the primary purpose and usage call-to-action. The long enumeration of combined capabilities and the slight redundancy between 'Returns...' and 'Returns JSON: {...}' keep it from being perfectly concise, but every sentence contributes needed context.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite the absence of an output schema, the description states the JSON shape, the verdict values, the provenance of ranked targets, scaffold hints, and recommended next steps. Combined with the explicit usage instruction and alternative tool pointer, an agent has what it needs to invoke the tool correctly on a new task.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, and each parameter already has a clear description with defaults and constraints. The tool description does not meaningfully add parameter-level guidance, so it meets the baseline but does not exceed it.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Opening-move router for new tasks' and lists a concrete deliverable set: verdict, confidence, ranked targets with provenance, scaffold hints, and recommended next tool calls. It also explicitly distinguishes itself from get_task_context, so an agent can select it correctly among siblings.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description states exactly when to use it: 'Call this FIRST on a new task to break the empty-result hallucination chain.' It also names the alternative, get_task_context, and the condition for preferring that instead ('broader task context with source code'), giving clear, actionable routing guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
query_decisionsARead-onlyIdempotent
Query the decision knowledge graph. Filter by type, subproject, code symbol, file path, tag, or time — answers "why was this architecture chosen?" with the actual decision record. Defaults to approved decisions; use include_pending or review_status for other tiers. Read-only. Returns JSON: { decisions: [{ id, title, type, content, tags, review_status, cluster_ids? }], clusters_summary?, total_results }. Supports output_format: "toon".
| Name | Required | Description | Default |
|---|---|---|---|
| tag | No | Filter by tag | |
| type | No | Filter by decision type | |
| as_of | No | Only decisions active at this ISO timestamp | |
| limit | No | Max results (default: 50) | |
| search | No | Full-text search query (FTS5 with porter stemming) | |
| verify | No | Verify linked code is still fresh (default true); stale rows are flagged `stale: true`. false skips the check. | |
| order_by | No | Result ordering: "recency" (default), "created_at", or "heat" (popular + fresh; falls back to recency when disabled) | |
| file_path | No | Filter by linked file path | |
| symbol_id | No | Filter by linked symbol FQN | |
| git_branch | No | Branch filter: "current" (default), "all", or a branch name | |
| index_only | No | Omit full `content` (default false) — pick ids cheaply, then pull content with `get_decision` | |
| service_name | No | Filter by subproject name (e.g., "auth-api") | |
| verification | No | Filter by verification verdict (implies verify=true): "stale" or "ok" | |
| output_format | No | Output format: "json" (default), "markdown" (tool-specific), or "toon" (30-60% fewer tokens on tabular data). | |
| review_status | No | Restrict to one review tier ("pending" = review queue; overrides default) | |
| include_pending | No | Also return pending review-queue decisions (default: approved only) | |
| include_invalidated | No | Include invalidated decisions (default: false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The annotations already declare read-only, idempotent, and non-destructive, so the bar is lower. The description adds valuable behavior: defaults to approved decisions, returns JSON with specific fields, supports the 'toon' output format for token savings, and explains the verify flag's effect on stale rows. It could further clarify the interpretation of 'active' for as_of and the fallback behavior of 'heat', but these are minor.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single paragraph that front-loads the purpose and key filters, then covers defaults, output format, and read-only nature. It is fairly concise, but the field list in the output JSON is a bit dense and could be trimmed without losing value. Still, it earns its sentences with concrete details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a complex tool with 17 parameters, the description effectively summarizes the major filter categories, defaults, output format, and read-only nature. The schema covers individual parameter semantics, so the description need not repeat every detail. The mention of the toon format is a nice touch, given the presence of output_format. Complete for the intended use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The schema has 100% coverage, so the baseline is 3. The description adds context beyond the schema: it explains that decisions are part of a knowledge graph, mentions the 'toon' output format's benefit, and clarifies that 'verify' defaults to true (schema only lists the flag). This goes slightly beyond the schema, justifying a 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool queries the decision knowledge graph, lists the filter dimensions (type, subproject, code symbol, file path, tag, time), and explains the core use case (answering 'why was this architecture chosen?'). It is specific about the resource and distinguishes it from siblings like search, get_symbol, and remember_decision.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description implies when to use it (querying decision records) and mentions defaults (approved decisions, include_pending/review_status for other tiers), but it does not explicitly state when not to use it or name alternative tools for similar purposes. The 'read-only' note and the output format hint provide context, but exclusions are missing.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
register_editAIdempotent
Notify trace-mcp that a file was edited. Reindexes the single file and invalidates search caches. Call after Edit/Write to keep index fresh — much lighter than full reindex. Also flags duplicate symbols — if _duplication_warnings appears, you may be recreating existing logic; review them. Each one is reported once per file, not on every edit; check_duplication re-asks. Mutates the index; idempotent. Returns JSON: { status, file, totalFiles, indexed, _duplication_warnings? }.
| Name | Required | Description | Default |
|---|---|---|---|
| file_path | Yes | Relative path to the edited file |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already carry idempotentHint=true, readOnlyHint=false, and destructiveHint=false; the description confirms these ('Mutates the index; idempotent') without contradicting them. It adds real value beyond annotations: search-cache invalidation, duplicate-symbol flagging, the once-per-file reporting constraint, and the exact JSON return shape. Minor gap: no mention of error behavior when the file path is invalid.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core purpose, then usage guidance, then the duplication caveat, then the return format. Each sentence earns its place and the information is ordered by importance. Slightly dense, but not padded — no redundancy with the schema or annotations.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Complete for a single-parameter mutation tool: it covers what happens (reindex, cache invalidation), when to use it (after Edit/Write, lighter than full reindex), its idempotency, a caveat (duplication warnings, reported once per file), and the return structure in the absence of an output schema. Nothing an agent needs to call it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100% — the single parameter file_path is fully documented as 'Relative path to the edited file.' The description references the file only implicitly through the return JSON ({ file }), adding no semantic detail beyond the schema. Baseline 3 is appropriate since the schema carries the parameter meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Notify trace-mcp that a file was edited'), and precisely scopes what happens: reindex the single file and invalidate search caches. It clearly differentiates from the read-oriented siblings (get_index_health, get_symbol, search) by being a mutation, and positions itself as the lightweight single-file counterpart to a full reindex.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use guidance: 'Call after Edit/Write to keep index fresh — much lighter than full reindex.' It names the alternative (full reindex) and the deciding condition. It also routes a follow-up question to a sibling ('check_duplication re-asks'), so an agent knows which tool to pick next.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
remember_decisionA
Live agent write into the decision knowledge graph. Confidence-scores the input and routes it through the memoir review queue: high-confidence rows enter the active graph immediately, mid-confidence rows queue for human approval, low-confidence rows are dropped without persistence. Per-session dedup + rate-limit. Use during a session to capture decisions in real time. For manual high-confidence writes use add_decision; for post-hoc extraction from session logs use mine_sessions. Returns JSON: { id, review_status, confidence, deduplicated? }.
| Name | Required | Description | Default |
|---|---|---|---|
| tags | No | Tags for categorization (e.g., ["auth", "security"]) | |
| type | Yes | Decision type | |
| title | Yes | Short summary of the decision | |
| content | Yes | Full decision text — reasoning, context, tradeoffs | |
| file_path | No | File path this decision is about | |
| symbol_id | No | Symbol FQN this decision is about (e.g., "src/auth/provider.ts::AuthProvider#class") | |
| git_branch | No | Git branch this decision belongs to. Omit to auto-detect, or pass null to make it branch-agnostic. | |
| session_id | No | Session identifier for dedup/rate-limit (default: "_default") | |
| service_name | No | Subproject name this decision is about (e.g., "auth-api", "user-service") |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations only cover the safety profile (readOnlyHint=false, idempotentHint=false, destructiveHint=false). The description goes well beyond them: confidence-based routing with three explicit outcomes (immediate graph entry, human review queue, silent drop without persistence), plus per-session dedup and rate limiting. This is exactly the behavioral detail an agent needs before writing.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with the core action, then routing behavior, then constraints, then alternatives, then return shape. Dense but every sentence carries distinct information with no repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 9-parameter mutation tool with no output schema, the description supplies the missing return shape ({ id, review_status, confidence, deduplicated? }), the persistence semantics, and the dedup/rate-limit constraints. An agent has everything needed to call it correctly and predict the outcome.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents all 9 parameters including session_id's dedup/rate-limit role. The description restates session-level dedup but adds no new syntax or meaning for tags, file_path, symbol_id, git_branch, or service_name, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb+resource ("write into the decision knowledge graph") and immediately scopes it as a live, in-session write. It explicitly distinguishes itself from the two nearest siblings, add_decision and mine_sessions, so an agent can route without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ("during a session to capture decisions in real time") plus named alternatives for the two adjacent cases: add_decision for manual high-confidence writes, mine_sessions for post-hoc log extraction. Nothing is left to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
searchARead-onlyIdempotent
Search symbols by name, kind, or text. Use instead of Grep for functions, classes, methods, variables. For raw text/comment search use search_text; for references to a known symbol use find_usages. Read-only. Returns JSON: { items: [{ symbol_id, name, kind, fqn, signature, file, line, score }], total, search_mode } — mode-specific shape when mode!=single. Supports output_format: "toon".
| Name | Required | Description | Default |
|---|---|---|---|
| kind | No | Filter by symbol kind (class, method, function, etc.) | |
| mode | No | single (default): top-K. tiered: high/medium/low buckets. drill: scoped to drill_from. flat: raw FTS. get: exact lookup. Omit to auto-pick. | |
| fuzzy | No | Typo-tolerant search. Auto-enabled when exact search returns 0 results. | |
| limit | No | Max results (default 20) | |
| query | Yes | Search query | |
| fusion | No | Enable Signal Fusion — WRR ranking across lexical, structural, similarity, and identity channels. Weights come from `tune_weights`. | |
| offset | No | Offset for pagination | |
| extends | No | Filter to classes/interfaces extending this name | |
| language | No | Filter by language | |
| semantic | No | auto (default): hybrid if AI available. on: force hybrid. off: lexical-only. only: pure vector. Non-"off" needs an AI provider. | |
| decorator | No | Filter to symbols carrying this decorator/annotation/attribute | |
| retriever | No | Run one named retrieval algorithm instead of the mode dispatcher. Ignores mode/filters/fuzzy/fusion. | |
| drill_from | No | [mode="drill"] File path or symbol_id to restrict results to. | |
| implements | No | Filter to classes implementing this interface | |
| detail_level | No | Output verbosity. "minimal" saves ~40-60% tokens (drops scores, fqn, signatures, summaries). Use to pick a candidate before get_symbol. Default: "default". | |
| file_pattern | No | Filter by file path pattern | |
| output_format | No | "json" (default) or "toon" (lossless, 30-60% fewer tokens). "markdown" behaves as json here. | |
| fuzzy_threshold | No | [fuzzy] Min trigram similarity (default 0.3) | |
| semantic_weight | No | [semantic] 0 = lexical only, 0.5 = balanced (default), 1 = vector only. | |
| max_edit_distance | No | [fuzzy] Max edit distance (default 3) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already convey read-only/idempotent/non-destructive traits. The description adds value by disclosing the JSON return shape, mode-specific shape when mode!=single, and the toon output format. This goes beyond annotation coverage and helps the agent understand what to expect.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is compact and front-loaded: purpose, exclusions, and return shape in three sentences. Minor redundancy exists, such as restating 'Read-only' despite annotations and noting output_format already in the schema, but the overall size is appropriate.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
This is a complex 20-parameter tool with no output schema, but the schema itself is richly documented and annotations cover safety. The description supplies the missing output contract and key routing guidance, leaving only deeper mode-selection advice to the schema's own explanations.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds minimal per-parameter meaning beyond the schema—mainly highlighting output_format toon—but the schema already documents all parameters thoroughly.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Description states a specific verb and resource: 'Search symbols by name, kind, or text.' It explicitly distinguishes the tool from Grep and names the sibling alternatives search_text and find_usages, so an agent can tell them apart immediately.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Description gives explicit when-to-use guidance: use instead of Grep for functions, classes, methods, variables; use search_text for raw text/comment search; use find_usages for references to a known symbol. This clearly routes the agent to the correct alternative.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_textARead-onlyIdempotent
Full-text search across all indexed files. Supports regex, glob file patterns, language filter. Use for finding strings, comments, TODOs, config values, error messages — anything not captured as a symbol. Skips .env files: secret values are never searched or returned, use get_env_vars for env keys. For symbol search (functions, classes) use search instead. Read-only. Returns JSON: { files: [{ file, language, hits: [{ line, column, match, context }] }], total_matches } — hits grouped per file, so a long path is paid once. Pass grouping: "flat" for the ungrouped matches[] shape.
| Name | Required | Description | Default |
|---|---|---|---|
| query | Yes | Search string or regex pattern | |
| grouping | No | Payload shape. "by_file" (default) groups hits per file, so a long path is paid once; "flat" is a single matches[] array. | by_file |
| is_regex | No | Treat query as regex (default false) | |
| language | No | Filter by language (e.g. "typescript", "python") | |
| timeout_ms | No | Wall-clock budget in ms — caps a catastrophic-backtracking regex. Default 2000; 0 disables. | |
| max_results | No | Max matches to return (default 50) | |
| file_pattern | No | Glob filter, e.g. "src/**/*.ts" | |
| context_lines | No | Lines of context before/after each match (default 0 — set higher if you need surrounding code) | |
| case_sensitive | No | Case-sensitive search (default false) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare read-only and idempotent, but the description adds substantial behavioral detail: it skips .env files (secrets never searched/returned), describes the return JSON shape and grouping behavior, and mentions the timeout parameter's purpose (caps catastrophic-backtracking regex). This goes beyond the annotations and provides context an agent needs.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Although the description is lengthier than average, every sentence earns its place. It front-loads the purpose, then usage, then behavioral and return details, with no redundancy. The structure is logical and each sentence adds substantive information, making it efficient despite its length.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (9 parameters, no output schema), the description is remarkably complete. It covers the return JSON shape (since no output schema exists), usage scenarios, exclusions, alternatives, and even explains grouping and timeout behavior. An agent can confidently invoke this tool correctly without additional context.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds meaning beyond the schema by explaining the 'grouping' parameter's trade-off ('hits grouped per file, so a long path is paid once') and the 'timeout_ms' rationale (caps catastrophic-backtracking regex). While not all parameters get extra semantics, the added context is valuable, earning a 4.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource ('Full-text search across all indexed files') and explicitly differentiates from the sibling 'search' by stating 'For symbol search (functions, classes) use search instead.' It also enumerates the types of content it targets (strings, comments, TODOs, config values, error messages), making the purpose unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives explicit usage guidance: 'Use for finding strings, comments, TODOs, config values, error messages — anything not captured as a symbol.' It also specifies when not to use it (symbol search) and directs the agent to 'get_env_vars' for environment keys, covering exclusions and alternatives clearly.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
suggest_queriesARead-onlyIdempotent
Onboarding helper: shows top imported files, most connected symbols (PageRank), language stats, and example tool calls. Call this first when exploring an unfamiliar project. For a structured project map use get_project_map instead. Read-only. Returns JSON: { topFiles, topSymbols, languageStats, exampleQueries }.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
The description confirms read-only behavior, consistent with the readOnlyHint/idempotentHint annotations, and adds a concrete response shape even without an output schema. It does not mention edge cases like rate limits or authentication, but these are less critical for a read-only onboarding helper.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Every sentence earns its place: the onboarder identity is front-loaded, usage timing is clear, the alternative is named, and the JSON return shape is compactly listed. There is no filler or redundant schema repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a zero-parameter read-only tool with no output schema, the description covers purpose, output fields, usage timing, and the relevant sibling. Nothing an agent needs to call this correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool has zero parameters, so schema coverage is trivially 100% and the baseline is 4. The description adds no parameter details because there are none; instead it clarifies what the no-input call returns, which is sufficient.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('shows') and names the resource types it exposes: top imported files, connected symbols, language stats, and example tool calls. It also explicitly differentiates from get_project_map, so an agent can distinguish it from siblings without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states the exact condition for use — 'Call this first when exploring an unfamiliar project' — and names the alternative, get_project_map, for a structured project map. This gives clear when-to-use and when-not-to-use guidance with an explicit replacement.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v3.31.3- Changed
query_decisions8 fields changed- changed
Input schema / properties / git_branch / descriptionPrevious value: -"Branch filter: \"current\" (default) = current branch + branch-agnostic; \"all\" = every branch; any other value = that branch + branch-agnostic."New value: +"Branch filter: \"current\" (default), \"all\", or a branch name" - changed
Input schema / properties / include_pending / descriptionPrevious value: -"Also return decisions in the review queue (review_status=\"pending\"). Default: false — only auto-approved and approved rows are returned."New value: +"Also return pending review-queue decisions (default: approved only)" - changed
Input schema / properties / index_only / descriptionPrevious value: -"Progressive disclosure (default false). true omits full `content` — just id, title, type, anchors, tags, ~1-line `summary`. Pick ids cheaply, then pull full content with `get_decision`."New value: +"Omit full `content` (default false) — pick ids cheaply, then pull content with `get_decision`" - changed
Input schema / properties / order_by / descriptionPrevious value: -"Result ordering: \"recency\" (default, valid_from DESC), \"created_at\" DESC, or \"heat\" (time-decay favoring frequently-recalled + fresh; degrades to recency if disabled in config)."New value: +"Result ordering: \"recency\" (default), \"created_at\", or \"heat\" (popular + fresh; falls back to recency when disabled)" - changed
Input schema / properties / output_format / descriptionPrevious value: -"Output format. \"json\" (default), \"markdown\" (LLM-friendly fenced markdown, tool-specific), or \"toon\" (Token-Oriented Object Notation — 30-60% fewer tokens on tabular data, lossless)."New value: +"Output format: \"json\" (default), \"markdown\" (tool-specific), or \"toon\" (30-60% fewer tokens on tabular data)." - changed
Input schema / properties / review_status / descriptionPrevious value: -"Restrict to a single review tier (overrides default + include_pending). Use \"pending\" to fetch the review queue."New value: +"Restrict to one review tier (\"pending\" = review queue; overrides default)" - changed
Input schema / properties / verification / descriptionPrevious value: -"Filter by verification verdict (implies verify=true). \"stale\" = any flagged row; \"ok\" = verified-fresh only. Omit to return all rows annotated in place."New value: +"Filter by verification verdict (implies verify=true): \"stale\" or \"ok\"" - changed
Input schema / properties / verify / descriptionPrevious value: -"Staleness verification (default true). Checks each `symbol_id`-linked decision against the live index + git history; deleted/renamed/materially-changed code is flagged `verification` + `stale: true`. false skips the check."New value: +"Verify linked code is still fresh (default true); stale rows are flagged `stale: true`. false skips the check."
- Changed
search4 fields changed- changed
Input schema / properties / fusion / descriptionPrevious value: -"Enable Signal Fusion — multi-channel WRR ranking across lexical (BM25), structural (PageRank), similarity (embeddings), and identity match. Weights come from `tune_weights`."New value: +"Enable Signal Fusion — WRR ranking across lexical, structural, similarity, and identity channels. Weights come from `tune_weights`." - changed
Input schema / properties / mode / descriptionPrevious value: -"single (default): top-K. tiered: high/medium/low buckets. drill: scoped to drill_from. flat: raw FTS, no PageRank. get: exact lookup. Omit to auto-pick."New value: +"single (default): top-K. tiered: high/medium/low buckets. drill: scoped to drill_from. flat: raw FTS. get: exact lookup. Omit to auto-pick." - changed
Input schema / properties / retriever / descriptionPrevious value: -"Run one named retrieval algorithm instead of the mode dispatcher. Ignores mode/filters/fuzzy/fusion; returns { retriever, items, total }."New value: +"Run one named retrieval algorithm instead of the mode dispatcher. Ignores mode/filters/fuzzy/fusion." - changed
Input schema / properties / semantic / descriptionPrevious value: -"auto (default): hybrid if AI available. on: force hybrid. off: lexical-only. only: pure vector. Non-\"off\" needs an AI provider + one embed_repo run."New value: +"auto (default): hybrid if AI available. on: force hybrid. off: lexical-only. only: pure vector. Non-\"off\" needs an AI provider."
7 tool updates
v3.31.0- Changed
find_usages1 field changed- changed
Input schema / properties / detail_level / descriptionPrevious value: -"Output verbosity. \"minimal\" returns ~40-60% fewer tokens (drops scores, fqn, signatures, summaries — keeps name/file/line). Use when you only need to pick a candidate before drilling in with get_symbol. Default: \"default\"."New value: +"Output verbosity. \"minimal\" saves ~40-60% tokens (drops scores, fqn, signatures, summaries). Use to pick a candidate before get_symbol. Default: \"default\"."
- Changed
get_change_impact2 fields changed- added
Input schema / properties / bundleAdded value: +{ + "description": "Handle; @N = page N.", + "type": "string" +} - added
Input schema / properties / compactAdded value: +{ + "description": "Paged recall (opt-in).", + "type": "boolean" +}
- Changed
get_diagnostics1 field changed- added
Input schema / properties / reduce_outputAdded value: +{ + "description": "Shrink long output to a receipt", + "type": "boolean" +}
- Changed
get_feature_context1 field changed- changed
Input schema / properties / detail_level / descriptionPrevious value: -"Output verbosity. \"minimal\" returns ~40-60% fewer tokens (drops scores, fqn, signatures, summaries — keeps name/file/line). Use when you only need to pick a candidate before drilling in with get_symbol. Default: \"default\"."New value: +"Output verbosity. \"minimal\" saves ~40-60% tokens (drops scores, fqn, signatures, summaries). Use to pick a candidate before get_symbol. Default: \"default\"."
- Changed
get_outline1 field changed- changed
Input schema / properties / detail_level / descriptionPrevious value: -"Output verbosity. \"minimal\" returns ~40-60% fewer tokens (drops scores, fqn, signatures, summaries — keeps name/file/line). Use when you only need to pick a candidate before drilling in with get_symbol. Default: \"default\"."New value: +"Output verbosity. \"minimal\" saves ~40-60% tokens (drops scores, fqn, signatures, summaries). Use to pick a candidate before get_symbol. Default: \"default\"."
- Changed
get_task_context1 field changed- changed
Input schema / properties / detail_level / descriptionPrevious value: -"Output verbosity. \"minimal\" returns ~40-60% fewer tokens (drops scores, fqn, signatures, summaries — keeps name/file/line). Use when you only need to pick a candidate before drilling in with get_symbol. Default: \"default\"."New value: +"Output verbosity. \"minimal\" saves ~40-60% tokens (drops scores, fqn, signatures, summaries). Use to pick a candidate before get_symbol. Default: \"default\"."
- Changed
search1 field changed- changed
Input schema / properties / detail_level / descriptionPrevious value: -"Output verbosity. \"minimal\" returns ~40-60% fewer tokens (drops scores, fqn, signatures, summaries — keeps name/file/line). Use when you only need to pick a candidate before drilling in with get_symbol. Default: \"default\"."New value: +"Output verbosity. \"minimal\" saves ~40-60% tokens (drops scores, fqn, signatures, summaries). Use to pick a candidate before get_symbol. Default: \"default\"."
1 tool update
v3.25.0- Added
get_diagnostics
11 tool updates
v3.22.0- Changed
find_usages2 fields changed- added
Input schema / properties / limitAdded value: +{ + "description": "Max references returned (default 50).", + "maximum": 1000, + "minimum": 1, + "type": "integer" +} - removed
Input schema / requiredRemoved value: -[ - "symbol_id", - "fqn", - "file_path" -]
- Changed
get_call_graph1 field changed- removed
Input schema / requiredRemoved value: -[ - "symbol_id", - "fqn" -]
- Changed
get_change_impact1 field changed- removed
Input schema / requiredRemoved value: -[ - "file_path", - "symbol_id" -]
- Changed
get_context_bundle1 field changed- removed
Input schema / requiredRemoved value: -[ - "symbol_id", - "fqn" -]
- Changed
get_session_analytics1 field changed- removed
Input schema / requiredRemoved value: -[ - "session_id" -]
- Changed
get_symbol1 field changed- removed
Input schema / requiredRemoved value: -[ - "symbol_id", - "fqn" -]
- Changed
load_tools1 field changed- removed
Input schema / requiredRemoved value: -[ - "preset" -]
- Changed
query_decisions2 fields changed- changed
Input schema / properties / output_format / descriptionPrevious value: -"Output format. \"json\" (default) returns JSON, \"markdown\" returns LLM-friendly fenced markdown (tool-specific), \"toon\" returns Token-Oriented Object Notation — 30-60% fewer tokens on tabular data, fully lossless."New value: +"Output format. \"json\" (default), \"markdown\" (LLM-friendly fenced markdown, tool-specific), or \"toon\" (Token-Oriented Object Notation — 30-60% fewer tokens on tabular data, lossless)." - removed
Input schema / requiredRemoved value: -[ - "symbol_id", - "file_path", - "tag", - "as_of" -]
- Changed
remember_decision1 field changed- changed
Input schema / requiredPrevious value: -[ - "title", - "content", - "type", - "file_path" -]New value: +[ + "title", + "content", + "type" +]
- Changed
search1 field changed- changed
Input schema / requiredPrevious value: -[ - "query", - "language", - "file_pattern" -]New value: +[ + "query" +]
- Changed
search_text3 fields changed- changed
Input schema / properties / grouping / defaultPrevious value: -"flat"New value: +"by_file" - changed
Input schema / properties / grouping / descriptionPrevious value: -"Payload shape. \"flat\" (default) is a single matches[] array; \"by_file\" groups hits per file — saves tokens on long paths with many hits."New value: +"Payload shape. \"by_file\" (default) groups hits per file, so a long path is paid once; \"flat\" is a single matches[] array." - changed
Input schema / requiredPrevious value: -[ - "query", - "file_pattern" -]New value: +[ + "query" +]
56 tool updates
v3.3.0- Removed
apply_codemod - Removed
assess_change_risk - Changed
batch1 field changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#"
- Removed
check_duplication - Removed
check_quality_gates - Removed
check_rename - Removed
detect_antipatterns - Changed
find_usages1 field changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#"
- Changed
get_call_graph1 field changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#"
- Changed
get_change_impact1 field changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#"
- Removed
get_changed_symbols - Removed
get_circular_imports - Removed
get_complexity_report - Removed
get_complexity_trend - Changed
get_context_bundle1 field changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#"
- Removed
get_control_flow - Removed
get_coupling - Removed
get_coupling_trend - Changed
get_coverage_report1 field changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#"
- Removed
get_dead_code - Removed
get_dead_exports - Removed
get_env_vars - Changed
get_feature_context3 fields changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#" - added
Input schema / properties / detail_levelAdded value: +{ + "description": "Output verbosity. \"minimal\" returns ~40-60% fewer tokens (drops scores, fqn, signatures, summaries — keeps name/file/line). Use when you only need to pick a candidate before drilling in with get_symbol. Default: \"default\".", + "enum": [ + "minimal", + "default", + "full" + ], + "type": "string" +} - changed
Input schema / properties / output_format / descriptionPrevious value: -"Output format. \"json\" (default) returns structured items; \"markdown\" returns LLM-friendly fenced code blocks (~15-20% token savings, easier for the model to read); \"toon\" returns Token-Oriented Object Notation — 30-60% fewer tokens, lossless."New value: +"\"json\" (default, structured items), \"markdown\" (fenced code blocks, ~15-20% cheaper), or \"toon\" (lossless, 30-60% fewer tokens)."
- Removed
get_implementations - Changed
get_index_health1 field changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#"
- Changed
get_optimization_report1 field changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#"
- Changed
get_outline3 fields changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#" - changed
Input schema / properties / nested / descriptionPrevious value: -"When true, walks the body of each top-level symbol whose LOC exceeds min_loc_for_nesting and emits inner function-like declarations as additional rows carrying `parentId` + `depth`. Default false — fully backward compatible."New value: +"Walk the body of each top-level symbol past min_loc_for_nesting and emit inner declarations as extra rows carrying `parentId` + `depth`. Default false." - changed
Input schema / properties / output_format / descriptionPrevious value: -"Output format. \"json\" (default) returns JSON; \"toon\" returns Token-Oriented Object Notation — 30-60% fewer tokens, lossless. \"markdown\" is unsupported here and behaves as json."New value: +"\"json\" (default) or \"toon\" (lossless, 30-60% fewer tokens). \"markdown\" is unsupported here and behaves as json."
- Changed
get_preset_info1 field changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#"
- Changed
get_project_map1 field changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#"
- Changed
get_real_savings1 field changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#"
- Removed
get_related_symbols - Changed
get_session_analytics1 field changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#"
- Removed
get_session_resume - Changed
get_session_stats1 field changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#"
- Changed
get_symbol2 fields changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#" - changed
Input schema / properties / verify_against_git / descriptionPrevious value: -"When true, compare the indexed source against the current git HEAD slice for that file and line range. If they differ, the response includes `git_mismatch: true` indicating the index may be stale. Read-only — never writes. Silently skipped when git is unavailable or the file is not tracked."New value: +"Compare the indexed source against the current git HEAD slice; mismatches set `git_mismatch: true` in the response (index may be stale). Read-only. Silently skipped when git is unavailable or the file is untracked."
- Removed
get_symbol_complexity_trend - Changed
get_task_context3 fields changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#" - added
Input schema / properties / detail_levelAdded value: +{ + "description": "Output verbosity. \"minimal\" returns ~40-60% fewer tokens (drops scores, fqn, signatures, summaries — keeps name/file/line). Use when you only need to pick a candidate before drilling in with get_symbol. Default: \"default\".", + "enum": [ + "minimal", + "default", + "full" + ], + "type": "string" +} - changed
Input schema / properties / output_format / descriptionPrevious value: -"Output format. \"json\" (default) returns structured fields; \"markdown\" returns a single LLM-optimized document with code fences (~15-20% token savings)."New value: +"\"json\" (default, structured fields) or \"markdown\" (single LLM-optimized document with code fences, ~15-20% cheaper)."
- Removed
get_tech_debt - Removed
get_tests_for - Changed
get_usage_trends1 field changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#"
- Removed
get_workspace_map - Changed
invalidate_decision1 field changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#"
- Added
load_tools - Changed
mine_sessions5 fields changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#" - changed
Input schema / properties / incremental_cursor / descriptionPrevious value: -"Per-call override for `memory.mining.incrementalCursor`. When true (default), reuse byte-offset cursors so appended turns get re-processed; when false, fall back to legacy binary mined/unmined semantics."New value: +"Per-call override for `memory.mining.incrementalCursor`. true (default) reuses byte-offset cursors for appended turns; false falls back to legacy mined/unmined semantics." - changed
Input schema / properties / reject_threshold / descriptionPrevious value: -"Memoir reject floor (default: decisions.reject_threshold from config, fallback 0.45). Decisions in [reject_threshold, review_threshold) go into the review queue; below reject_threshold they are dropped."New value: +"Reject floor (default: config decisions.reject_threshold, fallback 0.45). Decisions in [reject_threshold, review_threshold) queue for review; below it, dropped." - changed
Input schema / properties / review_threshold / descriptionPrevious value: -"Memoir auto-approve cutoff (default: decisions.review_threshold from config, fallback 0.75). Decisions ≥ this enter the active knowledge graph immediately."New value: +"Auto-approve cutoff (default: config decisions.review_threshold, fallback 0.75). Decisions ≥ this enter the active graph immediately." - changed
Input schema / properties / strategy / descriptionPrevious value: -"Extraction strategy. regex (default): free, fast, low recall. llm: uses AI provider, costs tokens, higher recall. hybrid: regex + LLM safety net (recommended when AI configured). Falls back to regex with a warning if llm/hybrid is requested but no AI provider is configured."New value: +"Extraction strategy: regex (default, free/fast/low recall), llm (AI provider, costs tokens, higher recall), hybrid (regex + LLM safety net). Falls back to regex with a warning if no AI provider is configured."
- Changed
plan_turn1 field changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#"
- Removed
predict_bugs - Changed
query_decisions6 fields changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#" - changed
Input schema / properties / git_branch / descriptionPrevious value: -"Branch filter. \"current\" (default) → current branch + branch-agnostic decisions. \"all\" → every branch. Any other value → that specific branch + branch-agnostic decisions."New value: +"Branch filter: \"current\" (default) = current branch + branch-agnostic; \"all\" = every branch; any other value = that branch + branch-agnostic." - changed
Input schema / properties / index_only / descriptionPrevious value: -"Progressive disclosure (default: false). When true, each decision is returned WITHOUT its full `content` — just id, title, type, code anchors, tags, and a ~1-line `summary`. Pick the relevant ids cheaply, then pull full content with `get_decision`. Pure token-saver."New value: +"Progressive disclosure (default false). true omits full `content` — just id, title, type, anchors, tags, ~1-line `summary`. Pick ids cheaply, then pull full content with `get_decision`." - changed
Input schema / properties / order_by / descriptionPrevious value: -"Result ordering. \"recency\" (default): valid_from DESC. \"created_at\": created_at DESC. \"heat\": time-decay scoring biased toward frequently-recalled + fresh decisions. When heat is disabled in config, \"heat\" gracefully degrades to \"recency\"."New value: +"Result ordering: \"recency\" (default, valid_from DESC), \"created_at\" DESC, or \"heat\" (time-decay favoring frequently-recalled + fresh; degrades to recency if disabled in config)." - changed
Input schema / properties / verification / descriptionPrevious value: -"Filter by verification verdict (implies verify). \"stale\" returns any flagged row (symbol_missing OR code_changed); \"ok\" returns only verified-fresh rows. Omit to return all rows annotated in place."New value: +"Filter by verification verdict (implies verify=true). \"stale\" = any flagged row; \"ok\" = verified-fresh only. Omit to return all rows annotated in place." - changed
Input schema / properties / verify / descriptionPrevious value: -"Staleness verification (default: true). When true, each decision linked to a `symbol_id` is checked against the live index + git history; rows whose code was deleted/renamed or materially changed since `created_at` are flagged with `verification` (\"symbol_missing\" | \"code_changed\") and `stale: true`. Pass false to skip the check entirely."New value: +"Staleness verification (default true). Checks each `symbol_id`-linked decision against the live index + git history; deleted/renamed/materially-changed code is flagged `verification` + `stale: true`. false skips the check."
- Changed
register_edit1 field changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#"
- Removed
reindex - Changed
remember_decision1 field changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#"
- Removed
remove_dead_code - Removed
scan_security - Changed
search14 fields changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#" - changed
Input schema / properties / decorator / descriptionPrevious value: -"Filter to symbols with this decorator/annotation/attribute (e.g. \"Injectable\", \"Route\", \"Transactional\")"New value: +"Filter to symbols carrying this decorator/annotation/attribute" - changed
Input schema / properties / drill_from / descriptionPrevious value: -"Drill scope for mode=\"drill\" — a file path or symbol_id. Results are restricted to the subtree rooted here."New value: +"[mode=\"drill\"] File path or symbol_id to restrict results to." - changed
Input schema / properties / fusion / descriptionPrevious value: -"Enable Signal Fusion Pipeline — multi-channel WRR ranking across lexical (BM25), structural (PageRank), similarity (embeddings), and identity (exact/prefix/segment match). Produces better results than single-channel search."New value: +"Enable Signal Fusion — multi-channel WRR ranking across lexical (BM25), structural (PageRank), similarity (embeddings), and identity match. Weights come from `tune_weights`." - removed
Input schema / properties / fusion_debugRemoved value: -{ - "description": "Include per-channel rank contributions in fusion results.", - "type": "boolean" -} - removed
Input schema / properties / fusion_weightsRemoved value: -{ - "description": "Per-channel weights for fusion (auto-normalized). Defaults: lexical=0.4, structural=0.25, similarity=0.2, identity=0.15.", - "properties": { - "identity": { - "maximum": 1, - "minimum": 0, - "type": "number" - }, - "lexical": { - "maximum": 1, - "minimum": 0, - "type": "number" - }, - "similarity": { - "maximum": 1, - "minimum": 0, - "type": "number" - }, - "structural": { - "maximum": 1, - "minimum": 0, - "type": "number" - } - }, - "type": "object" -} - changed
Input schema / properties / fuzzy / descriptionPrevious value: -"Enable fuzzy search (trigram + Levenshtein). Auto-enabled when exact search returns 0 results."New value: +"Typo-tolerant search. Auto-enabled when exact search returns 0 results." - changed
Input schema / properties / fuzzy_threshold / descriptionPrevious value: -"Minimum Jaccard trigram similarity (default 0.3)"New value: +"[fuzzy] Min trigram similarity (default 0.3)" - changed
Input schema / properties / max_edit_distance / descriptionPrevious value: -"Maximum Levenshtein edit distance (default 3)"New value: +"[fuzzy] Max edit distance (default 3)" - changed
Input schema / properties / mode / descriptionPrevious value: -"Memoir-style retrieval mode: single (default — top-K), tiered (high/medium/low buckets), drill (scoped to drill_from), flat (raw FTS, no PageRank), get (exact lookup). Omit to auto-pick (path-shaped query → get, otherwise → single)."New value: +"single (default): top-K. tiered: high/medium/low buckets. drill: scoped to drill_from. flat: raw FTS, no PageRank. get: exact lookup. Omit to auto-pick." - changed
Input schema / properties / output_format / descriptionPrevious value: -"Output format. \"json\" (default) returns JSON; \"toon\" returns Token-Oriented Object Notation — 30-60% fewer tokens, lossless. \"markdown\" is unsupported here and behaves as json."New value: +"\"json\" (default) or \"toon\" (lossless, 30-60% fewer tokens). \"markdown\" behaves as json here." - added
Input schema / properties / retrieverAdded value: +{ + "description": "Run one named retrieval algorithm instead of the mode dispatcher. Ignores mode/filters/fuzzy/fusion; returns { retriever, items, total }.", + "enum": [ + "lexical", + "semantic", + "hybrid", + "summary", + "feeling_lucky", + "graph_completion" + ], + "type": "string" +} - changed
Input schema / properties / semantic / descriptionPrevious value: -"Semantic mode: auto (default — hybrid if AI available), on (force hybrid), off (lexical-only), only (pure vector). Requires AI provider + embed_repo for non-\"off\" modes."New value: +"auto (default): hybrid if AI available. on: force hybrid. off: lexical-only. only: pure vector. Non-\"off\" needs an AI provider + one embed_repo run." - changed
Input schema / properties / semantic_weight / descriptionPrevious value: -"Hybrid fusion weight in [0,1]. 0 = lexical only, 0.5 = balanced (default), 1 = semantic only."New value: +"[semantic] 0 = lexical only, 0.5 = balanced (default), 1 = vector only."
- Changed
search_text3 fields changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#" - changed
Input schema / properties / grouping / descriptionPrevious value: -"Payload shape. \"flat\" returns a single matches[] array (default). \"by_file\" groups hits under each file — saves tokens on long paths with many hits."New value: +"Payload shape. \"flat\" (default) is a single matches[] array; \"by_file\" groups hits per file — saves tokens on long paths with many hits." - changed
Input schema / properties / timeout_ms / descriptionPrevious value: -"Wall-clock budget in milliseconds. Catastrophic-backtracking regex cannot pin a worker beyond this. Default 2000. Set 0 to disable."New value: +"Wall-clock budget in ms — caps a catastrophic-backtracking regex. Default 2000; 0 disables."
- Removed
self_audit - Changed
suggest_queries1 field changed- removed
Input schema / $schemaRemoved value: -"http://json-schema.org/draft-07/schema#"
99 tool updates
v1.47.1- Removed
add_decision - Removed
analyze_perf - Removed
apply_move - Removed
apply_rename - Removed
approve_decision - Removed
audit_config - Removed
benchmark_project - Removed
build_corpus - Removed
build_decision_clusters - Removed
change_signature - Removed
check_architecture - Removed
check_claudemd_drift - Removed
check_edit_safe - Removed
check_embedding_drift - Removed
compare_branches - Removed
consolidate_decisions - Removed
delete_corpus - Removed
detect_ast_clones - Removed
detect_communities - Removed
detect_drift - Removed
diff_graph_snapshots - Removed
discover_hermes_sessions - Removed
embed_repo - Removed
export_decisions - Removed
export_graph - Removed
export_security_context - Removed
extract_function - Removed
generate_docs - Removed
generate_insights_report - Removed
generate_sbom - Removed
get_api_surface - Removed
get_artifacts - Removed
get_cluster_decisions - Removed
get_co_changes - Removed
get_code_owners - Removed
get_communities - Removed
get_community - Removed
get_cross_domain_deps - Removed
get_cross_workspace_impact - Removed
get_dataflow - Removed
get_decision - Removed
get_decision_clusters - Removed
get_decision_stats - Removed
get_decision_timeline - Removed
get_dependency_diagram - Removed
get_domain_context - Removed
get_domain_map - Removed
get_edge_bottlenecks - Removed
get_file_health_timeline - Removed
get_git_churn - Removed
get_graph_timeline - Removed
get_health_trends - Removed
get_import_graph - Removed
get_minimal_context - Removed
get_package_deps - Removed
get_pagerank - Removed
get_plugin_registry - Removed
get_project_health - Removed
get_project_memo - Removed
get_refactor_candidates - Removed
get_risk_hotspots - Removed
get_session_journal - Removed
get_session_snapshot - Removed
get_suggested_questions - Removed
get_surprises - Removed
get_symbol_owners - Removed
get_type_hierarchy - Removed
get_untested_exports - Removed
get_untested_symbols - Removed
get_wake_up - Removed
graph_query - Removed
index_sessions - Removed
list_bundles - Removed
list_corpora - Removed
list_graph_snapshots - Removed
list_pins - Removed
pack_context - Removed
pin_file - Removed
pin_symbol - Removed
plan_batch_change - Removed
plan_refactoring - Removed
query_by_intent - Removed
query_corpus - Removed
refresh_co_changes - Removed
regenerate_project_memo - Removed
reject_decision - Removed
repair_index - Removed
scan_code_smells - Removed
search_bundles - Removed
search_sessions - Removed
search_with_mode - Removed
snapshot_graph - Removed
taint_analysis - Removed
traverse_graph - Removed
tune_decision_weights - Removed
tune_weights - Removed
unpin - Removed
verify_index - Removed
visualize_graph
TDQS
Scored across 29 tools
Several clusters overlap heavily: get_task_context vs get_feature_context vs get_context_bundle vs plan_turn all deliver 'context', and get_session_analytics vs get_session_stats vs get_usage_trends vs get_real_savings vs get_optimization_report all report token/cost analytics. The descriptions mitigate this with explicit 'use X instead of Y' cross-references, but the boundaries remain genuinely fuzzy and require the agent to read carefully.
Predominantly a consistent snake_case verb_noun pattern (get_symbol, search_text, query_decisions, find_usages, plan_turn), with a strong 'get_' convention for readers. A few single-word outliers (search, batch) break the pattern slightly, but overall it is predictable and readable.
29 tools is on the heavy side for the apparent scope, with visible redundancy across the context and analytics clusters. The preset/load_tools deferral and batch mechanism soften this by allowing tools to be hidden, but the registered surface is still large enough to strain selection.
The surface is broad and largely self-contained: search (symbol + text), retrieval (outline, symbol, bundle, task/feature context), graph analysis (call graph, usages, change impact), diagnostics, decision-graph read/write/mine, and index/session management. Minor gaps remain (e.g. no explicit symbol rename/refactor or test-runner operation), but core code-intelligence workflows are covered.
Maintenance
Related MCP Connectors
Code intelligence platform for AI agents. 20 tools for architecture, security & impact analysis.
Hosted code graph over MCP: exact callers, dependencies, and cross-repo blast radius for AI agents.
Enterprise code intelligence for M&A, security audits, and tech debt. Hosted server with 200k free.
Code intelligence for coding agents: semantic, AST, graph, and full-text search. 279+ languages.
Related MCP Servers
- AlicenseAqualityCmaintenanceCross-repository code knowledge graph MCP server for Java, Kotlin, JavaScript, and TypeScript. Indexes source code into embedded KuzuDB via tree-sitter and exposes 30+ tools for call-flow tracing, multi-hop taint analysis (OWASP/CWE/PCI/STIG), entry-point reachability filtering, performance hotspot detection, and license compliance — without reading source files. 95% fewer tokens vs source-read331MIT
- AlicenseBqualityDmaintenanceInstant codebase knowledge graph MCP server. It auto-detects languages, indexes functions, classes, and call chains, enabling LLMs to navigate code in milliseconds.24MIT
- AlicenseNot gradedqualityAmaintenanceA persistent code-intelligence MCP server that builds a queryable knowledge graph of your codebase, enabling AI assistants to perform cross-file structural reasoning, dependency analysis, and blast radius detection.9MIT
- AlicenseNot gradedqualityBmaintenanceMulti-language code intelligence MCP server providing structured code analysis including symbol search, references, hierarchies, and change impact. Supports 25 languages with persistent indexing and LSP integration.56 npmMIT