predictive-debugger
Use Predictive Debugger as an MCP server to find likely runtime failures in JavaScript/TypeScript projects with seven tools, mostly local and deterministic, plus optional model predictions through an installed Claude Code, Codex or Copilot CLI.
Rank JavaScript/TypeScript source files by risk density, excluding tests by default and optionally including them (
scan_project).Analyze one file for static complexity metrics and a heuristic risk score (
analyze_file).Get TypeScript compiler diagnostics for selected files using the local project's declarations and settings (
check_types).Map a file's local imports, reverse imports and connected tests with source-line evidence (
map_dependencies).Find anomalous log lines ranked by severity and unusual wording (
analyze_logs).Get an independent model verdict on the most likely runtime failure with a line number, reason and confidence, including batch predictions across files (
predict_failures).Check which supported CLI providers are installed and signed in (
list_providers).Ask agents to verify newly written code with a fresh prediction or sub-agent check, and omit dependency context with
calleeContext: false.
Allows the server to use an installed GitHub Copilot CLI's model access to provide independent runtime-failure predictions with a line number, reason, and confidence score.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@predictive-debuggerFind the riskiest files in src/"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Predictive Debugger
Website and docs: predictivedebugger.dev
MCP server for finding likely runtime failures in JavaScript and TypeScript. Seven tools help your coding agent check types, rank risky files, trace dependencies, inspect logs and get an independent model review with a line number and reason.
Uses the Claude Code, Codex or GitHub Copilot CLI you already have installed. Prediction calls use that CLI's model access and usage allowance. No separate API key is needed. Your provider's terms apply to those calls; check them before use, especially on a Claude subscription. See Provider terms.
Setup
Requires Node.js 22 or later. For model predictions, install and sign in to at least one supported CLI. Python 3 is optional for log analysis.
1. Add the MCP server
Choose your agent below. npx downloads and runs the package automatically.
Run in your project on macOS, Linux or WSL:
claude mcp add --scope project predictive-debugger -- npx -y predictive-debugger@latestOn native Windows, from PowerShell:
claude mcp add --scope project predictive-debugger -- cmd /d /c npx -y predictive-debugger@latestUse --scope user to make it available in every project.
For a user-level setup on macOS, Linux or WSL:
codex mcp add predictive-debugger -- npx -y predictive-debugger@latestOn native Windows, from PowerShell:
codex mcp add predictive-debugger -- cmd /d /c npx -y predictive-debugger@latestFor project-only setup or a longer startup timeout, see Codex configuration.
Add to .mcp.json in your project:
{
"mcpServers": {
"predictive-debugger": {
"command": "npx",
"args": ["-y", "predictive-debugger@latest"],
"tools": ["*"]
}
}
}On native Windows, use "command": "cmd" and
"args": ["/d", "/c", "npx", "-y", "predictive-debugger@latest"].
For every project, see Copilot user-level setup.
2. Check the connection
Restart your agent and check /mcp for predictive-debugger and its seven tools.
To check that the package downloads and print its version:
npx -y predictive-debugger@latest --versionRunning without --version starts a stdio server that waits for your agent's
messages. See setup help for local builds and troubleshooting.
3. Ask your agent about the code
Use Predictive Debugger to find the riskiest files in src/.
Show the imports and tests connected to src/services/orders.ts.
Check src/services/orders.ts for likely runtime failures.
Find unusual entries in logs/app.log.Related MCP server: Semantic JS MCP
Tools
Tool | What it does | Model call |
| Rank source files by risk density. Excludes tests by default. | No |
| Return complexity metrics, risk scores and contributing signals. | No |
| Return selected files' TypeScript compiler diagnostics using local project settings. | No |
| Find imports, reverse imports and connected test files, with source-line evidence. | No |
| Return log anomalies, ranked by severity and unusual wording. | No |
| Get an independent model verdict with a line number, reason and confidence. Supports batches. | Yes |
| Check which supported CLIs are installed and their sign-in status. | No |
Start with scan_project and check_types, then read the files they highlight. Use
map_dependencies to find related files and predict_failures when you want a
second opinion. Pass several paths as files so small files share bounded model
calls and avoid repeating the CLI context for every file.
See the tool reference for parameters, result fields and limits.
The server's MCP instructions ask agents to check code they wrote in the current session from a fresh context. The routing depends on file count:
Change | Requested check |
One file, including a feature contained in one file | A fresh |
Several files | A sub-agent scoped to the changed files and intended behavior, where the host supports it |
Mechanical correction with one clear answer | Neither check required |
Grouped predictions retain per-file verdicts. They do not verify a feature's requirements. The benefit of the sub-agent rule has not been measured.
Privacy and limits
Static analysis, dependency maps and log analysis run locally.
predict_failuressends source and bounded dependency context to your CLI's model provider. SetcalleeContext: falseto omit dependency context.Credentials stay with the CLI. MCP tools can read paths the server process can access; project-scoped setup does not restrict file access. See the security model.
JavaScript and TypeScript are supported, including JSX, TSX and decorators. Vue and Svelte single-file components are not supported. Files above 4 MB are rejected; large predictions may cover only selected declarations.
Risk scores and predictions can be wrong. The benchmarks use development cases and do not establish accuracy on arbitrary repositories.
Manually verified on Windows. CI covers Windows, macOS and Linux on Node 22 and 24; real CLI installations on macOS and Linux have not been manually verified.
VS Code preview
An unfinished extension can show findings in the Problems panel. It requires a local build and is not available on the Marketplace or as a prebuilt VSIX. See trying the extension.
Documentation
predictivedebugger.dev: website, quickstart and guides.
Advanced setup: provider login, project scope, updates and local builds.
Tool reference: parameters, prediction results, dependency context and scoring.
VS Code preview: development host, commands and settings.
Benchmarks and method.
Development and contributing. Bug reports are welcome; code contributions are not open yet.
Security policy. Report vulnerabilities privately.
License
MIT.
Available Tools
7 toolsanalyze_fileAnalyze one source fileARead-only
Return static complexity metrics and a heuristic risk score (0-1) for a single JavaScript or TypeScript file. Deterministic and fast — no model call. Call this when reviewing a file to find out where the structural risk sits (nested loops, long functions, async boundaries, unguarded mutation) before reading the whole file yourself.
| Name | Required | Description | Default |
|---|---|---|---|
| file | Yes | Absolute path to a .js/.jsx/.ts/.tsx file |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false, and openWorldHint=false, so safety is covered. The description adds genuine value beyond that: determinism ('Deterministic and fast — no model call') and the nature of the output (metrics plus a 0-1 risk score), which sets expectations for cost and reliability. It omits any file-size or performance limits, keeping it just shy of 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the return value, then the mechanism, then the use case. Every sentence carries distinct information with no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With one parameter, no output schema, and full coverage of the return shape (metrics + risk score) in the text, an agent has enough to invoke it correctly. Minor gaps around output structure specifics remain, but not enough to impede correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100% and the single 'file' parameter is fully documented as 'Absolute path to a .js/.jsx/.ts/.tsx file'. The description adds no syntax or format detail beyond the schema, so the baseline of 3 for high coverage applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Return static complexity metrics and a heuristic risk score') scoped precisely to 'a single JavaScript or TypeScript file'. This makes it immediately distinguishable from siblings like scan_project or map_dependencies, which operate at project or dependency scope.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives clear when-to-use context: 'Call this when reviewing a file to find out where the structural risk sits ... before reading the whole file yourself.' It also enumerates what to look for (nested loops, long functions, async boundaries, unguarded mutation). It stops short of naming the sibling to use for whole-project analysis, so no explicit exclusion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
analyze_logsFind anomalous log linesARead-only
Score a log file's lines by severity and how unusual their wording is, and return the anomalies, worst first. Deterministic — no model call, no API key. Call this when you have a log file and want the handful of lines worth reading rather than the whole file.
| Name | Required | Description | Default |
|---|---|---|---|
| logFile | Yes | Absolute path to the log file | |
| threshold | No | Anomaly score cutoff, 0-1 (default 0.5) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, non-destructive and closed-world, so the safety profile is covered. The description adds real value beyond that by disclosing that scoring is deterministic with no model call or API key, and that results are ordered worst-first. It omits any note on result limits or volume handling.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, with the core action and ranking front-loaded in the first sentence. The second sentence carries distinct usage and determinism information rather than restating the name.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a two-parameter read-only tool with no output schema, the description covers purpose, ranking, ordering and determinism. The remaining gap is absence of any detail on how many anomalies are returned or any cap on result size, which an agent might want to know.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, with both the path and the threshold range/default fully documented in the schema. The description adds no extra meaning to either parameter, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (score a log file's lines), the ranking method (severity + word rarity), and the output ordering (anomalies, worst first). This clearly separates it from siblings like analyze_file and scan_project.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives an explicit when-to-use trigger ('you have a log file and want the handful of lines worth reading rather than the whole file'), which frames the alternative behavior. It stops short of naming a specific sibling tool to use instead, so it is clear context without explicit alternatives.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
check_typesCheck selected files with TypeScriptARead-only
Return TypeScript compiler diagnostics for selected JavaScript/TypeScript files. Uses the full local project's declarations and settings, but reports selected files only. No model, emit, plugins or project code execution. Without a config, uses inferred null checking with JS checking. Reports skipped files, incomplete context and limits; no diagnostics is not proof of runtime safety.
| Name | Required | Description | Default |
|---|---|---|---|
| files | Yes | Absolute paths of up to 20 files in one project | |
| project | No | Explicit tsconfig.json/jsconfig.json path; otherwise discovered |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Adds substantial behavioral context beyond the readOnly/openWorld/destructive annotations: no model, emit, plugins, or project code execution; inferred null checking plus JS checking when no config is present; and it discloses that skipped files, incomplete context, and limits are reported. The closing caveat that "no diagnostics is not proof of runtime safety" is exactly the kind of nuance annotations cannot express.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, front-loaded with what is returned, then scope, then side-effect-free guarantees, then caveats. No filler; each sentence adds a constraint or capability an agent would otherwise have to guess.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description still explains what comes back (compiler diagnostics) and that skipped files, incomplete context, and limits are surfaced. Combined with annotations covering safety, nothing material is missing for correct invocation.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3, but the description adds real semantics for the project parameter by stating what happens when no config is supplied (inferred null checking with JS checking). The files parameter's semantics are left to the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Names a specific verb and resource ("Return TypeScript compiler diagnostics for selected JavaScript/TypeScript files") and scopes it ("reports selected files only", "uses the full local project's declarations"). This implicitly separates it from scan_project/analyze_file, but no sibling is named explicitly, so it falls just short of the top band.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Usage is implied by the scoping language (selected files, full project context), but there is no explicit when-to-use vs. when-not, no mention of the sibling tools, and no stated prerequisites such as needing an installed compiler. Adequate but leaves routing to inference.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
list_providersShow available CLI providersARead-only
Report which supported CLIs (Claude Code, Codex, GitHub Copilot) are installed and signed in. Call this to diagnose why predict_failures is failing.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, destructiveHint=false and a closed (non-open-world) scope, so the safety profile is covered. The description adds that the check covers both installation and sign-in state, which is useful context, but says nothing about behavior with zero providers found or what state is reported.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two tight sentences: the first front-loads what is reported, the second gives the diagnostic use case. No filler or redundancy.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no parameters and no output schema, the description adequately tells the agent what the tool determines and why to reach for it. Only minor gaps remain, such as the shape or granularity of the reported provider status.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters, so per the rubric the baseline is 4. There is nothing for the description to clarify beyond what the empty schema already conveys.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: report which supported CLIs are installed and signed in, naming the concrete set (Claude Code, Codex, GitHub Copilot). It also distinguishes itself from the sibling predict_failures by framing itself as its diagnostic counterpart.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Call this to diagnose why predict_failures is failing,' giving a clear triggering condition tied to a named sibling. It stops short of stating when NOT to call it (e.g., routine health checks), so it falls just short of the top score.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
map_dependenciesMap a file's dependency neighborhoodARead-only
Find a file's local imports, reverse imports and tests connected by imports. Each relationship includes a source path and line as evidence. Static file relationships, not runtime callers or test coverage. Scans JavaScript/TypeScript including tests, with bounded work and explicit unresolved imports and scan limits. Deterministic; no provider call.
| Name | Required | Description | Default |
|---|---|---|---|
| file | Yes | Source file, absolute or relative to directory | |
| depth | No | Import hops in each direction, default 1 | |
| limit | No | Total neighboring files to return, default 50 | |
| maxFiles | No | Maximum source files to discover, default 1000 | |
| directory | Yes | Project directory; source outside this directory is excluded |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint/destructiveHint/openWorldHint, so the safety profile is covered; the description adds genuinely new traits: determinism ('no provider call'), bounded work with surfaced 'unresolved imports and scan limits', and the evidence format (source path + line per relationship). This goes meaningfully beyond the annotations, though it doesn't detail failure modes or how unresolved imports are reported.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four tight sentences, front-loaded with the purpose, then scope, behavior and determinism. No sentence is redundant: each adds a distinguishing constraint (edge types, static-only, language scope, bounded/deterministic execution).
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description takes on the burden of explaining returns and does so ('Each relationship includes a source path and line as evidence'), plus it discloses static vs runtime semantics, language support, bounded scanning and determinism. Nothing an agent needs to decide or invoke correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so depth, limit, maxFiles and directory are already documented in the schema, including defaults and bounds. The description only loosely gestures at these via 'bounded work and explicit ... scan limits' and adds no syntax or interaction detail, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb and resource ('Find a file's local imports, reverse imports and tests connected by imports') and names the exact edge types returned, so the tool's scope is unambiguous. It does not explicitly contrast itself with siblings like analyze_file or scan_project, so it stops short of full sibling differentiation.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly bounds when this applies ('Static file relationships, not runtime callers or test coverage') and states the language scope ('JavaScript/TypeScript including tests'), which tells an agent when this tool is the right fit. It still never names an alternative tool to use instead for runtime callers or coverage.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
predict_failuresPredict the most likely runtime failure in a fileARead-only
Combine static analysis with a second-opinion verdict from the signed-in Claude Code, Codex, or GitHub Copilot CLI, returning the most likely runtime failure with a line number and reason. status distinguishes actionable, uncertain, no-finding, and unavailable results. checked lists the bug categories the model reports having considered, so a clean file weighed against the whole catalogue is distinguishable from one where it stopped early; it is a self-report, and an empty list means no coverage was reported. Pass multi: true to get every finding the model can demonstrate, ranked, in a findings array instead of one verdict — experimental, and more findings per call is also more surface for false positives per call. actionable applies the measured score >= 0.7 precision gate. uncertain names a candidate the reviewer could not confirm: read its cited lines before dismissing it, and report it only if you confirm it yourself. This spawns another model and takes 5-15 seconds, so only call it when you specifically want an independent second opinion. If you are yourself reviewing the code, use analyze_file and read the source instead. Reviewing several files? Pass them all as files in one call rather than calling once per file: small files share bounded model calls to reduce repeated CLI context. Groups run concurrently, with individual verdicts and explicit failures.
| Name | Required | Description | Default |
|---|---|---|---|
| file | No | Absolute path to a .js/.jsx/.ts/.tsx file | |
| files | No | Absolute paths to review in one call. Small files share bounded model calls, reducing repeated CLI context; large files run alone. Replies carry a `results` array in the order given. Supersedes `file`. At most 100; only JavaScript/TypeScript source is sent to the model. | |
| model | No | Model override passed to the CLI | |
| multi | No | Return every finding the model can demonstrate, ranked by score, rather than the single most likely one (default false). Experimental: the precision gate was measured on one-finding replies, so `actionable` is less well characterised here. | |
| logFile | No | Optional log file to fold into the combined score | |
| verbose | No | Include the static metric counts and the full log breakdown (default false) | |
| provider | No | Which CLI to ask (default: whichever is installed) | |
| concurrency | No | Model calls in flight at once (default 4), each covering up to eight files. Lower it if the provider starts rate-limiting. | |
| calleeContext | No | Also send bounded imported definitions and referenced type contracts so the model can check dependency behavior (default true). Turning this off reduces input tokens but removes that evidence. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Goes well beyond the readOnly/openWorld annotations by disclosing cost and latency (spawns another model, 5-15 seconds), rate-limit behavior (lower `concurrency` if the provider throttles), the precision-gate semantics of `actionable`, the experimental status of `multi`, and the false-positive surface tradeoff.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but front-loaded, leading with the core purpose before the status/checked/multi details and the usage/alt-tool guidance. Every sentence carries information, though the volume is high enough that some trimming is possible.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, yet the description explains return semantics (status categories, checked self-report, findings array, per-file verdicts with explicit failures), which is exactly what an agent needs for a 9-param, no-required-param tool.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With 100% schema coverage the baseline is 3, but the description adds real meaning: `files` supersedes `file` with a results array, `multi` changes the return shape, `actionable` applies a measured >=0.7 gate, and `concurrency` defaults to 4 with rate-limit guidance. It does not fully cover `model`, `provider`, or `logFile`, so it is not a 5.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (predict) and resource (runtime failure in a file) and names the mechanism (static analysis plus a second-opinion CLI verdict). It clearly distinguishes itself from sibling analyze_file, which it explicitly routes simple self-review cases to.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit when-to-use ('only call it when you specifically want an independent second opinion'), when-not ('if you are yourself reviewing the code, use analyze_file and read the source instead'), and a batching alternative (pass all files in one call rather than calling per file).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scan_projectRank a project's files by riskARead-only
Walk a directory and rank its JavaScript/TypeScript files by risk density — how concentrated the failure-prone code is, not how big the file is. Test files are left out by default; pass includeTests to rank them too. Deterministic and fast — no model call. Call this at the start of a code review to decide which files are worth your attention, instead of reading the tree in arbitrary order.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Maximum number of files to return (default 50) | |
| verbose | No | Include the raw metric counts for every file (default false) | |
| directory | Yes | Absolute path to the directory to scan | |
| includeTests | No | Rank test files too — *.spec.*, *.test.*, and anything under __tests__/test/tests/spec/__mocks__ (default false). They rank high for a structural reason rather than a real one: mocked awaits read as async complexity. Turn this on to audit a suite's own complexity. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover the safety profile (readOnly, non-destructive), and the description adds meaningful traits beyond them: it is deterministic and requires no model call, test files are excluded by default, and the ranking metric is density-based rather than size-based. The return format itself is not described, which keeps it from a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, each earning its place: the core behavior and metric come first, then the test-file caveat, then the determinism note, then the workflow recommendation. No filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With no output schema, the description conveys what the tool conceptually returns (files ranked by risk density) and how to act on it, and annotations cover the safety profile. A brief note on output shape (list vs. scored entries) would close the remaining gap.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters including includeTests' exclusion patterns are already documented in the schema. The description only restates the includeTests default and purpose, adding little beyond the structured field, so the baseline of 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (walk/rank) and resource (JavaScript/TypeScript files in a project directory) with a precise metric definition: risk density, explicitly contrasted with file size. This clearly separates it from siblings like analyze_file and map_dependencies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit timing guidance -- 'call this at the start of a code review' -- and frames the alternative as 'instead of reading the tree in arbitrary order'. It does not name a sibling tool as the alternative, but the workflow placement is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.9.0- Added
check_types - Changed
predict_failures3 fields changed- changed
Input schema / properties / concurrency / descriptionPrevious value: -"Verdicts in flight at once for a batch (default 4). Lower it if the provider starts rate-limiting."New value: +"Model calls in flight at once (default 4), each covering up to eight files. Lower it if the provider starts rate-limiting." - changed
Input schema / properties / files / descriptionPrevious value: -"Absolute paths to review in one call, run concurrently. Prefer this over one call per file when checking a change set: the verdicts are independent, so a batch bills the same as the same files one at a time but finishes in roughly the time of the slowest one. Replies carry a `results` array in the order given. Supersedes `file`."New value: +"Absolute paths to review in one call. Small files share bounded model calls, reducing repeated CLI context; large files run alone. Replies carry a `results` array in the order given. Supersedes `file`. At most 100; only JavaScript/TypeScript source is sent to the model." - added
Input schema / properties / files / maxItemsAdded value: +100
6 tool updates
v0.8.2- First observed
analyze_file - First observed
analyze_logs - First observed
list_providers - First observed
map_dependencies - First observed
predict_failures - First observed
scan_project
TDQS
Scored across 7 tools
Each tool targets a distinct granularity or data source: check_types (compiler diagnostics), analyze_file (single-file metrics), scan_project (whole-tree ranking), map_dependencies (import graph), analyze_logs (log lines), list_providers (environment probe), and predict_failures (model verdict). The descriptions explicitly clarify the boundary between analyze_file and predict_failures, and between scan_project and analyze_file, removing overlap.
All seven tools follow a consistent snake_case verb_noun pattern (check_types, map_dependencies, analyze_file, scan_project, analyze_logs, list_providers, predict_failures). Verbs are predictable and nouns name the resource clearly.
Seven tools is well-scoped for a predictive debugging server, covering static analysis, dependency mapping, logs, environment probing, and model-based prediction without redundancy. Each tool earns its place with a clear role in the workflow.
The surface covers the debugging lifecycle well: project triage, per-file analysis, dependency context, logs, provider diagnostics, and independent failure prediction. Minor gaps exist, such as no way to read raw file content or validate configuration directly, but agents can work around these with standard file access.
Maintenance
Related MCP Connectors
Synthetic checks, nightly regression replay and model-drift alerts for AI agents
Find your AI agent's likely failure mode, get runtime settings, and clarify ambiguous prompts.
Change-aware CI validation and affected-test guidance for coding agents.
Change-aware CI validation and affected-test guidance for coding agents.
Related MCP Servers
- FlicenseAqualityDmaintenanceEnables LLMs to automatically diagnose coding errors through codebase search, test execution, and live debugger integration (DAP/V8 CDP). Provides a secure, policy-gated environment for investigating failures while preventing destructive operations.9-
- AlicenseNot gradedqualityAmaintenanceProvides structured semantic context for JavaScript/TypeScript codebases, enabling coding agents to navigate, review, and change code with explicit uncertainty.25 npm1MIT
- FlicenseNot gradedqualityBmaintenanceEnables AI-powered code review and analysis of asynchronous JavaScript/TypeScript control flow, identifying unawaited promises and race conditions through AST taint-flow tracing and ESLint-style audits within MCP-compliant clients.8-
- AlicenseNot gradedqualityBmaintenanceEnables AI coding agents to verify, compare, and safely upgrade npm packages by scoring packages from public signals and sandboxed co-installation, detecting breaking API surface changes, and rewriting code that upgrades break.143 npm8MIT