DevPilot MCP
Provides read-only access to GitHub issues, including fetching issue details, comments, labels, linked PRs, and analyzing issue content for error messages, stack frames, and suggested search queries.
Enables running the project's Jest test suite and summarizing failures, with parsing of Jest JSON output to identify test names, files, messages, and likely source files.
Enables running the project's pytest test suite and summarizing failures, with parsing of pytest JUnit XML output to group and report test failures.
Enables running the project's Vitest test suite and summarizing failures, with parsing of Vitest JSON output to identify test names, files, messages, and likely source files.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@DevPilot MCPtriage issue #12 in Ashbruh22/devpilot_mcp and propose a fix"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
DevPilot MCP
Give your coding agent the issue, the code, and the failing tests in one loop.
DevPilot is a Model Context Protocol server (TypeScript, MCP SDK v2) that gives
AI coding agents like Claude Code structured, scoped access to GitHub issues, a codebase, its docs, and its test
suite. It runs locally over stdio (npx) or remotely over Streamable HTTP.
Live server: devpilot-mcp.onrender.com (the landing page, with
/mcpas the MCP endpoint). See Deployment.
Why
Most of the time an agent spends on a bug goes to finding context: reading the issue, guessing which files matter, grepping, running the whole test suite, and scrolling through raw output. DevPilot collapses that into six structured, size-capped tool calls:
get_issue → analyze_issue → search_codebase → run_tests → summarize_test_failures → fixThe server makes no LLM calls. Extraction is deterministic, which keeps it cheap, fast, and testable, and the client's model already does the reasoning. See Benchmark for how the time savings are measured.
Related MCP server: GitHub MCP Server
Tools
Tool | Input | Output |
|
| title, state, labels, author, dates, body, comments |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
Plus the triage_issue prompt (owner, repo, issue_number), which walks the agent through the whole loop and ends with a proposed fix.
Every tool validates its input with zod (every field is described for the model), declares an outputSchema, returns
structured content plus a one-line summary, truncates with an explicit …[truncated N chars] marker, and reports
errors as actionable tool errors (Issue #99 not found in owner/repo) rather than stack traces.
owner/repo are optional. Locally they default to the workspace's origin remote, and remotely to the only
allowlisted repo. In remote mode with several allowlisted repos, the workspace tools also take repo: "owner/repo".
Quick start
Local (stdio)
Run it inside the repository you want the agent to work on:
claude mcp add devpilot -- npx -y @ashbruh22/devpilot-mcp
# optional, for private repos and higher rate limits:
claude mcp add devpilot -e GITHUB_TOKEN=github_pat_xxx -- npx -y @ashbruh22/devpilot-mcpThen ask Claude Code to "triage issue #12", or run the prompt directly: /mcp__devpilot__triage_issue <owner> <repo> 12.
Any other MCP client works too:
{ "mcpServers": { "devpilot": { "command": "npx", "args": ["-y", "@ashbruh22/devpilot-mcp"] } } }Remote (Streamable HTTP)
claude mcp add --transport http devpilot https://devpilot-mcp.onrender.com/mcp
# if the server sets DEVPILOT_API_KEY:
claude mcp add --transport http devpilot https://devpilot-mcp.onrender.com/mcp --header "Authorization: Bearer <key>"The hosted server works on the demo project bundled in this repo (demo/devpilot-demo), which
has two real bugs. Try: "Use devpilot to run the tests, summarize the failures, and propose a fix for the largest
group." Or triage one of the demo issues end to
end: /mcp__devpilot__triage_issue Ashbruh22 Devpilot_MCP 1.
Architecture
flowchart LR
C["MCP client<br/>(Claude Code, Inspector)"] <--> S["stdio<br/>src/stdio.ts"]
C <--> H["Streamable HTTP<br/>src/http.ts<br/>auth · rate limit · CORS"]
S <--> T["Tool layer<br/>createServer(deps)"]
H <--> T
T --> G["GitHub API<br/>(read-only, 60 s cache)"]
T --> R["ripgrep<br/>(path-guarded)"]
T --> X["Test runner<br/>(no shell, timeout, tree kill)"]Two transports, one core.
createServer(deps)builds anMcpServerwith every tool registered; it knows nothing about transports.stdio.tsserves it withserveStdio, andhttp.tsserves it statelessly withcreateMcpHandler+toNodeHandleron Express, building one server per request. Long-lived state (GitHub cache, run store, workspaces) lives indepsand is shared.Token budget by design. Outputs are structured, every string is capped, search results shrink to fit
MAX_OUTPUT_CHARS, test output keeps only the tail, and failures are grouped by normalized error message (numbers, strings, and paths become placeholders), so 40 failures with one root cause read as one group.likely_source_filesare the non-test project files in each failure's stack. When the stack only reaches the test file (typical for assertion failures), DevPilot falls back to naming:src/pricing.test.ts→src/pricing.ts.
src/
├─ server.ts createServer(deps) + createDeps(config)
├─ stdio.ts local entry (bin)
├─ http.ts remote entry: Express, Streamable HTTP, landing page, /healthz
├─ config.ts env parsing with zod, fail-fast
├─ tools/ one file per tool
├─ prompts/ triage_issue
├─ lib/ github, workspace, pathGuard, exec, runStore, truncate, search, docs, failures, …
│ └─ parsers/ vitest/jest JSON, pytest JUnit XML, raw output, stack traces (JS + Python)
└─ public/index.html landing pageSecurity model
Running tests on a public server is remote code execution by design, so remote mode is locked down:
Allowlist only. Tools accept only repos in
ALLOWED_REPOS. At boot, each is shallow-cloned (git clone --depth 1 --branch <ref>) into/tmp/workspaces/<owner>__<repo>, and its dependencies are installed once withnpm ci --ignore-scripts.Fixed commands. Test commands come from
TEST_COMMANDS, never from tool input.filtermust match^[\w\-./ ]{1,100}$, can't start with-(so it can't inject flags), and can't contain... Commands run without a shell, and configured commands may not contain shell operators.Resource limits. Each run has a
TEST_TIMEOUT_MStimeout that kills the whole process group, output is capped, and each repo allows at most one concurrent test run ("busy, retry in N s")./mcpis rate-limited per IP (60/min by default) and request bodies are capped at 1 MB.Least privilege. The container runs as the non-root
nodeuser withtinias PID 1. The GitHub token should be a fine-grained, read-only PAT for public repos.GITHUB_TOKEN,DEVPILOT_API_KEY, and anything that looks like a secret (*TOKEN*,*SECRET*,*PASSWORD*,*API_KEY*, …) are stripped from child-process environments.Path guard. Every client-supplied path is resolved with
path.resolve, must stay under the workspace root, and is re-checked withrealpathso symlinks can't escape. Absolute paths,.., and NUL bytes are rejected. Search results go through the same check.No arbitrary network fetches. No tool fetches a URL from its input. GitHub is the only outbound API, and clone URLs are built from the allowlist.
The HTTP server also uses the SDK's Host-header validation (DNS-rebinding protection) and Origin validation on
localhost binds, an optional constant-time bearer-token check on /mcp (DEVPILOT_API_KEY), and CORS that exposes
only the MCP headers. Without a key, the server runs as a public read-only demo limited to the allowlist.
Configuration
All configuration comes from environment variables, validated at startup. See .env.example.
Var | Mode | Default | Purpose |
| both |
|
|
| both | – | Fine-grained PAT. Read-only, public repos only for the hosted deployment |
| local |
| Repo to operate on |
| local | auto-detect | Override the detected test command |
| remote | – (required) |
|
| remote |
| JSON |
| remote | – | Bearer token for |
| remote |
| Listen address ( |
| remote | Render hostname | Host-header allowlist for public binds |
| remote |
| Per-IP limit on |
| both |
| Test run timeout |
| both |
| Cap on any tool's text response |
| remote |
| Where to clone, and whether to install deps |
Test runner auto-detection (local mode) prefers machine-readable output:
vitest → vitest run --reporter=json, jest → jest --json, pytest → pytest --junitxml=… -q, otherwise
npm test with best-effort text parsing. Binaries are resolved from node_modules/.bin (walking up for monorepos)
or run with npx --no, so nothing is ever downloaded.
Development
npm install
npm run typecheck && npm run lint && npm test # unit + integration tests
npm run build # tsc → dist/, keeps the shebang, chmod +x
npm run dev:stdio # run from source over stdio
npm run dev:http # run from source over HTTP on :3000
npm run inspector # MCP Inspector against dist/stdio.js
npx @modelcontextprotocol/inspector # then connect to http://localhost:3000/mcpThe tests cover:
Unit: path guard (traversal, absolute paths, symlink escape), parsers (vitest/jest JSON, pytest JUnit XML, raw output), stack traces (Node, browser, Python), issue analysis, truncation, exec (tree kill on timeout, secret stripping, no shell), config validation.
Integration: a real MCP client (
@modelcontextprotocol/client) over in-memory and Streamable HTTP transports. It calls every tool againsttest/fixtures/sample-repo(real vitest runs with real failures) with GitHub mocked bynock. HTTP tests cover/healthz, 401 without a token, 429 rate limiting, CORS, Host/Origin validation, and body limits. Remote-mode tests clone from a local bare git repo and check allowlist enforcement and fixed commands.
Deployment
Docker → Render. The repo includes a multi-stage Dockerfile (non-root, tini, git) and a
render.yaml blueprint:
In Render, go to New → Blueprint and pick this repo. Render reads
render.yaml, builds the Docker image, and creates a free web service.When prompted, set
GITHUB_TOKEN: a fine-grained PAT with read-only access to public repositories. It's needed forget_issue/analyze_issue, since unauthenticated calls share a 60/hour limit. LeaveDEVPILOT_API_KEYempty for a public demo, or set it to require a bearer token.On boot, the server shallow-clones this repo, uses
demo/devpilot-demoas its workspace (ALLOWED_REPOS=Ashbruh22/Devpilot_MCP:demo/devpilot-demo), and installs that project's dependencies./healthzreports"ready"after a few seconds. Render's hostname is added to the Host allowlist automatically.Connect with
claude mcp add --transport http devpilot https://devpilot-mcp.onrender.com/mcp.Free instances sleep when idle: the first request can take 30–60 s. To keep it warm, set the repository variable
DEVPILOT_URLto the service URL, which enables thekeepaliveworkflow (a 10-minute ping). Check Render's current free-tier terms first.
npm. Tag a release (git tag v1.0.0 && git push --tags) and the release
workflow publishes with provenance (it needs an NPM_TOKEN secret). Or run npm publish --access public by hand, then
check it with npx -y @ashbruh22/devpilot-mcp in a clean folder.
Benchmark
bench/ has a reproducible method: 10 tasks on the demo repo ("find the root cause of issue #N and the file
to change"), each timed once without DevPilot (the agent reads files and runs commands itself) and once with the
triage_issue flow. It records time-to-context and tool calls or steps. Results go in
bench/results.csv, and npm run bench:summary prints the medians.
Status: the harness and tasks are in place. The results table is empty until the runs are done, so no time-saved number is claimed yet. Quote only the measured median.
Available Tools
6 toolsanalyze_issueAnalyze GitHub issueARead-onlyIdempotent
Deterministically extract triage signals from an issue (no LLM): a type guess (bug/feature/question/docs), error messages, parsed stack frames (JS and Python), mentioned file paths, fenced code blocks, and suggested queries to pass to search_codebase.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | GitHub repository name. Optional, see "owner". | |
| owner | No | GitHub owner (user or org). Optional: defaults to the workspace's "origin" remote (local mode) or the only allowlisted repo (remote mode). | |
| issue_number | Yes | Issue (or PR) number, e.g. 42. |
Output Schema
| Name | Required | Description |
|---|---|---|
| repo | Yes | |
| owner | Yes | |
| title | Yes | |
| type_guess | Yes | |
| code_blocks | Yes | |
| issue_number | Yes | |
| stack_frames | Yes | |
| type_signals | Yes | |
| error_messages | Yes | |
| mentioned_paths | Yes | |
| suggested_search_queries | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, and openWorld, so the safety profile is covered. The description adds a genuine behavioral trait beyond that: the extraction is 'deterministic' and uses 'no LLM', telling the agent results are stable and reproducible. It doesn't discuss rate limits or failure modes, but the determinism claim is real added value.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single front-loaded sentence that leads with the verb and the determinism guarantee, then enumerates outputs. The enumeration is long but each item is informative and non-redundant; only the parenthetical nesting makes it slightly dense.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need no explanation, yet the description helpfully names the signal categories anyway. Combined with annotations covering the safety profile and 100% schema coverage, an agent has everything needed to invoke it correctly, though nothing is said about behavior on invalid issue numbers or missing repos.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so owner, repo, and issue_number are already fully documented in the schema, including the origin-remote fallback behavior. The description adds no parameter-level syntax or format guidance, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('extract triage signals from an issue') and enumerates the exact outputs (type guess, error messages, stack frames, file paths, code blocks, suggested queries). The parenthetical '(no LLM)' plus the enumerated artifacts make it impossible to confuse with get_issue or get_docs.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly signals the workflow context by noting the suggested queries are meant 'to pass to search_codebase', giving the agent a downstream routing cue. It stops short of stating when NOT to use it (e.g., vs. get_issue for raw issue body retrieval), so it is clear context without explicit exclusions.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_docsGet project docsARead-onlyIdempotent
Read project documentation (README*, CONTRIBUTING*, root .md, docs/**/.md). With "path": return that file, or just the section whose heading matches "topic". With only "topic": return the 3 best-matching sections ranked by keyword hits. With neither: return the README intro and the list of docs.
| Name | Required | Description | Default |
|---|---|---|---|
| path | No | A specific doc file relative to the workspace root, e.g. "docs/setup.md". | |
| repo | No | Remote mode only: which allowlisted repo ("owner/repo") to operate on. Optional when only one repo is allowlisted; ignored in local mode. | |
| topic | No | Keywords to look for, e.g. "discount rounding" or "configuration". | |
| max_chars | No | Total character budget for returned section content. Default 8000. |
Output Schema
| Name | Required | Description |
|---|---|---|
| sections | Yes | |
| workspace | Yes | |
| available_docs | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already cover safety (readOnly, idempotent, non-destructive, closed world), so the bar is lower. The description nonetheless adds real behavioral detail: the fallback behavior when no arguments are supplied, the ranked-by-keyword-hits selection, and the fixed cap of 3 sections for topic-only calls. It doesn't discuss truncation interplay with max_chars in any depth, which is the only notable gap.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three tight sentences, front-loaded with the resource scope, then the three invocation modes in a predictable path/topic/neither progression. Every clause carries a distinct constraint; nothing is filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return-value documentation is not required, yet the description still specifies what each mode yields. With 4 optional parameters, zero required, and full schema coverage, an agent has everything needed to call it correctly in any mode.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description goes beyond the per-field schema text by explaining the interaction semantics: path alone returns the file or just the heading-matched section when topic is also present, and topic alone switches to ranked multi-section retrieval. That combinatorial meaning is not derivable from the schema fields individually.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource (read project documentation) and enumerates the exact file patterns covered (README*, CONTRIBUTING*, root *.md, docs/**/*.md), which is unusually precise. It does not, however, explicitly contrast itself with the sibling search_codebase, so the boundary between doc reading and code search is left to inference.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives explicit mode-selection rules: path+provision to fetch a file or matching section, topic only to get the 3 best keyword-ranked sections, neither to get the README intro plus docs list. That is clear operational guidance for invoking the tool. It stops short of naming when a sibling (e.g. search_codebase) would be preferable.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
get_issueGet GitHub issueARead-onlyIdempotent
Fetch a GitHub issue with its discussion: title, state, labels, author, dates, body, recent comments, and linked pull requests. Read-only. Start here when triaging an issue.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | GitHub repository name. Optional, see "owner". | |
| owner | No | GitHub owner (user or org). Optional: defaults to the workspace's "origin" remote (local mode) or the only allowlisted repo (remote mode). | |
| issue_number | Yes | Issue (or PR) number, e.g. 42. | |
| max_comments | No | Maximum number of comments to return (most recent first kept). Default 20. |
Output Schema
| Name | Required | Description |
|---|---|---|
| url | Yes | |
| body | Yes | |
| repo | Yes | |
| owner | Yes | |
| state | Yes | |
| title | Yes | |
| author | Yes | |
| labels | Yes | |
| number | Yes | |
| comments | Yes | |
| created_at | Yes | |
| linked_prs | Yes | |
| updated_at | Yes | |
| total_comments | Yes | |
| is_pull_request | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false. The description repeats 'Read-only' and lists returned content, but adds no behavioral context such as auth needs, rate limits, or failure modes.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two short sentences, front-loaded with the action and resource, followed by scope and usage cue. No wasted wording.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With rich annotations, 100% schema description coverage, and an output schema present, the description supplies enough context for correct invocation. It does not need to explain return values further.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all four parameters are already documented in the schema. The description adds no parameter-level meaning beyond what the schema provides, making 3 the correct baseline.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: fetch a GitHub issue with its discussion. It also lists the returned fields and positions the tool as the starting point for triage, distinguishing it from sibling analyze_issue.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicitly says 'Start here when triaging an issue,' giving clear usage context. It does not name when not to use it or explicitly route to alternatives like analyze_issue, so it stops short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
run_testsRun testsA
Run the project's test suite (auto-detects vitest, jest, pytest, or npm test; in remote mode only the server's fixed, configured command runs). Returns pass/fail counts and a run_id; pass the run_id to summarize_test_failures for grouped failures and the source files to fix.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | Remote mode only: which allowlisted repo ("owner/repo") to operate on. Optional when only one repo is allowlisted; ignored in local mode. | |
| filter | No | Optional test name or file pattern, e.g. "pricing" or "src/pricing.test.ts". Letters, digits, _ - . / and spaces only. |
Output Schema
| Name | Required | Description |
|---|---|---|
| failed | Yes | |
| passed | Yes | |
| run_id | Yes | |
| runner | Yes | |
| command | Yes | |
| skipped | Yes | |
| exit_code | Yes | |
| timed_out | Yes | |
| workspace | Yes | |
| duration_ms | Yes | |
| output_tail | Yes | |
| structured_report | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=false, openWorldHint=false, idempotentHint=false, and destructiveHint=false. The description adds valuable context beyond annotations: it auto-detects vitest/jest/pytest/npm test, notes that remote mode runs only a fixed configured command, and returns pass/fail counts plus a run_id. It does not contradict any annotation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, front-loaded sentence that covers purpose, auto-detection, remote-mode behavior, and return-value routing without waste. Every clause earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a test-running tool with an output schema and annotations, the description supplies the missing operational context: what gets auto-detected, how remote mode differs, and what the run_id is for. Nothing an agent needs to call or use the tool correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the schema already documents both 'repo' and 'filter' thoroughly. The description does not add parameter-level semantics beyond what the schema provides, which matches the baseline 3 for high-coverage schemas.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource: 'Run the project's test suite', and names the auto-detected frameworks. It also distinguishes itself from the sibling summarize_test_failures by explaining the run_id handoff, so an agent can tell the tools apart without opening schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Provides clear workflow guidance: use run_tests to execute the suite, then pass the returned run_id to summarize_test_failures for grouped failures. It does not explicitly state when not to use this tool or contrast it with other siblings like search_codebase, but the primary usage context is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
search_codebaseSearch codebaseARead-onlyIdempotent
Search the workspace with ripgrep (respects .gitignore; skips node_modules, dist, lockfiles and binaries). Returns matching lines with surrounding context and paths relative to the workspace root.
| Name | Required | Description | Default |
|---|---|---|---|
| repo | No | Remote mode only: which allowlisted repo ("owner/repo") to operate on. Optional when only one repo is allowlisted; ignored in local mode. | |
| query | Yes | Text to search for. Treated literally unless is_regex is true. | |
| is_regex | No | Interpret query as a (Rust-flavored) regular expression. | |
| path_glob | No | Restrict to files matching this glob, relative to the root, e.g. "src/**/*.ts" or "*.py". | |
| max_results | No | Maximum matches to return. Default 30. | |
| context_lines | No | Lines of context before and after each match. Default 2. |
Output Schema
| Name | Required | Description |
|---|---|---|
| matches | Yes | |
| truncated | Yes | |
| workspace | Yes | |
| total_matches | Yes | |
| files_with_matches | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnly, idempotent, non-destructive, closed-world, so the safety profile is covered. The description adds genuinely new behavioral detail: .gitignore is respected and node_modules, dist, lockfiles and binaries are skipped, plus the shape of the return (matching lines with surrounding context, workspace-relative paths). It does not mention the remote-vs-local mode distinction or any result truncation behavior.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, zero filler, with the core action front-loaded before the exclusion rules and return shape. Every clause carries information.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
An output schema exists, so return values need not be explained, and annotations plus a fully covered schema leave few gaps. What is missing is the local-vs-remote execution mode implied by the repo parameter and any note on result capping, but the description is otherwise sufficient to call the tool correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so all six parameters are already documented in the schema, including the literal-vs-regex behavior and the repo allowlist rule. The description adds no parameter-level meaning beyond that, so the baseline 3 applies.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ("Search the workspace") and immediately qualifies the mechanism (ripgrep) and scope. An agent can tell this is a full-text code search, clearly distinct from siblings like get_docs, run_tests, or analyze_issue.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
There is no explicit when-to-use, when-not-to-use, or named alternative. The description says what it searches but never routes the agent away from get_docs or get_issue when those would be better. Usage is only implied by the word "workspace".
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
summarize_test_failuresSummarize test failuresARead-onlyIdempotent
Summarize the failures of a previous run_tests call: each failing test with its message, top project stack frames, and likely_source_files (non-test files from the stack — where the fix probably goes), plus groups of failures sharing the same normalized error, largest first.
| Name | Required | Description | Default |
|---|---|---|---|
| run_id | Yes | The run_id returned by run_tests. |
Output Schema
| Name | Required | Description |
|---|---|---|
| groups | Yes | |
| run_id | Yes | |
| failures | Yes | |
| truncated | Yes | |
| workspace | Yes | |
| total_failures | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already declare readOnlyHint=true, idempotentHint=true and destructiveHint=false (and closed-world), so safety is covered. The description adds interpretive context absent from the schema, notably that likely_source_files are non-test files from the stack 'where the fix probably goes' and that failures are grouped by normalized error, largest first — ordering behavior the structured fields do not convey.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
A single dense sentence, front-loaded with the verb and the source of the input, followed by the contents. It is long but every clause carries new information; the parenthetical is slightly heavy but earns its place by explaining field meaning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
With a full output schema present, the description need not enumerate return values in detail, and annotations cover the safety profile. Combined, an agent has everything needed to select and call this tool on a valid run_id.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
There is a single parameter at 100% schema description coverage, with the pattern and origin (run_id returned by run_tests) documented in the schema itself. The description adds only the implicit linkage that the run_id comes from a prior run_tests call, so the baseline 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description gives a specific verb (summarize) and resource (failures of a previous run_tests call), and it enumerates the returned artifacts (failing test, message, top project stack frames, likely_source_files, error groups). It is clearly distinguishable from get_issue, analyze_issue, search_codebase and run_tests.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It anchors usage to a prerequisite — this operates on the output of a previous run_tests call — so an agent knows it must run tests first. There is no explicit when-not-use or named alternative, which keeps it short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
6 tool updates
v1.0.0- First observed
analyze_issue - First observed
get_docs - First observed
get_issue - First observed
run_tests - First observed
search_codebase - First observed
summarize_test_failures
TDQS
Scored across 6 tools
Each tool targets a distinct purpose: get_issue fetches raw issue data, analyze_issue extracts deterministic signals, get_docs reads documentation, search_codebase searches code, run_tests executes tests, and summarize_test_failures analyzes test output. The descriptions make the boundaries clear, and while get_issue and analyze_issue both relate to issues, one is read-only fetching and the other is analysis, so misselection is unlikely.
All tool names use snake_case with a consistent verb_noun pattern (get_issue, analyze_issue, get_docs, search_codebase, run_tests, summarize_test_failures). There are no naming outliers or mixed conventions, making the set predictable and easy to scan.
Six tools is well-scoped for a developer triage and test-diagnosis assistant. Each tool earns its place by covering a distinct step in the workflow, with no redundant or filler tools.
The surface covers issue fetching and analysis, docs reading, code search, and test running plus failure summarization, forming a coherent triage loop. However, there is no way to list or search issues, which is a minor gap for a triage-focused server where an agent may need to discover which issue to work on.
Maintenance
Related MCP Connectors
An MCP server that gives your AI access to the source code and docs of all public github repos
Nifty's MCP server — exposes tasks, projects, messages, and files as tools for AI agents.
MCP server for agentverse documentation, generated by doc2mcp.
Related MCP Servers
- AlicenseBqualityCmaintenanceMCP server that exposes GitHub operations as tools for AI agents, enabling code search, issue management, and PR review.12MIT
- AlicenseNot gradedqualityDmaintenanceAn MCP server that enables AI agents to directly manage GitHub repositories, including PRs, issues, and code search, using natural language.MIT
- AlicenseNot gradedqualityDmaintenanceAn MCP server that provides tools for interacting with the GitHub API, enabling AI assistants to query repositories, pull requests, issues, commits, users, and more.379 npmISC
- AlicenseNot gradedqualityBmaintenanceA local MCP server that gives AI agents access to developer tooling — GitHub (read-only), documentation search, and web research — via stdio transport.MIT