Skip to main content
Glama

DevPilot MCP

Give your coding agent the issue, the code, and the failing tests in one loop.

CI npm node

DevPilot is a Model Context Protocol server (TypeScript, MCP SDK v2) that gives AI coding agents like Claude Code structured, scoped access to GitHub issues, a codebase, its docs, and its test suite. It runs locally over stdio (npx) or remotely over Streamable HTTP.


Why

Most of the time an agent spends on a bug goes to finding context: reading the issue, guessing which files matter, grepping, running the whole test suite, and scrolling through raw output. DevPilot collapses that into six structured, size-capped tool calls:

get_issue → analyze_issue → search_codebase → run_tests → summarize_test_failures → fix

The server makes no LLM calls. Extraction is deterministic, which keeps it cheap, fast, and testable, and the client's model already does the reasoning. See Benchmark for how the time savings are measured.

Related MCP server: GitHub MCP Server

Tools

Tool

Input

Output

get_issue

owner?, repo?, issue_number, max_comments=20

title, state, labels, author, dates, body, comments [{author, body, created_at}], linked PRs, URL. Cached 60 s.

analyze_issue

owner?, repo?, issue_number

type_guess (bug/feature/question/docs), error_messages[], stack_frames[], mentioned_paths[], code_blocks[], suggested_search_queries[]

search_codebase

query, is_regex=false, path_glob?, max_results=30, context_lines=2

matches[{path, line, preview, context_before[], context_after[]}], total_matches, truncated

get_docs

topic?, path?, max_chars=8000

sections[{path, heading, content}], available_docs[]

run_tests

filter?

run_id, exit_code, duration_ms, passed, failed, skipped, timed_out, output_tail

summarize_test_failures

run_id

failures[{test_name, file, line?, message, top_frames[], likely_source_files[]}], groups[] (largest first)

Plus the triage_issue prompt (owner, repo, issue_number), which walks the agent through the whole loop and ends with a proposed fix.

Every tool validates its input with zod (every field is described for the model), declares an outputSchema, returns structured content plus a one-line summary, truncates with an explicit …[truncated N chars] marker, and reports errors as actionable tool errors (Issue #99 not found in owner/repo) rather than stack traces.

owner/repo are optional. Locally they default to the workspace's origin remote, and remotely to the only allowlisted repo. In remote mode with several allowlisted repos, the workspace tools also take repo: "owner/repo".

Quick start

Local (stdio)

Run it inside the repository you want the agent to work on:

claude mcp add devpilot -- npx -y @ashbruh22/devpilot-mcp
# optional, for private repos and higher rate limits:
claude mcp add devpilot -e GITHUB_TOKEN=github_pat_xxx -- npx -y @ashbruh22/devpilot-mcp

Then ask Claude Code to "triage issue #12", or run the prompt directly: /mcp__devpilot__triage_issue <owner> <repo> 12.

Any other MCP client works too:

{ "mcpServers": { "devpilot": { "command": "npx", "args": ["-y", "@ashbruh22/devpilot-mcp"] } } }

Remote (Streamable HTTP)

claude mcp add --transport http devpilot https://devpilot-mcp.onrender.com/mcp
# if the server sets DEVPILOT_API_KEY:
claude mcp add --transport http devpilot https://devpilot-mcp.onrender.com/mcp --header "Authorization: Bearer <key>"

The hosted server works on the demo project bundled in this repo (demo/devpilot-demo), which has two real bugs. Try: "Use devpilot to run the tests, summarize the failures, and propose a fix for the largest group." Or triage one of the demo issues end to end: /mcp__devpilot__triage_issue Ashbruh22 Devpilot_MCP 1.

Architecture

flowchart LR
  C["MCP client<br/>(Claude Code, Inspector)"] <--> S["stdio<br/>src/stdio.ts"]
  C <--> H["Streamable HTTP<br/>src/http.ts<br/>auth · rate limit · CORS"]
  S <--> T["Tool layer<br/>createServer(deps)"]
  H <--> T
  T --> G["GitHub API<br/>(read-only, 60 s cache)"]
  T --> R["ripgrep<br/>(path-guarded)"]
  T --> X["Test runner<br/>(no shell, timeout, tree kill)"]
  • Two transports, one core. createServer(deps) builds an McpServer with every tool registered; it knows nothing about transports. stdio.ts serves it with serveStdio, and http.ts serves it statelessly with createMcpHandler + toNodeHandler on Express, building one server per request. Long-lived state (GitHub cache, run store, workspaces) lives in deps and is shared.

  • Token budget by design. Outputs are structured, every string is capped, search results shrink to fit MAX_OUTPUT_CHARS, test output keeps only the tail, and failures are grouped by normalized error message (numbers, strings, and paths become placeholders), so 40 failures with one root cause read as one group.

  • likely_source_files are the non-test project files in each failure's stack. When the stack only reaches the test file (typical for assertion failures), DevPilot falls back to naming: src/pricing.test.ts → src/pricing.ts.

src/
├─ server.ts          createServer(deps) + createDeps(config)
├─ stdio.ts           local entry (bin)
├─ http.ts            remote entry: Express, Streamable HTTP, landing page, /healthz
├─ config.ts          env parsing with zod, fail-fast
├─ tools/             one file per tool
├─ prompts/           triage_issue
├─ lib/               github, workspace, pathGuard, exec, runStore, truncate, search, docs, failures, …
│  └─ parsers/        vitest/jest JSON, pytest JUnit XML, raw output, stack traces (JS + Python)
└─ public/index.html  landing page

Security model

Running tests on a public server is remote code execution by design, so remote mode is locked down:

  1. Allowlist only. Tools accept only repos in ALLOWED_REPOS. At boot, each is shallow-cloned (git clone --depth 1 --branch <ref>) into /tmp/workspaces/<owner>__<repo>, and its dependencies are installed once with npm ci --ignore-scripts.

  2. Fixed commands. Test commands come from TEST_COMMANDS, never from tool input. filter must match ^[\w\-./ ]{1,100}$, can't start with - (so it can't inject flags), and can't contain ... Commands run without a shell, and configured commands may not contain shell operators.

  3. Resource limits. Each run has a TEST_TIMEOUT_MS timeout that kills the whole process group, output is capped, and each repo allows at most one concurrent test run ("busy, retry in N s"). /mcp is rate-limited per IP (60/min by default) and request bodies are capped at 1 MB.

  4. Least privilege. The container runs as the non-root node user with tini as PID 1. The GitHub token should be a fine-grained, read-only PAT for public repos. GITHUB_TOKEN, DEVPILOT_API_KEY, and anything that looks like a secret (*TOKEN*, *SECRET*, *PASSWORD*, *API_KEY*, …) are stripped from child-process environments.

  5. Path guard. Every client-supplied path is resolved with path.resolve, must stay under the workspace root, and is re-checked with realpath so symlinks can't escape. Absolute paths, .., and NUL bytes are rejected. Search results go through the same check.

  6. No arbitrary network fetches. No tool fetches a URL from its input. GitHub is the only outbound API, and clone URLs are built from the allowlist.

The HTTP server also uses the SDK's Host-header validation (DNS-rebinding protection) and Origin validation on localhost binds, an optional constant-time bearer-token check on /mcp (DEVPILOT_API_KEY), and CORS that exposes only the MCP headers. Without a key, the server runs as a public read-only demo limited to the allowlist.

Configuration

All configuration comes from environment variables, validated at startup. See .env.example.

Var

Mode

Default

Purpose

DEVPILOT_MODE

both

local

local or remote

GITHUB_TOKEN

both

–

Fine-grained PAT. Read-only, public repos only for the hosted deployment

WORKSPACE_ROOT

local

cwd

Repo to operate on

TEST_COMMAND

local

auto-detect

Override the detected test command

ALLOWED_REPOS

remote

– (required)

owner/repo[@ref][:subdir],…

TEST_COMMANDS

remote

{}

JSON {"owner/repo": "command"}

DEVPILOT_API_KEY

remote

–

Bearer token for /mcp. Unset means a public demo

PORT / HOST

remote

3000 / 0.0.0.0

Listen address (127.0.0.1 in local mode)

ALLOWED_HOSTS

remote

Render hostname

Host-header allowlist for public binds

RATE_LIMIT_PER_MIN

remote

60

Per-IP limit on /mcp

TEST_TIMEOUT_MS

both

120000

Test run timeout

MAX_OUTPUT_CHARS

both

20000

Cap on any tool's text response

WORKSPACES_DIR / INSTALL_DEPS

remote

/tmp/workspaces / true

Where to clone, and whether to install deps

Test runner auto-detection (local mode) prefers machine-readable output: vitest → vitest run --reporter=json, jest → jest --json, pytest → pytest --junitxml=… -q, otherwise npm test with best-effort text parsing. Binaries are resolved from node_modules/.bin (walking up for monorepos) or run with npx --no, so nothing is ever downloaded.

Development

npm install
npm run typecheck && npm run lint && npm test   # unit + integration tests
npm run build                                   # tsc → dist/, keeps the shebang, chmod +x
npm run dev:stdio                               # run from source over stdio
npm run dev:http                                # run from source over HTTP on :3000
npm run inspector                               # MCP Inspector against dist/stdio.js
npx @modelcontextprotocol/inspector             # then connect to http://localhost:3000/mcp

The tests cover:

  • Unit: path guard (traversal, absolute paths, symlink escape), parsers (vitest/jest JSON, pytest JUnit XML, raw output), stack traces (Node, browser, Python), issue analysis, truncation, exec (tree kill on timeout, secret stripping, no shell), config validation.

  • Integration: a real MCP client (@modelcontextprotocol/client) over in-memory and Streamable HTTP transports. It calls every tool against test/fixtures/sample-repo (real vitest runs with real failures) with GitHub mocked by nock. HTTP tests cover /healthz, 401 without a token, 429 rate limiting, CORS, Host/Origin validation, and body limits. Remote-mode tests clone from a local bare git repo and check allowlist enforcement and fixed commands.

Deployment

Docker → Render. The repo includes a multi-stage Dockerfile (non-root, tini, git) and a render.yaml blueprint:

  1. In Render, go to New → Blueprint and pick this repo. Render reads render.yaml, builds the Docker image, and creates a free web service.

  2. When prompted, set GITHUB_TOKEN: a fine-grained PAT with read-only access to public repositories. It's needed for get_issue/analyze_issue, since unauthenticated calls share a 60/hour limit. Leave DEVPILOT_API_KEY empty for a public demo, or set it to require a bearer token.

  3. On boot, the server shallow-clones this repo, uses demo/devpilot-demo as its workspace (ALLOWED_REPOS=Ashbruh22/Devpilot_MCP:demo/devpilot-demo), and installs that project's dependencies. /healthz reports "ready" after a few seconds. Render's hostname is added to the Host allowlist automatically.

  4. Connect with claude mcp add --transport http devpilot https://devpilot-mcp.onrender.com/mcp.

  5. Free instances sleep when idle: the first request can take 30–60 s. To keep it warm, set the repository variable DEVPILOT_URL to the service URL, which enables the keepalive workflow (a 10-minute ping). Check Render's current free-tier terms first.

npm. Tag a release (git tag v1.0.0 && git push --tags) and the release workflow publishes with provenance (it needs an NPM_TOKEN secret). Or run npm publish --access public by hand, then check it with npx -y @ashbruh22/devpilot-mcp in a clean folder.

Benchmark

bench/ has a reproducible method: 10 tasks on the demo repo ("find the root cause of issue #N and the file to change"), each timed once without DevPilot (the agent reads files and runs commands itself) and once with the triage_issue flow. It records time-to-context and tool calls or steps. Results go in bench/results.csv, and npm run bench:summary prints the medians.

Status: the harness and tasks are in place. The results table is empty until the runs are done, so no time-saved number is claimed yet. Quote only the measured median.

Available Tools

6 tools
analyze_issueAnalyze GitHub issueA
Read-onlyIdempotent

Deterministically extract triage signals from an issue (no LLM): a type guess (bug/feature/question/docs), error messages, parsed stack frames (JS and Python), mentioned file paths, fenced code blocks, and suggested queries to pass to search_codebase.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoNoGitHub repository name. Optional, see "owner".
ownerNoGitHub owner (user or org). Optional: defaults to the workspace's "origin" remote (local mode) or the only allowlisted repo (remote mode).
issue_numberYesIssue (or PR) number, e.g. 42.

Output Schema

ParametersJSON Schema
NameRequiredDescription
repoYes
ownerYes
titleYes
type_guessYes
code_blocksYes
issue_numberYes
stack_framesYes
type_signalsYes
error_messagesYes
mentioned_pathsYes
suggested_search_queriesYes

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, and openWorld, so the safety profile is covered. The description adds a genuine behavioral trait beyond that: the extraction is 'deterministic' and uses 'no LLM', telling the agent results are stable and reproducible. It doesn't discuss rate limits or failure modes, but the determinism claim is real added value.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence that leads with the verb and the determinism guarantee, then enumerates outputs. The enumeration is long but each item is informative and non-redundant; only the parenthetical nesting makes it slightly dense.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need no explanation, yet the description helpfully names the signal categories anyway. Combined with annotations covering the safety profile and 100% schema coverage, an agent has everything needed to invoke it correctly, though nothing is said about behavior on invalid issue numbers or missing repos.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so owner, repo, and issue_number are already fully documented in the schema, including the origin-remote fallback behavior. The description adds no parameter-level syntax or format guidance, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('extract triage signals from an issue') and enumerates the exact outputs (type guess, error messages, stack frames, file paths, code blocks, suggested queries). The parenthetical '(no LLM)' plus the enumerated artifacts make it impossible to confuse with get_issue or get_docs.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly signals the workflow context by noting the suggested queries are meant 'to pass to search_codebase', giving the agent a downstream routing cue. It stops short of stating when NOT to use it (e.g., vs. get_issue for raw issue body retrieval), so it is clear context without explicit exclusions.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_docsGet project docsA
Read-onlyIdempotent

Read project documentation (README*, CONTRIBUTING*, root .md, docs/**/.md). With "path": return that file, or just the section whose heading matches "topic". With only "topic": return the 3 best-matching sections ranked by keyword hits. With neither: return the README intro and the list of docs.

ParametersJSON Schema
NameRequiredDescriptionDefault
pathNoA specific doc file relative to the workspace root, e.g. "docs/setup.md".
repoNoRemote mode only: which allowlisted repo ("owner/repo") to operate on. Optional when only one repo is allowlisted; ignored in local mode.
topicNoKeywords to look for, e.g. "discount rounding" or "configuration".
max_charsNoTotal character budget for returned section content. Default 8000.

Output Schema

ParametersJSON Schema
NameRequiredDescription
sectionsYes
workspaceYes
available_docsYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover safety (readOnly, idempotent, non-destructive, closed world), so the bar is lower. The description nonetheless adds real behavioral detail: the fallback behavior when no arguments are supplied, the ranked-by-keyword-hits selection, and the fixed cap of 3 sections for topic-only calls. It doesn't discuss truncation interplay with max_chars in any depth, which is the only notable gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences, front-loaded with the resource scope, then the three invocation modes in a predictable path/topic/neither progression. Every clause carries a distinct constraint; nothing is filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value documentation is not required, yet the description still specifies what each mode yields. With 4 optional parameters, zero required, and full schema coverage, an agent has everything needed to call it correctly in any mode.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description goes beyond the per-field schema text by explaining the interaction semantics: path alone returns the file or just the heading-matched section when topic is also present, and topic alone switches to ranked multi-section retrieval. That combinatorial meaning is not derivable from the schema fields individually.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (read project documentation) and enumerates the exact file patterns covered (README*, CONTRIBUTING*, root *.md, docs/**/*.md), which is unusually precise. It does not, however, explicitly contrast itself with the sibling search_codebase, so the boundary between doc reading and code search is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit mode-selection rules: path+provision to fetch a file or matching section, topic only to get the 3 best keyword-ranked sections, neither to get the README intro plus docs list. That is clear operational guidance for invoking the tool. It stops short of naming when a sibling (e.g. search_codebase) would be preferable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_issueGet GitHub issueA
Read-onlyIdempotent

Fetch a GitHub issue with its discussion: title, state, labels, author, dates, body, recent comments, and linked pull requests. Read-only. Start here when triaging an issue.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoNoGitHub repository name. Optional, see "owner".
ownerNoGitHub owner (user or org). Optional: defaults to the workspace's "origin" remote (local mode) or the only allowlisted repo (remote mode).
issue_numberYesIssue (or PR) number, e.g. 42.
max_commentsNoMaximum number of comments to return (most recent first kept). Default 20.

Output Schema

ParametersJSON Schema
NameRequiredDescription
urlYes
bodyYes
repoYes
ownerYes
stateYes
titleYes
authorYes
labelsYes
numberYes
commentsYes
created_atYes
linked_prsYes
updated_atYes
total_commentsYes
is_pull_requestYes

TDQS

A4.1/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, openWorldHint=true, and destructiveHint=false. The description repeats 'Read-only' and lists returned content, but adds no behavioral context such as auth needs, rate limits, or failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the action and resource, followed by scope and usage cue. No wasted wording.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With rich annotations, 100% schema description coverage, and an output schema present, the description supplies enough context for correct invocation. It does not need to explain return values further.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all four parameters are already documented in the schema. The description adds no parameter-level meaning beyond what the schema provides, making 3 the correct baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: fetch a GitHub issue with its discussion. It also lists the returned fields and positions the tool as the starting point for triage, distinguishing it from sibling analyze_issue.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly says 'Start here when triaging an issue,' giving clear usage context. It does not name when not to use it or explicitly route to alternatives like analyze_issue, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

run_testsRun testsA

Run the project's test suite (auto-detects vitest, jest, pytest, or npm test; in remote mode only the server's fixed, configured command runs). Returns pass/fail counts and a run_id; pass the run_id to summarize_test_failures for grouped failures and the source files to fix.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoNoRemote mode only: which allowlisted repo ("owner/repo") to operate on. Optional when only one repo is allowlisted; ignored in local mode.
filterNoOptional test name or file pattern, e.g. "pricing" or "src/pricing.test.ts". Letters, digits, _ - . / and spaces only.

Output Schema

ParametersJSON Schema
NameRequiredDescription
failedYes
passedYes
run_idYes
runnerYes
commandYes
skippedYes
exit_codeYes
timed_outYes
workspaceYes
duration_msYes
output_tailYes
structured_reportYes

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=false, openWorldHint=false, idempotentHint=false, and destructiveHint=false. The description adds valuable context beyond annotations: it auto-detects vitest/jest/pytest/npm test, notes that remote mode runs only a fixed configured command, and returns pass/fail counts plus a run_id. It does not contradict any annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, front-loaded sentence that covers purpose, auto-detection, remote-mode behavior, and return-value routing without waste. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a test-running tool with an output schema and annotations, the description supplies the missing operational context: what gets auto-detected, how remote mode differs, and what the run_id is for. Nothing an agent needs to call or use the tool correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents both 'repo' and 'filter' thoroughly. The description does not add parameter-level semantics beyond what the schema provides, which matches the baseline 3 for high-coverage schemas.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: 'Run the project's test suite', and names the auto-detected frameworks. It also distinguishes itself from the sibling summarize_test_failures by explaining the run_id handoff, so an agent can tell the tools apart without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides clear workflow guidance: use run_tests to execute the suite, then pass the returned run_id to summarize_test_failures for grouped failures. It does not explicitly state when not to use this tool or contrast it with other siblings like search_codebase, but the primary usage context is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

search_codebaseSearch codebaseA
Read-onlyIdempotent

Search the workspace with ripgrep (respects .gitignore; skips node_modules, dist, lockfiles and binaries). Returns matching lines with surrounding context and paths relative to the workspace root.

ParametersJSON Schema
NameRequiredDescriptionDefault
repoNoRemote mode only: which allowlisted repo ("owner/repo") to operate on. Optional when only one repo is allowlisted; ignored in local mode.
queryYesText to search for. Treated literally unless is_regex is true.
is_regexNoInterpret query as a (Rust-flavored) regular expression.
path_globNoRestrict to files matching this glob, relative to the root, e.g. "src/**/*.ts" or "*.py".
max_resultsNoMaximum matches to return. Default 30.
context_linesNoLines of context before and after each match. Default 2.

Output Schema

ParametersJSON Schema
NameRequiredDescription
matchesYes
truncatedYes
workspaceYes
total_matchesYes
files_with_matchesYes

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly, idempotent, non-destructive, closed-world, so the safety profile is covered. The description adds genuinely new behavioral detail: .gitignore is respected and node_modules, dist, lockfiles and binaries are skipped, plus the shape of the return (matching lines with surrounding context, workspace-relative paths). It does not mention the remote-vs-local mode distinction or any result truncation behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, zero filler, with the core action front-loaded before the exclusion rules and return shape. Every clause carries information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained, and annotations plus a fully covered schema leave few gaps. What is missing is the local-vs-remote execution mode implied by the repo parameter and any note on result capping, but the description is otherwise sufficient to call the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all six parameters are already documented in the schema, including the literal-vs-regex behavior and the repo allowlist rule. The description adds no parameter-level meaning beyond that, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ("Search the workspace") and immediately qualifies the mechanism (ripgrep) and scope. An agent can tell this is a full-text code search, clearly distinct from siblings like get_docs, run_tests, or analyze_issue.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use, when-not-to-use, or named alternative. The description says what it searches but never routes the agent away from get_docs or get_issue when those would be better. Usage is only implied by the word "workspace".

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

summarize_test_failuresSummarize test failuresA
Read-onlyIdempotent

Summarize the failures of a previous run_tests call: each failing test with its message, top project stack frames, and likely_source_files (non-test files from the stack — where the fix probably goes), plus groups of failures sharing the same normalized error, largest first.

ParametersJSON Schema
NameRequiredDescriptionDefault
run_idYesThe run_id returned by run_tests.

Output Schema

ParametersJSON Schema
NameRequiredDescription
groupsYes
run_idYes
failuresYes
truncatedYes
workspaceYes
total_failuresYes

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true and destructiveHint=false (and closed-world), so safety is covered. The description adds interpretive context absent from the schema, notably that likely_source_files are non-test files from the stack 'where the fix probably goes' and that failures are grouped by normalized error, largest first — ordering behavior the structured fields do not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single dense sentence, front-loaded with the verb and the source of the input, followed by the contents. It is long but every clause carries new information; the parenthetical is slightly heavy but earns its place by explaining field meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With a full output schema present, the description need not enumerate return values in detail, and annotations cover the safety profile. Combined, an agent has everything needed to select and call this tool on a valid run_id.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

There is a single parameter at 100% schema description coverage, with the pattern and origin (run_id returned by run_tests) documented in the schema itself. The description adds only the implicit linkage that the run_id comes from a prior run_tests call, so the baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb (summarize) and resource (failures of a previous run_tests call), and it enumerates the returned artifacts (failing test, message, top project stack frames, likely_source_files, error groups). It is clearly distinguishable from get_issue, analyze_issue, search_codebase and run_tests.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It anchors usage to a prerequisite — this operates on the output of a previous run_tests call — so an agent knows it must run tests first. There is no explicit when-not-use or named alternative, which keeps it short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 6 tool updatesv1.0.0
    • First observedanalyze_issue
    • First observedget_docs
    • First observedget_issue
    • First observedrun_tests
    • First observedsearch_codebase
    • First observedsummarize_test_failures

TDQS

A4.2/5.0

Scored across 6 tools

Disambiguation5/5

Each tool targets a distinct purpose: get_issue fetches raw issue data, analyze_issue extracts deterministic signals, get_docs reads documentation, search_codebase searches code, run_tests executes tests, and summarize_test_failures analyzes test output. The descriptions make the boundaries clear, and while get_issue and analyze_issue both relate to issues, one is read-only fetching and the other is analysis, so misselection is unlikely.

Naming Consistency5/5

All tool names use snake_case with a consistent verb_noun pattern (get_issue, analyze_issue, get_docs, search_codebase, run_tests, summarize_test_failures). There are no naming outliers or mixed conventions, making the set predictable and easy to scan.

Tool Count5/5

Six tools is well-scoped for a developer triage and test-diagnosis assistant. Each tool earns its place by covering a distinct step in the workflow, with no redundant or filler tools.

Completeness4/5

The surface covers issue fetching and analysis, docs reading, code search, and test running plus failure summarization, forming a coherent triage loop. However, there is no way to list or search issues, which is a minor gap for a triage-focused server where an agent may need to discover which issue to work on.

Maintenance

ActivityMaintained
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    An MCP server that enables AI agents to directly manage GitHub repositories, including PRs, issues, and code search, using natural language.
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    An MCP server that provides tools for interacting with the GitHub API, enabling AI assistants to query repositories, pull requests, issues, commits, users, and more.
    379 npm
    ISC
  • A
    license
    Not graded
    quality
    B
    maintenance
    A local MCP server that gives AI agents access to developer tooling — GitHub (read-only), documentation search, and web research — via stdio transport.
    MIT