DevPilot MCP
Drives a local Git repository as the version-control backend for a workspace: git status snapshots, checkpoint creation, diff review with risk classification, and guarded rollback that never destroys uncommitted user work.
Supports Gradle as a (secondary) build and test backend for JVM projects, running the build and the compiled test suite and parsing the results into a structured envelope.
Runs Jest test suites for Node/JS/TS projects and parses the output into structured counts and failure details, with optional filters, per-file selection and fail-fast.
Runs pytest test suites for Python projects and parses the output into structured pass/fail counts and evidence, with optional filters, per-file selection and fail-fast.
Detects and analyzes local Python projects: project profiling, ignore-aware file scanning, symbol/reference indexing via a Python lexical extractor, and running their test suites through pytest/unittest.
Runs Vitest test suites for JS/TS projects and parses the output into structured results and failure evidence.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@DevPilot MCPrun the tests and diagnose any failures"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
DevPilot MCP
A local software engineering runtime for AI coding agents.
DevPilot gives AI agents the ability to understand, run, test, diagnose, review, and validate real software projects through the Model Context Protocol.
Understand → Analyze → Modify → Build → Run → Test → Diagnose → Review → Benchmark → Commit / RollbackKeywords: MCP · AI Coding Agent · Software Engineering · Local Development · Code Intelligence · Testing · Git · Benchmark
What DevPilot is
DevPilot is the execution and verification layer an agent calls between "I think this fix is right" and "here is the evidence". It provides:
Capability | Meaning |
Accurate engineering context | project type, entrypoints, symbols, references, dependency/impact graph |
Safe execution environment | workspace-scoped paths, command policy, timeouts, output caps |
Reliable verification | build, test, run, failure diagnosis, benchmark, structured evidence |
Version control | checkpoints, diff review, guarded rollback that never destroys user work |
Related MCP server: MCP Agent Toolkit
What DevPilot is not
Not an IDE, not a VS Code/Cursor clone, not a chat UI, not a Copilot clone, not a code generator, not a cloud service, not "another filesystem MCP".
Filesystem MCP : read file / write file / list directory
GitHub MCP : remote repo / issue / PR / Actions / remote metadata
DevPilot MCP : understand project / analyze impact / run / test / diagnose / review / rollback (local)DevPilot is an agent backend, never an agent frontend. It does not replace the LLM's reasoning: it supplies context, safety and evidence, and it returns structured results instead of dumping 10,000 log lines on the agent.
Path boundaries (hard rule)
DevPilot itself D:\tools\DevPilot-MCP ← source, tests, docs, this Git repo
managed target project D:\Projects\<project> ← opened as a Workspace only
per-workspace data <workspace>\.devpilot\ ← config / cache / logs / checkpoints / dbThe three are never mixed:
DevPilot source code is never copied into a target project.
src/server,src/workspace,src/runnerare never created inside a target project..devpilot/holds only that workspace's metadata, cache, index, logs and checkpoints.
Requirements
Node.js >= 20 (developed on 22.x)
Git on
PATH(required for checkpoint / diff / rollback)Optional per language: Python + pytest, Maven/Gradle + JDK, npm/pnpm
Quick start
cd D:\tools\DevPilot-MCP
npm install
npm run build
npm test
npm run dev -- serve # start the MCP server on stdioCLI (the MCP server is the product; the CLI is for humans, and for isolating a fault from the bridge):
devpilot init [dir] create <dir>\.devpilot\ (--write-gitignore to also edit .gitignore)
devpilot status [dir] profile + git snapshot (--json)
devpilot scan [dir] scan into .devpilot/cache/project.json (--force, --json)
devpilot test [dir] run the detected suite, printed as parsed counts
(--filter=<expr>, --file=<path>, --fail-fast, --json)
devpilot diagnose [dir] classify the last failed job (--job=<id>, --log=<file>, --max-evidence=<n>)
devpilot doctor [dir] report the local toolchain; never installs anything (--verbose, --json)
devpilot serve run the MCP server on stdio (what DeepSeek Harness spawns)
devpilot version | helpExit codes: 0 success · 1 tests failed / none collected / run unverified · 2 a typed
DevPilot error (devpilot <CODE>: message, never a stack trace).
Tool set
19 tools, each returning the same envelope:
workspace open_workspace get_workspace_status close_workspace
understand scan_project get_project_map find_symbol find_references
analyze impact_analysis dependency_audit doctor
execute build_project run_project run_tests run_test
diagnose diagnose_failure
git get_git_status create_checkpoint rollback_checkpoint review_diff10 reliable tools beat 40 half-finished ones. Gradle support is secondary; Python, Maven, Node come first. Everything returns one stable envelope:
{
"success": true,
"summary": "2 test failures found",
"data": { "total": 42, "passed": 40, "failed": 2 },
"artifacts": { "log": ".devpilot/logs/test-20250101-120000.log" },
"warnings": []
}Failures return { "success": false, "error": { "code": "TEST_FAILED", "message": "...", "details": {}, "hint": "..." } }
with codes an agent can branch on (WORKSPACE_NOT_OPEN, COMMAND_TIMEOUT, BUILD_FAILED, ...).
Design principles
Agent oriented — optimize for an agent calling the tool correctly, not for human CLI comfort.
Local first — no cloud, no account, code/Git/environment data never uploaded.
Safe by default — workspace-restricted paths, command allow/deny policy, timeouts, output limits, file-change limits.
Reversible — checkpoints, diffs, guarded rollback; never
git reset --hardover user work.Evidence based — build output, tests, diff and logs instead of "should work".
No wheel reinvention — Git, Maven, Gradle, pytest, npm are driven as backends. The language parsers sit behind a
LanguageParserseam, so Tree-sitter or ripgrep can replace today's lexical extractors without changing a single tool contract.
Documentation
Doc | Content |
layering, modules, adapters, extension points, key decisions | |
core entities, TypeScript types, SQLite schema, incremental indexing | |
MCP tool schemas, envelopes, error codes, permission levels | |
workspace state machine, | |
Phase 1–10 plan, per-phase gates, V1 acceptance criteria | |
the gate log: real commands, real output, and the mistakes that were corrected | |
wiring into DeepSeek Harness, and how to reload the entry after a rebuild | |
how to verify DevPilot yourself: four channels, expected output, failure codes, boundaries | |
what changed per release, with the evidence for each fix |
Status
All ten phases are implemented and gated; V1 is usable, not just startable. Each phase was closed only after a real build and test run passed — docs/GATES.md records the commands, the observed output and the corrections, including one restore mechanism that unit tests blessed and an end-to-end run proved wrong.
tsc -p tsconfig.json --noEmit clean (strict); src and tests are both typechecked
vitest run 50 test files / 388 tests green (~20 s) (2026-10-05)
npm run smoke SMOKE PASS, self-hosting: 207 files scanned, 4,089 symbols,
13,965 refs; repeated call parsed=0 reused=177
stack acceptance node 13/13 and maven 13/13, exit 0 each; the Python loop was also
run on a real Git project through the DeepSeek Harness bridge
doctor Git, Node, Python+pytest, JDK 17 + Maven all detected on this machineImplemented: strict TS project · typed error model + result envelope · DevPilot home and
per-workspace .devpilot\ layout · zod-validated config.yml · security layer (workspace
confinement, command policy, change budgets, secret redaction, protected files) · workspace manager
(open / status / close, registry, project detection, git snapshot) · ignore-aware file walker
(own .gitignore engine) · project scanner + project map · symbol index (Python / Java / TS / JS
lexical extractors behind a LanguageParser seam, per-file mtime+size incrementality, SQLite
store with a JSON fallback) · process runner with timeouts and output caps · build runner · run
runner · test runner (pytest, unittest, Surefire/Gradle, jest, vitest, node:test) with structured
parsing · failure diagnosis · git diff review with risk classification · checkpoints and guarded
rollback · impact analysis · environment doctor · dependency audit · MCP server with 19 tools ·
CLI · fixtures and acceptance drivers under tools/.
Not in V1: benchmark, semantic/embedding search, Gradle real-machine acceptance, non-Windows acceptance. docs/VERIFY.md §7 lists the boundaries explicitly, so a green run is never read as more than it is.
License
MIT — see LICENSE.
This server cannot be deployed
Maintenance
Related MCP Connectors
Context and control for any AI. Agent identity, scoped access, and proof across tools.
Persistent cloud workspaces for AI agents: run commands, edit files, use git and a browser.
Cross-agent artifact workspace with provenance across Claude Code, Codex, Cursor, LangGraph.
Codebase graphs, caller impact analysis, and recorded project context for AI coding agents.
Related MCP Servers
- AlicenseNot gradedqualityFmaintenanceA secure, container-based implementation of the Model Context Protocol (MCP) that provides sandboxed environments for AI systems to safely execute code, run commands, access files, and perform web operations.22Apache 2.0
- AlicenseNot gradedqualityBmaintenanceEnables AI agents to securely interact with local files, live internet search, databases, and development tools through the Model Context Protocol, turning them into autonomous production-ready assistants.5MIT
- AlicenseAqualityBmaintenanceEnables AI coding agents to safely execute commands, run tests, and modify project files inside disposable, policy-enforced Docker sandboxes that are isolated from the host machine and its credentials.15MIT
- AlicenseAqualityBmaintenanceEnables AI agents to perform software engineering tasks inside an isolated, deterministic sandbox—exploring repositories, reproducing failures, applying patches, running tests, and verifying solutions against hidden suites through MCP tools.8MIT