Skip to main content
Glama

DevPilot MCP

A local software engineering runtime for AI coding agents.

DevPilot gives AI agents the ability to understand, run, test, diagnose, review, and validate real software projects through the Model Context Protocol.

Understand → Analyze → Modify → Build → Run → Test → Diagnose → Review → Benchmark → Commit / Rollback

Keywords: MCP · AI Coding Agent · Software Engineering · Local Development · Code Intelligence · Testing · Git · Benchmark


What DevPilot is

DevPilot is the execution and verification layer an agent calls between "I think this fix is right" and "here is the evidence". It provides:

Capability

Meaning

Accurate engineering context

project type, entrypoints, symbols, references, dependency/impact graph

Safe execution environment

workspace-scoped paths, command policy, timeouts, output caps

Reliable verification

build, test, run, failure diagnosis, benchmark, structured evidence

Version control

checkpoints, diff review, guarded rollback that never destroys user work

Related MCP server: MCP Agent Toolkit

What DevPilot is not

Not an IDE, not a VS Code/Cursor clone, not a chat UI, not a Copilot clone, not a code generator, not a cloud service, not "another filesystem MCP".

Filesystem MCP : read file / write file / list directory
GitHub MCP     : remote repo / issue / PR / Actions / remote metadata
DevPilot MCP   : understand project / analyze impact / run / test / diagnose / review / rollback   (local)

DevPilot is an agent backend, never an agent frontend. It does not replace the LLM's reasoning: it supplies context, safety and evidence, and it returns structured results instead of dumping 10,000 log lines on the agent.


Path boundaries (hard rule)

DevPilot itself          D:\tools\DevPilot-MCP          ← source, tests, docs, this Git repo
managed target project   D:\Projects\<project>          ← opened as a Workspace only
per-workspace data       <workspace>\.devpilot\         ← config / cache / logs / checkpoints / db

The three are never mixed:

  • DevPilot source code is never copied into a target project.

  • src/server, src/workspace, src/runner are never created inside a target project.

  • .devpilot/ holds only that workspace's metadata, cache, index, logs and checkpoints.


Requirements

  • Node.js >= 20 (developed on 22.x)

  • Git on PATH (required for checkpoint / diff / rollback)

  • Optional per language: Python + pytest, Maven/Gradle + JDK, npm/pnpm

Quick start

cd D:\tools\DevPilot-MCP
npm install
npm run build
npm test
npm run dev -- serve        # start the MCP server on stdio

CLI (the MCP server is the product; the CLI is for humans, and for isolating a fault from the bridge):

devpilot init [dir]        create <dir>\.devpilot\ (--write-gitignore to also edit .gitignore)
devpilot status [dir]      profile + git snapshot (--json)
devpilot scan [dir]        scan into .devpilot/cache/project.json (--force, --json)
devpilot test [dir]        run the detected suite, printed as parsed counts
                           (--filter=<expr>, --file=<path>, --fail-fast, --json)
devpilot diagnose [dir]    classify the last failed job (--job=<id>, --log=<file>, --max-evidence=<n>)
devpilot doctor [dir]      report the local toolchain; never installs anything (--verbose, --json)
devpilot serve             run the MCP server on stdio (what DeepSeek Harness spawns)
devpilot version | help

Exit codes: 0 success · 1 tests failed / none collected / run unverified · 2 a typed DevPilot error (devpilot <CODE>: message, never a stack trace).

Tool set

19 tools, each returning the same envelope:

workspace   open_workspace   get_workspace_status   close_workspace
understand  scan_project     get_project_map        find_symbol        find_references
analyze     impact_analysis  dependency_audit       doctor
execute     build_project    run_project            run_tests          run_test
diagnose    diagnose_failure
git         get_git_status   create_checkpoint      rollback_checkpoint  review_diff

10 reliable tools beat 40 half-finished ones. Gradle support is secondary; Python, Maven, Node come first. Everything returns one stable envelope:

{
  "success": true,
  "summary": "2 test failures found",
  "data": { "total": 42, "passed": 40, "failed": 2 },
  "artifacts": { "log": ".devpilot/logs/test-20250101-120000.log" },
  "warnings": []
}

Failures return { "success": false, "error": { "code": "TEST_FAILED", "message": "...", "details": {}, "hint": "..." } } with codes an agent can branch on (WORKSPACE_NOT_OPEN, COMMAND_TIMEOUT, BUILD_FAILED, ...).

Design principles

  1. Agent oriented — optimize for an agent calling the tool correctly, not for human CLI comfort.

  2. Local first — no cloud, no account, code/Git/environment data never uploaded.

  3. Safe by default — workspace-restricted paths, command allow/deny policy, timeouts, output limits, file-change limits.

  4. Reversible — checkpoints, diffs, guarded rollback; never git reset --hard over user work.

  5. Evidence based — build output, tests, diff and logs instead of "should work".

  6. No wheel reinvention — Git, Maven, Gradle, pytest, npm are driven as backends. The language parsers sit behind a LanguageParser seam, so Tree-sitter or ripgrep can replace today's lexical extractors without changing a single tool contract.

Documentation

Doc

Content

docs/ARCHITECTURE.md

layering, modules, adapters, extension points, key decisions

docs/DATA-MODEL.md

core entities, TypeScript types, SQLite schema, incremental indexing

docs/TOOLS.md

MCP tool schemas, envelopes, error codes, permission levels

docs/WORKSPACE-LIFECYCLE.md

workspace state machine, .devpilot layout, concurrency, Git safety

docs/ROADMAP.md

Phase 1–10 plan, per-phase gates, V1 acceptance criteria

docs/GATES.md

the gate log: real commands, real output, and the mistakes that were corrected

docs/DSH-INTEGRATION.md

wiring into DeepSeek Harness, and how to reload the entry after a rebuild

docs/VERIFY.md

how to verify DevPilot yourself: four channels, expected output, failure codes, boundaries

CHANGELOG.md

what changed per release, with the evidence for each fix

Status

All ten phases are implemented and gated; V1 is usable, not just startable. Each phase was closed only after a real build and test run passed — docs/GATES.md records the commands, the observed output and the corrections, including one restore mechanism that unit tests blessed and an end-to-end run proved wrong.

tsc -p tsconfig.json --noEmit   clean (strict); src and tests are both typechecked
vitest run                      50 test files / 388 tests green (~20 s)      (2026-10-05)
npm run smoke                   SMOKE PASS, self-hosting: 207 files scanned, 4,089 symbols,
                                13,965 refs; repeated call parsed=0 reused=177
stack acceptance                node 13/13 and maven 13/13, exit 0 each; the Python loop was also
                                run on a real Git project through the DeepSeek Harness bridge
doctor                          Git, Node, Python+pytest, JDK 17 + Maven all detected on this machine

Implemented: strict TS project · typed error model + result envelope · DevPilot home and per-workspace .devpilot\ layout · zod-validated config.yml · security layer (workspace confinement, command policy, change budgets, secret redaction, protected files) · workspace manager (open / status / close, registry, project detection, git snapshot) · ignore-aware file walker (own .gitignore engine) · project scanner + project map · symbol index (Python / Java / TS / JS lexical extractors behind a LanguageParser seam, per-file mtime+size incrementality, SQLite store with a JSON fallback) · process runner with timeouts and output caps · build runner · run runner · test runner (pytest, unittest, Surefire/Gradle, jest, vitest, node:test) with structured parsing · failure diagnosis · git diff review with risk classification · checkpoints and guarded rollback · impact analysis · environment doctor · dependency audit · MCP server with 19 tools · CLI · fixtures and acceptance drivers under tools/.

Not in V1: benchmark, semantic/embedding search, Gradle real-machine acceptance, non-Windows acceptance. docs/VERIFY.md §7 lists the boundaries explicitly, so a green run is never read as more than it is.

License

MIT — see LICENSE.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    F
    maintenance
    A secure, container-based implementation of the Model Context Protocol (MCP) that provides sandboxed environments for AI systems to safely execute code, run commands, access files, and perform web operations.
    22
    Apache 2.0
  • A
    license
    A
    quality
    B
    maintenance
    Enables AI coding agents to safely execute commands, run tests, and modify project files inside disposable, policy-enforced Docker sandboxes that are isolated from the host machine and its credentials.
    15
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Enables AI agents to perform software engineering tasks inside an isolated, deterministic sandbox—exploring repositories, reproducing failures, applying patches, running tests, and verifying solutions against hidden suites through MCP tools.
    8
    MIT