DevTwin MCP
# DevTwin MCP
**Give AI coding agents a live, structured understanding of your local development environment.**
DevTwin is a Model Context Protocol (MCP) server that answers one central question for an AI coding agent: *why is this developer's environment different, broken, or unhealthy?*
It detects project technology, checks installed runtime versions against what a project actually requires, inspects dependency and lockfile state, finds required local services (Postgres, Redis, ...) and whether they're running, checks ports and Git state, and turns all of that into structured, evidence-based diagnostics -- without ever sending your environment to a cloud backend, and without ever exposing secret values to the model.
## Quick Start (2 minutes)
**1. Install** (pick one):
```bash
uv pip install devtwin-mcp # fastest
# or
pipx install devtwin-mcp # simplest
# or
pip install devtwin-mcp # standard
```
**2. Add to your MCP client** (Claude Code, Claude Desktop, Cursor, etc.):
```json
{
"mcpServers": {
"devtwin": {
"command": "devtwin"
}
}
}
```
**3. Try it:**
```
Ask Claude: "Check my development environment"
or "Why does `npm test` fail?"
```
**That's it.** Next time you ask Claude about your project, it'll have access to real environment data instead of guessing.
### Why use DevTwin instead of just asking Claude to run bash commands?
**The problem:** Claude can run bash, but developers lose:
- **Secrets** — environment variable reads leak API keys and passwords into conversation
- **Parsing hell** — Claude has to guess the right 5–6 commands to run, parsing messy output each time
- **No safety** — anything goes, even destructive commands
- **Inconsistent** — every project handles environment checks differently (or not at all)
**DevTwin solves it:**
- **No secrets ever leak** — environment checks report presence only, values never read
- **One call, one answer** — `dev_health()` bundles 10+ checks into structured JSON
- **Safety by design** — only allowlisted operations; the three tools that run anything run only commands DevTwin recognized itself
- **Same check every time** — exact same detection logic across all projects
**Token efficiency:** ~800 tokens (one call + schema) vs. ~1500–2000 tokens (5–6 bash commands + parsing).
**See the side-by-side comparison:** [MCP vs raw Claude](https://claude.ai/code/artifact/4a85ce00-e3c0-45c7-8fe3-4dcf61b75ff8)
## What makes DevTwin different
| Feature | Raw Bash | DevTwin |
|---------|----------|---------|
| Safety | Any command possible | Only safe, allowlisted checks |
| Secrets | Risk of leaking keys/passwords | Never reads or returns secret values |
| Consistency | Different per project | Same checks across all projects |
| Tokens | 5–6 commands, ~1500–2000 tokens | 1 call, ~800–1200 tokens (after schema tax) |
| Works in Claude Code | ✓ (has shell) | ✓ (any MCP client) |
| Works in Claude Desktop | ✗ (no shell) | ✓ |
| Works in Cursor / Windsurf | (conditional) | ✓ |
## Contents
- [Why DevTwin exists](#why-devtwin-exists)
- [FAQ: Claude CLI already has a shell, so why an MCP at all?](#faq-claude-cli-already-has-a-shell-so-why-an-mcp-at-all)
- [Benefits](#benefits)
- [Token cost](#token-cost)
- [Honest tradeoffs](#honest-tradeoffs)
- [With vs. without DevTwin: a worked example](#with-vs-without-devtwin-a-worked-example)
- [Example questions this unlocks](#example-questions-this-unlocks)
- [Per-language examples](#per-language-examples)
- [Architecture](#architecture)
- [Supported ecosystems](#supported-ecosystems)
- [Installation](#installation)
- [MCP client configuration](#mcp-client-configuration)
- [Using it on another project](#using-it-on-another-project-for-other-developers)
- [Tool reference](#tool-reference)
- [Security model](#security-model)
- [Privacy model](#privacy-model)
- [Local-first architecture](#local-first-architecture)
- [Adoption & team setup](#adoption--team-setup)
- [Development](#development)
- [Contributing](#contributing)
- [Roadmap](#roadmap)
- [License](#license)
## Why DevTwin exists
- AI coding agents read code well, but are blind to the environment that
code actually runs in.
- "Why does `npm test` fail on my machine?" usually has nothing to do with
the code -- it's a Node version mismatch, a service that isn't running,
or dependencies that were never installed.
- DevTwin gives an agent the same signal a senior engineer would gather by
hand -- `node --version`, `git status`, `lsof -i :5432`, `docker ps` --
as structured tool calls instead of guesswork.
## FAQ: Claude CLI already has a shell, so why an MCP at all?
This is usually the first question a developer asks, and it's a fair one.
In a client like Claude Code that already has a Bash tool, you can just
ask it to run `node --version`, `docker ps`, `lsof -i :5432`, etc.
directly -- no MCP server required. **The gap DevTwin closes isn't "can
this be done at all" -- it's these:**
| Without DevTwin (raw Bash) | With DevTwin |
|---|---|
| The agent *can* run anything, including destructive commands, even unintentionally. | Zero arbitrary execution -- a fixed allowlist of read-only/safe checks only. See [Security model](#security-model). |
| Picks a different investigation each session; can miss ecosystem edge cases (Gradle wrapper vs. system Gradle, `.nvmrc` vs. `package.json` engines). | The same curated, tested check every time, for every ecosystem. |
| A command like `cat .env` can pull a real secret value straight into the conversation. | Structurally never returns secret values -- presence/absence only. See [Privacy model](#privacy-model). |
| Only works in clients that have a shell tool at all (not Claude Desktop, some IDE plugins). | Works in any MCP client, shell or no shell. |
| ~6 separate round-trips to diagnose one failure; ~1500–2000 tokens per check (5–6 bash commands, scattered output). | 1 call; ~800–1200 tokens per check. See the [worked example](#with-vs-without-devtwin-a-worked-example). |
**Honest answer for Claude CLI specifically:** since it already has Bash,
DevTwin's win there is smaller than "capability you didn't have" -- it's
safety guarantees and consistent, structured output, not brand-new
access. That's also why it isn't free -- see [Token cost](#token-cost)
for what connecting it actually costs, and when it's worth it.
**A few more questions worth asking before adopting this:**
**"Isn't this just a `doctor` script (`make doctor`, `bin/setup`) with extra
steps?"** Conceptually, yes -- plenty of mature repos already hand-write
one. DevTwin's difference is that most repos *don't* have one, writing a
good one per-ecosystem is real work, its output is structured JSON an
agent can reason over rather than plain text a human reads, and the same
13 tools work identically across every repo instead of a bespoke script
per project with its own conventions and blind spots.
**"Does this only work with Claude / Claude Code?"** No. DevTwin speaks
the standard Model Context Protocol -- any MCP-compatible client (Claude
Desktop, Cursor, Windsurf, etc.) can connect to it the same way. Nothing
about it is Claude-specific.
**"Is this safe to depend on -- is it actively maintained?"** It's
[Alpha status](pyproject.toml) and a young project -- read the code (it's
short) before trusting it in a workflow you depend on, same as you would
any new dev-tooling dependency.
**"Could it suggest something wrong, or run a bad recommendation
automatically?"** No tool here executes a `recommendations` string --
those are just text for the agent (or you) to read and decide on.
`dev_check`, `dev_build` and `dev_build_all` are the only tools that execute
anything, and only commands they recognized themselves against a fixed
allowlist -- see [Security model](#security-model).
**"Does it phone home or send telemetry anywhere?"** No. Zero network
calls of its own -- see [Local-first architecture](#local-first-architecture).
**"I don't want an MCP server running *any* commands on my machine."**
10 of the 13 tools are read-only (file reads, version checks). Only
`dev_check`, `dev_build` and `dev_build_all` execute project commands, and only
commands DevTwin itself recognized from project files, checked against an
allowlist, with `shell=False` and a timeout -- see
[Security model](#security-model) for exactly what that does and doesn't
allow.
## Benefits
- **Fewer wrong diagnoses.** Without DevTwin, an agent debugging a failure
can only read code and guess -- it will often propose a code fix for
what's actually a Node version mismatch or a stopped database. DevTwin
gives it ground truth instead of a guess.
- **One call instead of many.** A single `dev_health` call bundles ~10
underlying checks (runtime versions, dependency state, services, ports,
Git) into one structured, scored result -- instead of an agent making a
dozen separate shell round-trips and parsing raw CLI output each time.
- **Same check every time.** The exact checks per ecosystem (Gradle
wrapper vs. system Gradle, `.nvmrc` vs. `package.json` engines, ...) are
encoded once, so the diagnosis is consistent across sessions instead of
depending on what an agent happens to think to run.
- **Safer than handing an agent a shell.** No arbitrary command execution,
no destructive operations, ever -- see [Security model](#security-model).
- **Secrets never touched.** Environment variables that look secret are
checked for presence only; values are never read or returned -- see
[Privacy model](#privacy-model).
- **Works even where the agent has no shell.** MCP clients without a Bash
tool (some IDE assistants, restricted agents) get this capability at
all, not zero capability.
## Token cost
Real numbers, not an estimate -- measured directly from this server's own
MCP tool schemas (`mcp.list_tools()`) and a real `dev_health()` response,
using the standard ~4-characters-per-token approximation.
**Two different moments spend tokens, and they cost very differently:**
| When | What happens | Cost |
|---|---|---|
| **The moment the client connects** to DevTwin | All 13 tool schemas (name, description, parameters) are added to *every request* in that session -- whether or not any tool is ever called. This is true of any MCP server, not specific to DevTwin. | **≈1,900 tokens, every single turn** |
| **Only when a tool is actually called** | That one tool's JSON response is added to context, once. | **~120-200 tokens per call** (varies with how many issues are found) |
Per-tool schema breakdown (measured):
| Tool | Schema size | ≈ tokens |
|---|---|---|
| `dev_detect` | 440 chars | ~110 |
| `dev_health` | 500 chars | ~125 |
| `dev_health_all` | 593 chars | ~148 |
| `dev_drift` | 470 chars | ~117 |
| `dev_explain_failure` | 793 chars | ~198 |
| `dev_project_info` | 523 chars | ~130 |
| `dev_dependencies` | 507 chars | ~126 |
| `dev_services` | 507 chars | ~126 |
| `dev_check` | 771 chars | ~192 |
| `dev_build` | 792 chars | ~198 |
| `dev_build_all` | 604 chars | ~151 |
| `dev_prepare` | 645 chars | ~161 |
| `dev_precommit` | 481 chars | ~120 |
| **Total (all 13 tools)** | **7,626 chars** | **≈1,900** |
**The honest bottom line:** for a *single* one-off diagnosis in a session
that otherwise never touches an environment question, raw Bash can come
out cheaper in total tokens -- the ~1,900-token fixed schema tax often
outweighs the savings from replacing several shell commands with one call.
See the worked comparison below for real numbers on both sides.
DevTwin's case gets stronger the more environment questions come up in one
session (the fixed tax is paid once; every question after that is ~150
tokens on DevTwin vs. hundreds more on raw Bash each time) -- and its real
advantage isn't raw token count at all, it's consistency, safety, and
working in MCP clients that have no Bash tool. See
[Benefits](#benefits) and [Honest tradeoffs](#honest-tradeoffs).
**Practical implication:** register DevTwin per-project, not user-wide, so
the fixed tax is only paid in sessions where it's actually useful -- see
[Using it on another project](#using-it-on-another-project-for-other-developers).
## Honest tradeoffs
DevTwin is not a daily-use tool for a stable environment -- nobody needs
to re-check "is Postgres running" on every function they write. It's a
**break-glass tool**: high value at specific moments (fresh clone, a build
that mysteriously fails, right before a commit), and idle the rest of the
time. That's the intended usage pattern, not a shortcoming.
- Token overhead is paid on every turn the moment it's connected, whether
used or not -- see [Token cost](#token-cost) for real measured numbers.
- It doesn't reliably win on tokens for a single one-off question; it wins
on consistency, safety, and reach into clients with no shell -- see
[Benefits](#benefits).
- If an agent already has full shell access to a repo you fully control
and rarely has environment drift, you may not need DevTwin there at all.
- DevTwin earns its keep most on: shared/onboarding repos, less-trusted or
shell-less agent setups, and multi-ecosystem monorepos where "what do I
even check" is itself the hard part.
## With vs. without DevTwin: a worked example
Say you ask an agent "why does `npm test` fail?" and the real cause is a
Node version mismatch plus Postgres not running.
**Without DevTwin** (agent using raw Bash) -- it has to guess the right
sequence, one command at a time:
```
cat package.json # spot "engines": {"node": ">=20"}
node --version # v16.20.0 -- mismatch found
grep -i "pg\|postgres" package.json # spot the Postgres dependency
cat .env # risk: may print a real secret into context
lsof -i :5432 # nothing listening
docker ps # check if it's in a container instead
```
Six round-trips, an investigation path the agent had to invent, a real
chance of a secret leaking into the conversation at step 4, and roughly
**400-800 tokens** of command + output text (varies with file sizes and
how many Docker containers are running).
**With DevTwin**, one call:
```
dev_health()
```
```json
{
"status": "error",
"summary": "2 issues found: runtime drift, service down",
"issues": [
"Node 16.20.0 installed, project requires >=20 (from package.json engines)",
"Postgres required (found in docker-compose.yml) but not running on 5432"
],
"recommendations": [
"nvm install 20 && nvm use 20",
"docker compose up -d postgres"
]
}
```
Same conclusion, ~150 tokens for the response -- plus the ~1,900-token
fixed schema tax already paid that turn regardless (see
[Token cost](#token-cost)). One call instead of six, no possibility of
leaking a secret, and the exact same curated check every time instead of
a freehand investigation that varies session to session.
## Example questions this unlocks
- "Check my development environment."
- "Why is my Kotlin project failing to build?"
- "Is my Node version correct for this repo?"
- "Why can't my app connect to Postgres?"
- "Does my environment drift from what this repository expects?"
- "What should I run before I commit?"
- "I just cloned this repo -- what do I need to do to get it running?"
- "Check all ecosystems in this monorepo" (uses `dev_health_all` for Android/iOS/React/Python)
- "I just changed backend code — do Android, iOS, and React still build?" (uses `dev_build_all`)
- "Which of my backend/frontend/mobile apps is ready to ship?"
## Per-language examples
One row per supported ecosystem: a question you'd actually ask, what
DevTwin checks to answer it, and the test/build command it recognizes for
`dev_check`.
| Ecosystem | Example question | What gets checked | Recognized command(s) |
|---|---|---|---|
| Python | "Is my Python version right for this repo?" | `python`/`python3` vs. `.python-version` or `pyproject.toml [project.requires-python]`; uv/pip/poetry/pipenv + lockfile | `pytest`, `ruff check .`, `mypy .` |
| Node.js | "Why does `npm test` fail?" | `node` vs. `.nvmrc`/`.node-version`/`package.json engines`; npm/pnpm/yarn/bun + lockfile | `npm test` (or `pnpm test`/`yarn test`/`bun test`), `<mgr> run lint` |
| JVM (Java + Kotlin + Android) | "Why won't my Android app build after a fresh clone?" | `java`/`kotlinc` version; Gradle wrapper version vs. installed; Maven wrapper; **on Android projects specifically:** `ANDROID_HOME`/`ANDROID_SDK_ROOT`, or `local.properties`' `sdk.dir` and whether that path actually exists | `./gradlew test`, `./mvnw test` |
| Go | "Is my Go version correct for this repo?" | `go` vs. the version required in `go.mod` | `go test ./...`, `go build ./...` |
| Rust | "Why does `cargo build` fail?" | `rustc` vs. `rust-toolchain[.toml]` channel | `cargo test` |
| .NET | "Why does `dotnet build` fail?" | `dotnet` SDK presence and version | `dotnet test` |
| Swift (iOS/macOS) | "Why does my iOS build fail?" | `swift`/`xcodebuild` vs. `Package.swift` tools-version; CocoaPods/SPM lockfile state | `swift test` (SPM); `xcodebuild test` (macOS Xcode projects) |
| Ruby | "Why does `bundle exec rspec` fail?" | `ruby` vs. `.ruby-version`; Bundler + `Gemfile.lock` | `bundle exec rspec`, `bundle exec rake test` |
| PHP | "Why does my PHP app fail to boot?" | `php` vs. `composer.json`'s `require.php`; Composer + `composer.lock` | `composer test`, `vendor/bin/phpunit` |
| Generic (fallback) | "This repo isn't in any language above -- what can you tell me?" | `Makefile`/`Taskfile.yml`/`justfile`/`Dockerfile`/compose services | `make test`, `task test`, `just test` |
## Architecture
One MCP server, many ecosystem adapters -- not a separate server per
language.
```
MCP server -> core (workspace/detector/health/drift/diagnostics) ->
adapters (python/node/jvm/go/rust/dotnet/swift/ruby/php/generic) ->
system inspection (os/process/ports/env/fs/docker) ->
service detection (postgres/redis/generic)
```
Full details in [`docs/architecture.md`](docs/architecture.md). How to add
a new language adapter: [`docs/adapters.md`](docs/adapters.md).
## Supported ecosystems
| Ecosystem | Detected from | Runtime checked | Package managers |
|---|---|---|---|
| Python | `pyproject.toml`, `requirements.txt`, `uv.lock`, `poetry.lock`, `Pipfile`, `.python-version` | `python`/`python3` | uv, pip, poetry, pipenv |
| Node.js | `package.json`, lockfiles, `.nvmrc`, `.node-version` | `node` | npm, pnpm, yarn, bun |
| JVM (Java + Kotlin) | `pom.xml`, `build.gradle[.kts]`, `.java`/`.kt` sources | `java`, `kotlinc` | Gradle (wrapper-aware), Maven (wrapper-aware) |
| Go | `go.mod`, `go.sum`, `go.work` | `go` | go modules |
| Rust | `Cargo.toml`, `rust-toolchain[.toml]` | `rustc` | cargo |
| .NET | `*.csproj`/`*.fsproj`/`*.vbproj`, `*.sln`, `global.json` | `dotnet` | NuGet |
| Swift (iOS/macOS) | `Package.swift`, `*.xcodeproj`, `*.xcworkspace`, `Podfile` | `swift`, `xcodebuild` | SPM, CocoaPods |
| Ruby | `Gemfile`, `*.gemspec`, `.ruby-version` | `ruby` | Bundler |
| PHP | `composer.json` | `php` | Composer |
| Generic (fallback) | `Makefile`, `Taskfile.yml`, `justfile`, `Dockerfile`, compose files | -- | make/task/just/docker |
Any project not matching a specific adapter still gets useful output from
the generic adapter -- DevTwin never returns nothing for an unrecognized
project.
## Installation
**Prerequisites:**
- macOS/Linux (Windows: WSL)
- Python 3.10+
- One package manager: `uv`, `pipx`, or `pip`
**Pick one method:**
**Option 1: `uv` (fastest, recommended)**
```bash
# Install uv first (if not already installed)
brew install uv
# Then install devtwin
uv pip install devtwin-mcp
```
**Option 2: `pipx` (simplest, no venv needed)**
```bash
# Install pipx first (if not already installed)
brew install pipx
# Then install devtwin
pipx install devtwin-mcp
```
**Option 3: `pip` (standard, may need venv on newer macOS)**
```bash
pip install devtwin-mcp
# or with venv:
python3 -m venv ~/.devtwin-venv
source ~/.devtwin-venv/bin/activate
pip install devtwin-mcp
```
All methods install the `devtwin` binary globally so it works in any MCP client.
### Troubleshooting installation
**"zsh: command not found: uv"**
```bash
brew install uv
uv pip install devtwin-mcp
```
**"error: externally-managed-environment"** (on newer macOS)
Use `pipx` (simplest solution):
```bash
brew install pipx
pipx install devtwin-mcp
```
**"pip: command not found"**
Use `pipx` or `uv` (above), or create a venv:
```bash
python3 -m venv ~/.devtwin-venv
source ~/.devtwin-venv/bin/activate
pip install devtwin-mcp
```
**Verify installation:**
```bash
devtwin --version
# Should print: devtwin X.Y.Z
```
For local development against a clone of this repo, see
[`docs/development.md`](docs/development.md).
## MCP client configuration
Exact configuration syntax differs by client -- consult your client's docs.
Generically, DevTwin is a stdio MCP server invoked as:
```json
{
"mcpServers": {
"devtwin": {
"command": "devtwin"
}
}
}
```
For local development from a clone (without installing the package):
```json
{
"mcpServers": {
"devtwin": {
"command": "uv",
"args": ["run", "--directory", "/absolute/path/to/devtwin-mcp", "devtwin"]
}
}
}
```
Verify tool discovery with the MCP Inspector:
```bash
npx @modelcontextprotocol/inspector uv run devtwin
```
## Using it on another project (for other developers)
DevTwin is one binary -- point any number of projects at the same install,
no per-project reinstall needed. Two scopes:
| Scope | Loads | When to use |
|---|---|---|
| **Project** (recommended default) | Only in this repo | Default choice -- see [Token cost](#token-cost) for why |
| **User** | Every project, every session | Once you're reaching for DevTwin across most of your repos |
**Project scope** -- drop a `.mcp.json` in the project root:
```json
{
"mcpServers": {
"devtwin": {
"command": "/absolute/path/to/devtwin-mcp/.venv/bin/devtwin"
}
}
}
```
or with the Claude Code CLI:
```bash
claude mcp add devtwin /absolute/path/to/devtwin-mcp/.venv/bin/devtwin --scope project
```
**User scope**:
```bash
claude mcp add devtwin /absolute/path/to/devtwin-mcp/.venv/bin/devtwin --scope user
```
After adding it, restart the client (or reconnect the MCP server), then
just ask normal questions -- see
[Example questions this unlocks](#example-questions-this-unlocks).
**Monorepo tip:** in a repo mixing platforms (e.g. Android + iOS +
backend), point questions at the specific subfolder rather than the repo
root -- e.g. "check the health of the `android/` app". `dev_detect` at the
root of a mixed repo reports every ecosystem it finds, which is useful
once but noisy for a targeted check.
## Tool reference
All tools return `{status, summary, data, issues, recommendations}`.
`status` is one of `ok`, `warning`, `error`, `unknown`.
| Tool | Class | Description |
|---|---|---|
| `dev_detect` | read-only | Fast, file-based project/ecosystem detection with evidence. |
| `dev_health` | read-only | Full 0-100 health score combining runtime, dependency, service, and Git state. |
| `dev_health_all` | read-only | Scan subdirectories for multiple ecosystems; run comprehensive health checks on all with detailed per-ecosystem reports: health score, runtime versions, dependency state, services, issues (build errors, conflicts, misconfigurations), and recommendations. Perfect for monorepos. |
| `dev_drift` | read-only | Compares required vs. actually-installed runtime/tool versions. |
| `dev_explain_failure` | read-only | Diagnoses a given error message into ranked, evidence-backed root causes. |
| `dev_project_info` | read-only | Detailed project inspection: runtimes, build tools, commands, OS, Git. |
| `dev_dependencies` | read-only | Per-ecosystem dependency/lockfile state. |
| `dev_services` | read-only | Required local services (Postgres, Redis, compose services) and their running state. |
| `dev_check` | safe execution | Runs recognized test/lint commands (e.g. `pytest`, `./gradlew test`) with a timeout. |
| `dev_build` | safe execution | Runs recognized build/compile commands (e.g. `npm run build`, `./gradlew build`, `xcodebuild build`) with a longer timeout. |
| `dev_build_all` | safe execution | Scan subdirectories for multiple ecosystems and run builds on all; returns per-ecosystem build results. Perfect for monorepos to verify backend changes don't break Android, iOS, and frontend builds. |
| `dev_prepare` | plans only | Produces a preparation plan for a freshly-cloned repo; never executes it. |
| `dev_precommit` | read-only | Commit-readiness summary: Git state, health, staged-secret-looking files. |
### `dev_health_all` detailed output
For monorepos with multiple ecosystems, `dev_health_all` returns per-ecosystem details:
Each ecosystem report includes:
- **health_score** (0-100): Overall ecosystem health
- **status**: ok/warning/error
- **runtime_summary**: Actual vs. required versions (Java/Swift/Node/Python)
- **dependency_summary**: Lockfile state, conflicts, missing packages
- **service_summary**: Required services (Postgres, Redis, etc.) and running status
- **issues**: Detailed problems found:
- Build errors (Gradle, xcodebuild, npm, pip)
- Dependency conflicts
- Version mismatches
- Missing SDKs or tools
- **recommendations**: Specific fixes for each issue
**Example:** For Android, you get Gradle build errors, missing SDK paths, Java version mismatches. For iOS, you get CocoaPods errors, Swift version issues. For React, you get npm conflicts. For Python, you get pip version conflicts.
## Security model
- **No arbitrary command execution.** There is no `execute_shell` tool.
`dev_check`, `dev_build` and `dev_build_all` only run commands DevTwin
itself recognized from project files, checked against an allowlist, run
with `shell=False` and a timeout.
- **No destructive actions, ever.** DevTwin never runs `git reset --hard`,
`rm -rf`, `kill -9`, `docker compose down`, lockfile deletion, or
`.env` mutation.
- **`dev_prepare` only plans.** It classifies every proposed step
(`read_only`/`safe`/`requires_approval`/`dangerous`) and never executes
anything itself.
Full details: [`docs/security.md`](docs/security.md).
## Privacy model
- Environment variables are checked for **presence only** when their name
looks secret (`PASSWORD`, `TOKEN`, `SECRET`, `API_KEY`, `PRIVATE_KEY`,
`ACCESS_KEY`, `AUTH`, `CREDENTIAL`, ...) -- values are never returned.
- `.env` files are scanned for variable *names* only.
- `dev_precommit` flags secret-*looking* staged filenames without reading
or reporting their contents.
## Local-first architecture
- No server component, no account, no network calls of its own beyond the
local commands it inspects (`git`, `docker`, language toolchains).
- Everything it reports comes from files and processes already on the
machine it runs on.
- One exception worth knowing about: on an Xcode project, discovering the
build/test commands runs `xcodebuild -list`, which populates DerivedData and,
for a project with SwiftPM dependencies, may resolve packages over the
network. It runs at most once per adapter run and is the only read-only path
that is not purely a file read.
## MCPHub plugin
DevTwin is also packaged as an MCPHub plugin in [`mcphub/devtwin/`](mcphub/devtwin/) —
a TypeScript port of this server that satisfies the MCPHub plugin contract
(18 tools under the `devtwin_` prefix, pricing tiers, config schema, action
plans, SKILL.md). It exists for teams who consume tools through
`mcphub.indianic.in` rather than a local `.mcp.json`.
```bash
cd mcphub/devtwin
npm install && npm run verify && npm run smoke
```
The two distributions are not interchangeable. This Python server runs on the
developer's machine and reports that machine. The plugin runs on MCPHub's
server, so its live-state tools (`devtwin_services`, `devtwin_drift`,
`devtwin_check`, `devtwin_build`) describe the server they execute on; its
file-based tools (`devtwin_detect`, `devtwin_project_info`,
`devtwin_dependencies`, `devtwin_precommit`) read the workspace and are correct
either way. See [`mcphub/devtwin/README.md`](mcphub/devtwin/README.md) for the
full list of hosted-execution differences.
For local use, keep using this server.
## Adoption & team setup
**For team leads:** See [`ADOPTION.md`](ADOPTION.md) for per-project setup, FAQ, and how to announce DevTwin to your team.
**Copy-paste messaging:** See [`MESSAGING.md`](MESSAGING.md) for Slack, email, GitHub, and internal docs templates.
**Key idea:** Register DevTwin per-project in `.mcp.json` (so the fixed token tax only applies to sessions that use it). Individual developers install once (`uv pip install devtwin-mcp`), and every project they work on gets it automatically.
## Development
```bash
uv sync --all-extras
uv run pytest
uv run ruff check .
uv run mypy src
uv run devtwin
```
See [`docs/development.md`](docs/development.md) for the full workflow.
## Contributing
See [`CONTRIBUTING.md`](CONTRIBUTING.md). Adding a new language ecosystem
is the most common contribution -- see [`docs/adapters.md`](docs/adapters.md)
for a template, or [`src/devtwin/adapters/swift.py`](src/devtwin/adapters/swift.py),
[`ruby.py`](src/devtwin/adapters/ruby.py), and
[`php.py`](src/devtwin/adapters/php.py) for real, merged examples to
model yours after.
## Roadmap
- Additional ecosystem adapters: Elixir, Dart, Scala,
C/C++ (CMake/Bazel/Buck), Nix (see `docs/adapters.md` for how to add one)
- Additional service detectors (MySQL/MariaDB, MongoDB, Kafka, RabbitMQ)
- Richer drift comparison against CI configuration (e.g. GitHub Actions
runtime matrices)
- Optional local caching of expensive checks across tool calls within a
session
## License
Apache-2.0 -- see [`LICENSE`](LICENSE).
TDQS
Scored across 10 tools
Each tool targets a distinct aspect of the development environment: detection, health, drift, failure diagnosis, project info, dependencies, services, checks, preparation, and precommit. Even similar tools like dev_detect and dev_project_info are clearly differentiated by scope and speed. There is no ambiguous overlap that would cause an agent to select the wrong tool.
All tool names follow the consistent pattern `dev_` + lower_snake_case, using descriptive verbs or nouns (detect, health, drift, explain_failure, etc.). The naming convention is uniform and predictable, with no mixing of camelCase or inconsistent verb styles.
With 10 tools, the server is well-scoped and each tool serves a clear purpose within the domain of development environment analysis and preparation. The count is within the ideal range and avoids both bloat and insufficient coverage.
The tool set covers the full lifecycle for a diagnostics/preparation server: detection, health assessment, drift checking, failure explanation, dependency and service checks, test execution, preparation planning, and precommit readiness. No obvious gaps exist for the stated purpose, and the tools work together to provide comprehensive environment insight.