intent-diff-mcp
by ssh00n
README.md
# Intent-Diff MCP
> Catch **context drift** ā the scope an AI coding agent quietly adds beyond what you actually asked for.
An [MCP](https://modelcontextprotocol.io) server with two tools. You lock in your original request; the agent, before it declares "done", diffs the actual code changes against that request and reports anything it added on its own ā an unrequested library, a metrics layer, a new abstraction ā so you can **Keep** it or **Roll it back**.
The *diff* is deterministic ground truth (git doesn't lie). The *judgment* is a fresh Claude call that didn't write the code, so it has no stake in rationalizing the additions.
```
š¤ Intent Diff Report
Requested: Add a Postgres connection pool in db.ts ā ā
done
[+] Added on my own (drift candidates):
- New metrics.ts with Prometheus instrumentation ā you never asked for
observability or a new dependency (prom-client).
[?] Worth confirming:
- Should metrics.ts stay, or be removed?
ā Keep or Rollback?
```
## Tools
| Tool | When | What it does |
|------|------|--------------|
| `start_task` | A non-trivial feature/refactor is requested | Records your request + snapshots the current git HEAD. Cheap (one SHA). |
| `get_intent_diff` | Before saying "done" on substantial changes | Diffs everything since the snapshot, runs the judge, returns a drift report. |
## Install
Requires **Node ā„ 18** and either a logged-in **Claude Code** session *or* an `ANTHROPIC_API_KEY` (see [Auth](#auth)).
Add to your MCP client config (Claude Code `.mcp.json`, `mcporter.json`, etc.):
```json
{
"mcpServers": {
"intent-diff": {
"command": "npx",
"args": ["-y", "@ssh00n/intent-diff-mcp"]
}
}
}
```
That's the whole setup ā `npx` fetches and runs it.
## Recommended workflow (hybrid)
Intent-Diff is **not** meant to gate every edit ā that just taxes trivial work. The sweet spot: **always snapshot (cheap), judge on demand (costly)**. Drop this into your project's `CLAUDE.md` (a ready copy is in [`examples/CLAUDE.md`](examples/CLAUDE.md)):
```markdown
# Intent-Diff workflow (hybrid)
- When starting a non-trivial feature or refactor, call `start_task` with the
user's request verbatim. (It's cheap ā just a git snapshot.)
- Before declaring a substantial change "done" ā new feature, or edits spanning
multiple files ā run `get_intent_diff` and show the report.
- Treat the report as advisory: surface drift to the user and ask Keep/Rollback.
Skip it for one-line fixes and pure exploration.
```
Prefer softer or stricter? See [Tuning the rules](#tuning-the-rules).
## Auth
The judge needs to call Claude. Two ways, checked in this order:
1. **`ANTHROPIC_API_KEY`** (recommended for CI and contributors) ā normal API billing.
2. **Local Claude Code subscription** ā if no API key is set, the server reuses the
OAuth token your Claude Code login already stores (macOS Keychain or
`~/.claude/.credentials.json`), refreshing it when near expiry. No extra setup if
you're already logged into Claude Code.
Pick the judge model with `INTENT_DIFF_MODEL` (default `claude-sonnet-4-6`).
## Tuning the rules
The workflow above is the **hybrid** default. Adjust to taste:
- **Softer** ā drop the `get_intent_diff` line; call it only when *you* ask "check drift". The agent self-checks the rest of the time.
- **Stricter** ā make both calls mandatory ("always `start_task` first; never say done without `get_intent_diff`") and require an explicit Keep/Rollback answer. Good for team enforcement; higher overhead.
Note: a heavy-handed rule can bias an agent toward *under*-implementing (skipping necessary error handling for fear of "drift"). The judge is told to ignore refactors and necessary error handling, but keep the rule proportional to the stakes.
## How it works
`start_task` writes `<repo>/.mcp/intent_state.json`, keyed by git branch, with your
request and the HEAD SHA. `get_intent_diff` collects `git diff <baseSha>` plus any
untracked files (skipping binaries, lockfiles, and `.mcp/`), caps it, and sends
`{ original intent, actual diff }` to the judge, which returns structured JSON that's
rendered into the report.
## Development
```bash
npm install
npm run typecheck # tsc --noEmit
npm test # Tier-1 unit tests (no network)
npm run test:smoke # end-to-end via MCP client, stub judge (no network)
npm run eval # judge accuracy over labeled fixtures (needs auth ā real calls)
```
## Contributing
Contributions welcome ā see [CONTRIBUTING.md](CONTRIBUTING.md). Good first areas:
more judge fixtures, additional credential sources, non-git VCS support.
## License
[MIT](LICENSE) Ā© ssh00n
TDQS
A4.2/5.0
Scored across 2 tools
Disambiguation5/5
The two tools serve distinct purposes: start_task captures the initial state, while get_intent_diff compares changes against that state. No overlap or ambiguity.
Naming Consistency5/5
Both tools follow a consistent verb_noun snake_case pattern (start_task, get_intent_diff), making their actions clear and predictable.
Tool Count3/5
With only 2 tools, the server is thin, but it is focused on a narrow workflow (snapshot and diff). For its limited scope, the count is borderline but acceptable.
Completeness5/5
The tool set covers the full lifecycle needed: capturing the original intent (start_task) and reporting drift (get_intent_diff). No obvious gaps for its stated purpose.
Maintenance
ActivityInactive
ResponsivenessNo issues