diffctx
diffctx — smart diff context for LLM code review
diffctx selects the minimum code an LLM needs to review a git diff. Instead of pasting whole files, it walks the dependency graph outward from the changed lines and stops once more context stops paying for itself.
Coming from
treemapper? That name is deprecated. Every command, flag, and API call works unchanged:treemapper→diffctx,treemapper-mcp→diffctx-mcp.
Why not just use tree or repomix?
| repomix | Claude Code Review | diffctx | |
Primary use case | directory listing | full repo export | automated PR review | diff context for code review |
Smart diff context | ✗ | ✗ | ✓ | ✓ |
Works with any LLM | ✓ | ✓ | Claude only | ✓ |
Free / local / offline | ✓ | ✓ | $15–25/review | ✓ |
GitHub required | ✗ | ✗ | ✓ | ✗ |
Multiple output formats | ✗ | limited | — | YAML/JSON/MD/txt |
Python API | ✗ | ✗ | ✗ | ✓ |
MCP server | ✗ | ✗ | ✗ | ✓ |
Fuller positioning — measured results, and when a whole-repo packer or a persistent code-graph server fits better: COMPARISON.md.
Related MCP server: Code Reference Optimizer MCP Server
Install
uvx diffctx . --diff HEAD~1 # zero-install, run once via uv
pipx install diffctx # recommended: isolated CLI, no venv needed
pip install diffctx # or: into an active environment
pipx install 'diffctx[mcp]' # + MCP server for AI assistantsWithout Python:
cargo install diffctx # native CLI from crates.io
npx diffctx . --diff HEAD~1 # npm wrapper over the native binary
docker run --rm -v "$PWD:/repo" ghcr.io/nikolay-e/diffctx . --diff HEAD~1On Windows, via Scoop (this repository is the bucket):
scoop bucket add diffctx https://github.com/nikolay-e/diffctx
scoop install diffctx/diffctxThe image runs as a non-root user (uid 10001) and writes to stdout — the
native binary has no -o flag, so redirect to capture: ... --diff HEAD~1 > context.yaml.
Prebuilt binaries for linux (x86_64/aarch64), macOS (arm64) and Windows (x64)
are attached to every release.
The native binary covers diff mode with YAML/JSON output; tree mode, Markdown
output, the graph subcommand and the MCP server live in the Python package.
cargo add diffctx embeds the pipeline in a Rust project — the library is
imported as _diffctx (use _diffctx::pipeline::build_diff_context), since
the crate doubles as the Python extension module
(docs.rs).
Quick start
diffctx . --diff HEAD~1 # smart context for last commit → paste into Claude/ChatGPT
diffctx . -f md -c # full codebase export → clipboard in Markdown
diffctx . --diff HEAD~1 selects only the fragments an LLM needs to review the
last commit, instead of dumping every changed file in full.
Diff context mode
Finds the minimal set of fragments needed to understand a change — imports,
callers, type definitions, config dependencies — across 50+ file types. It
builds a code graph (imports, co-changes, type refs), propagates relevance
outward from the changed lines, and stops when relevance drops below --tau or
the --budget token cap is hit.
Flag | Default | Description |
|
|
|
| auto | Hard cap in o200k_base tokens (see Token counting): |
| 0.60 | PPR damping; higher = context clusters tighter around changes ( |
| 0.12 | Relevance threshold for full fragment content; lower-scoring fragments are stubbed or dropped (lower = more context) |
| false | Only the changed files, every fragment, no related-code context |
| 300 | Wall-clock deadline in seconds; on expiry diffctx exits 124 instead of hanging |
| false | Also embed git's raw unified diff ahead of the selected fragments — additive (selection unchanged), not charged to |
Calibration of --alpha, --tau, and the edge-weight priors:
docs/engineering/parameter-strategy.md.
Theory:
diffctx: Budgeted Typed-Graph Retrieval for Diff-Aware Code Context
Selection (Zenodo, 2026).
graph subcommand
Explore the underlying dependency graph directly, without a diff:
diffctx graph . # Mermaid graph of directory deps (default)
diffctx graph . --summary # cycles, hotspots, coupling metrics
diffctx graph . --level fragment -f json # fragment-level graph as JSON
diffctx graph . --level file -f graphml -o g.xml # file-level graph as GraphMLUsage
# full codebase export:
diffctx . # Markdown to stdout + token count
diffctx . -f md -c # Markdown → clipboard
diffctx . -f json -o tree.json # JSON → file
diffctx . --no-content # structure only, no file contents
diffctx . --max-depth 3 # limit depth
diffctx . -i custom.ignore # custom ignore patterns
# diff context mode (requires git repo):
diffctx . --diff # uncommitted changes (working tree vs HEAD)
diffctx . --diff HEAD~1 # context for last commit
diffctx . --diff main..feature # context for feature branch
diffctx . --diff HEAD~1 --budget 30000 # limit to ~30k tokens
diffctx . --diff HEAD~1 -c # diff context to clipboard
diffctx . --diff HEAD~1 --with-raw-diff # raw patch + selected contextEvery run reports token count and size on stderr — 12,847 tokens (o200k_base), 52.3 KB (tiktoken, the GPT-4o tokenizer; ~-prefixed above
1 MB). -c/--copy copies output via pbcopy (macOS), clip (Windows), or
wl-copy/xclip/xsel (Linux). Unreadable files become placeholders like
<binary file: N bytes>, <file too large: N bytes>, or
<unreadable content: not utf-8>.
Counts come from tiktoken's o200k_base encoder and are exact only for the
GPT-4o family; Claude, Gemini and others tokenize differently, so treat
--budget as an upper bound in o200k tokens and leave headroom. Details:
docs/product/token-budget.md.
Python API
from pathlib import Path
from diffctx import build_diff_context, map_directory, to_json, to_markdown, to_text, to_yaml
ctx = build_diff_context(
Path("."),
"HEAD~1..HEAD",
budget_tokens=None, # None = auto; 0 = strict-zero floor (empty); -1 = uncapped; N = hard cap
alpha=0.6,
tau=0.12,
full=False,
scoring_mode="ego",
timeout=300,
with_raw_diff=False, # True also embeds the raw unified diff (not charged to budget)
)
print(to_markdown(ctx))
tree = map_directory(
".",
max_depth=None,
no_content=False,
max_file_bytes=None,
ignore_file=None,
no_default_ignores=False,
whitelist_file=None,
)
print(to_yaml(tree))MCP server
diffctx includes an MCP server that lets AI
assistants (Claude Code, Cursor, Windsurf, etc.) call diff context analysis
automatically during code review. It is published in the official MCP registry
as io.github.nikolay-e/diffctx. One-line setup (zero-install via
uv):
# Claude Code
claude mcp add diffctx -- uvx --from 'diffctx[mcp]' diffctx-mcp
# Codex CLI
codex mcp add diffctx -- uvx --from 'diffctx[mcp]' diffctx-mcp
# Gemini CLI
gemini mcp add diffctx uvx -- --from 'diffctx[mcp]' diffctx-mcp
# VS Code
code --add-mcp '{"name":"diffctx","command":"uvx","args":["--from","diffctx[mcp]","diffctx-mcp"]}'With pip install 'diffctx[mcp]' already done, replace the
uvx --from 'diffctx[mcp]' diffctx-mcp tail with plain diffctx-mcp.
The server exposes three tools — get_diff_context, get_tree_map, and
get_file_context — that assistants call when reviewing PRs, explaining
changes, or investigating broken tests. Tool reference and JSON configs for
Claude Desktop, Cursor, Continue, Windsurf, and Zed:
src/diffctx/mcp/README.md.
Ignore patterns
Respects .gitignore and .diffctx/ignore automatically — hierarchically at
every directory level, with full gitignore semantics (negation !important.log,
anchored /root_only.txt). .diffctx/whitelist acts as an include-only filter,
and the output file is always auto-ignored. --no-default-ignores disables the
built-in patterns; --no-ignores disables all ignore rules (tree mode only).
An excluded path never appears in the output in any role: in diff mode it is
dropped both from changed_files and from the candidate universe, so it cannot
come back as a related-context fragment either (including under --full). The
same guarantee covers secret-like paths (id_rsa, *.pem, *.key, ...),
which are filtered even without an ignore entry.
Token cache
Diff mode caches per-blob tokenization under
~/Library/Caches/diffctx/token-cache (macOS),
$XDG_CACHE_HOME/diffctx/token-cache (Linux) or
%LOCALAPPDATA%\diffctx\token-cache (Windows). It is a pure speedup: deleting
it only costs one cold run.
Variable | Effect |
| Relocate the cache |
| Size cap, default |
Eviction is amortized: each run trims one of the cache's 256 shards back under its share of the cap, oldest entries first.
Exit codes
Code | Meaning |
| Success — output contains content |
| Runtime error (bad path, permission denied, etc.) |
| Usage error (invalid flags/arguments) |
| Environment error ( |
|
|
|
|
| Interrupted (Ctrl-C) |
| Broken pipe (e.g. piping into |
Repository layout
Rust product code lives in crates/diffctx-native/; the Python CLI and MCP
package live in src/diffctx/. Evaluation code, immutable datasets, and paper
artifacts are intentionally separated under eval/, datasets/, and paper/.
See repository ownership boundaries
for the lifecycle and entry point of each area.
License
Apache 2.0
Documentation site — the pipeline end to end: diff → fragments → graph → relevance → selection
GitHub Action — diff context as a CI step for LLM review
Token counting — which encoder, and what
--budgetmeans for non-GPT modelsComparison — measured results, and when a whole-repo packer or a persistent code-graph server fits better
Security policy — threat model and vulnerability reporting
Parameter strategy — how
--alpha,--tau, and edge weights are calibrated
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/nikolay-e/diffctx'
If you have feedback or need assistance with the MCP directory API, please join our Discord server