ripple
by xtelos
README.md
# ripple
**Know what breaks before you change it.** ripple is an MCP server that gives a coding agent
(Claude Code, Cursor, anything that speaks MCP) the call graph of a Python repository, so the agent
can ask "who calls this, and which tests reach it?" before it edits, and run just those tests after.
## Why
A coding agent dropped into unfamiliar code edits blind. It greps for a function name, sees a few
hits, and changes the signature. The callers that reach it through an import alias, a `self.` call in
a subclass or a re-export in a package `__init__` don't match the grep, and the first sign of trouble
is a failing test, or a failure in production. ripple answers the structural question directly and
says plainly where static analysis runs out.
## Quickstart
Requires Python 3.10+ and [uv](https://docs.astral.sh/uv/).
```sh
uv tool install git+https://github.com/xtelos/ripple # puts `ripple` on your PATH
```
**Claude Code**, from your project directory:
```sh
claude mcp add ripple -- ripple serve --repo "$PWD"
```
**Cursor**, in `.cursor/mcp.json`:
```json
{
"mcpServers": {
"ripple": { "command": "ripple", "args": ["serve", "--repo", "${workspaceFolder}"] }
}
}
```
Or from a terminal: `ripple index .`, then `ripple impact price_order`.
## The five tools
| Tool | What it answers | Read-only |
|---|---|---|
| `callers(symbol)` | Who calls this directly, with the file:line of each call site | yes |
| `callees(symbol)` | What this calls, plus the calls ripple could not resolve and why | yes |
| `impact(symbol, depth=5, limit=100)` | Every transitive caller grouped by file, the tests that reach it listed separately as pytest node ids, and same-named call sites ripple could not rule out | yes |
| `path(source, target)` | The shortest call chain from one symbol to another | yes |
| `check(tests=None)` | Runs the test command (default `python -m pytest -q`, using the repo's `.venv` if present), or just the given tests, and returns pass/fail, the summary line and the first failures, never the full log | no, it runs code |
A symbol can be a full dotted path (`shop.pricing.price_order`), a suffix (`OrderService.place_order`)
or a bare name (`price_order`). An ambiguous name returns the candidates rather than picking one.
The four graph tools advertise `readOnlyHint: true`, which lets a client run them without prompting.
`check` is marked destructive instead, because a test suite can do anything its tests do.
`impact` lists at most `limit` callers and `limit` tests (100 by default, up to 1000 over MCP). When
it found more, it says so: `truncated` is true and `omitted_callers` / `omitted_tests` give the
counts (the CLI prints "... and N more callers not listed"). Each test comes with its pytest node id
(`tests/test_x.py::test_y`, or `tests/test_x.py::TestZ::test_y` for a method), found by pytest's
default collection rules: `test*` functions at module level in `test_*.py` and `*_test.py` files,
`test*` methods of `Test*` classes that have no `__init__`, and every `test*` method of a
`unittest.TestCase` subclass whatever its name. A class runs the test methods it inherits, so a
method defined on a base is listed once for each class that runs it, under that class's node id
(`TestChild::test_shared`). Pass those node ids to `check` to run just those tests. Helpers in test
files that are not tests themselves are listed with the other callers, since they may need updating.
`check(tests=[...])` accepts only test file paths and node ids inside the repository. Anything that
starts with `-` or `@` (pytest reads options from an `@file`) or points outside the repo is refused,
so a caller cannot slip in a pytest option (`--basetemp=DIR` deletes `DIR`). Output is read from a temporary file and the timeout covers the
whole run, so a test that leaves a process running in the background cannot hold up the result.
The summary and each failure or output line `check` returns are cut to 240 characters, ending in
`[truncated N chars]` when something was cut.
Example, trimmed:
```
$ ripple impact tax_for --repo bench/fixtures/shop
impact of shop.tax.tax_for (depth 5): 6 callers
shop/pricing.py
shop.pricing.price_order line 8 (1 hop)
shop/service.py
shop.service.OrderService.place_order line 17 (2 hops)
shop.service.OrderService.reprice line 26 (2 hops)
tests that reach it (3):
tests/test_pricing.py::test_percent_discount_and_tax
...
$ ripple check tests/test_pricing.py::test_percent_discount_and_tax --repo bench/fixtures/shop
{
"passed": true,
"summary": "1 passed in 0.01s",
...
```
## How it works
Two passes over the standard library `ast` module, no other parser.
1. **Collect** (`ripple/indexer.py`): each file is read on its own. For every module, class and
function it records what names are bound (defs, imports, assignments and what they suggest about
a value's class) and every call site.
2. **Resolve** (`ripple/resolver.py`): with the whole repository in view, each call site either
becomes an edge, because the source proves its target, or is kept as *unresolved* with a reason.
ripple never guesses. "There is only one method named `save`, so it must be that one" is exactly
the kind of answer that looks right and is wrong in the cases an agent cannot check.
What counts as proof: names defined or imported in scope; `import x`, `from x import y`, aliases,
relative imports and package re-exports; `self.method()` through base classes; `super()`; calling a
class (edge to its `__init__`); and method calls on values whose class is known from a constructor
call, a type annotation, a return annotation or a `self.attr` assigned in the class. When a method is
called on `self` or a typed value, subclass overrides are linked too (class hierarchy analysis).
The index is cached outside your repo (`RIPPLE_CACHE_DIR` if set, else `$XDG_CACHE_HOME/ripple` or
`~/.cache/ripple`) and keyed on every file's mtime and size, so each query re-checks the files and
rebuilds only when something changed.
On a 105-file package (pydantic) a cold index took 0.8 s and a cached load 0.02 s on my machine.
## Benchmark
`python -m bench` scores ripple against hand-written ground truth: 39 caller, callee and impact
queries over three small fixture apps in `bench/fixtures` (an order and pricing service, a notes CLI,
a web handler layer). CI fails if any number drops below `bench/baseline.json`.
| Case | Queries | Exact | Precision | Recall |
|---|---:|---:|---:|---:|
| Imports, aliases, re-exports | 11 | 11 | 1.000 | 1.000 |
| Methods, `super()`, constructors | 6 | 6 | 1.000 | 1.000 |
| Typed receivers | 5 | 5 | 1.000 | 1.000 |
| Callees | 3 | 2 | 1.000 | 0.750 |
| Decorators | 4 | 2 | 0.667 | 0.500 |
| Dynamic (getattr, dispatch tables, untyped values) | 6 | 0 | 1.000 | 0.125 |
| Impact (transitive) | 4 | 1 | 1.000 | 0.459 |
| **Overall** | **39** | **27** | **0.985** | **0.670** |
How to read this honestly:
- **The truth was written before the code.** The fixtures and every expected answer were committed
before any indexer code existed (see the first commit), and the dynamic cases were included on
purpose to show where static analysis stops.
- **The truth is checked against runtime.** `bench/trace.py` runs each fixture's own test suite under
a profiler, records the calls that actually happen, and `tests/test_ground_truth.py` fails if any
expected answer disagrees with that trace. The tracer shares no code with ripple.
- **It is small.** 39 queries over three fixtures I wrote, each query aimed at one resolution rule or
one known limit. It shows which rules work, not what recall looks like on your codebase, where
more of the code is likely to be untyped and dynamic.
- **The one false positive is a design choice.** ripple treats a decorated function as called
directly, so it reports a test that calls `hello()` as a caller of `hello`. At runtime the test
calls the decorator's wrapper, which calls `hello`.
## Agent eval
The benchmark checks that ripple's answers are right. The eval checks whether a coding agent does
better with them: ten refactoring tasks, each run by Claude Code in headless mode, once with a
ripple MCP server and once without, three times each.
- **The tasks** (`eval/tasks/`) are renames and signature changes where grep is noisy or
incomplete: a method named `total` on four classes, a function reached only through a package
re-export alias, an override in another package, a `getattr` lookup that neither grep nor ripple
can see. Seven run over `eval/fixtures/ledger`, a small invoicing app written for the eval, and
three reuse the benchmark fixtures. Each task's `why` field says what makes it hard.
- **Grading.** An edit counts as a success when the fixture's own tests pass and still number at
least the task's minimum (`min_tests`), and a hidden test the agent never sees passes too. The
hidden test checks that the change was made and that every caller still works.
`tests/test_eval.py` shows each check fails on the untouched fixture and passes on a reference
solution, and CI runs it.
- **Same conditions.** Both modes get the same prompt, model, tools and permissions (edit files,
run pytest). The MCP server is the only difference. Each run starts from a fresh copy under
`/tmp`, in a directory whose name says nothing about the task or mode, and leaves out the user's
own settings, hooks and auto-memory. A with-ripple run whose ripple server fails to connect is
marked invalid and left out of the numbers.
```sh
.venv/bin/python -m eval.run # 10 tasks x 2 modes x 3 repeats, into eval/results/
.venv/bin/python -m eval.run --tasks 04-ledger-rename-invoice-total --repeats 1 --out /tmp/smoke.json
```
Each batch writes `eval/results/<date>-<time>-<id>.json` (unless `--out` says otherwise), a `.md`
table and a `-diffs/` directory beside it, and refuses to overwrite an earlier batch's files.
It calls the `claude` CLI on whatever account it is logged in to, so it never runs in CI.
### Results
First full run, 2026-09-26, Claude Sonnet 5, 60 runs (`eval/results/2026-09-26.json`):
| Mode | Success | Median input tokens | Median output tokens | Median turns | Median wall time |
|---|---:|---:|---:|---:|---:|
| with ripple | 30/30 | 145,461 | 3,372 | 19.5 | 25.2 s |
| without | 30/30 | 129,108 | 3,211 | 15.0 | 22.9 s |
**ripple did not help on these tasks.** Both modes solved every task every time, so there is no
difference in success to measure, and with ripple the agent spent about 13% more input tokens and
four or five more turns, mostly on loading and calling the tools (78 ripple calls across the 30
runs).
The likely reason is the tasks, not the tool: the fixtures are a few hundred lines, small enough
that the agent can grep and read everything it needs, which puts both modes at the ceiling. The
cases where a call graph should matter (a large unfamiliar codebase, callers the agent never opens)
are exactly the ones this first task set does not have. The next version of the eval needs tasks
over bigger, real repositories, where the without-ripple agent has a real chance to miss a caller.
In this run each arm's working directory name included its condition label (such as
`--with-ripple--`), which the agent can see; that is fixed now, and the result is unlikely to depend
on it since both arms scored 30/30.
## Known limits
Static analysis of a dynamic language misses things. ripple does not follow:
- method calls on values of unknown type: unannotated parameters, loop variables, results of
functions without a return annotation (reported as unresolved, with the reason)
- `getattr`, dict or registry dispatch, callbacks passed as arguments, `functools.partial`
- decorator wrappers (the decorated function is treated as called directly)
- calls on the result of a call (`make().run()`), property access, `__call__`, metaclasses,
monkeypatching, `importlib`, `exec`
- code imported under a name that differs from its path in the repo (modules are named by their
path from the repo root, with a leading `src/` dropped). Two files that get the same name
(`tests/test_x.py` and `src/tests/test_x.py`) share one set of symbol ids, so a name defined in
both is one symbol; `ripple index` warns when that happens and names the files.
- directories that hold dependencies or build output: hidden directories, `__pycache__`,
`node_modules`, `site-packages`, `venv` and any directory with a `pyvenv.cfg` anywhere, and `build`,
`dist` and `env` at the repo root only. `ripple index` lists the skipped directories, except
hidden ones and `__pycache__`.
- custom pytest collection settings (`python_files`, `python_functions`, `python_classes`): test node
ids follow pytest's defaults. ripple sees only the repo's own classes, so a class counts as a
`unittest.TestCase` subclass when it has a base from outside the repo whose name ends in `TestCase`
(`unittest.TestCase`, `django.test.TestCase`), and an `__init__` inherited from outside the repo
is not seen.
Override edges can over-approximate: a subclass override is linked even if that subclass never
reaches the call site. Method lookup walks bases depth-first, which matches Python's C3 order except
in diamond-shaped hierarchies.
To keep the blind spots visible, `callees` lists every unresolved call with its reason, and `impact`
lists unresolved call sites that use the same name as the target under `possible_missed_callers`.
A call counts as "code outside the repo" only when it must be: a method missing from a class's repo
bases is inherited from an external base such as `Exception`, unless a base could not be resolved
(`class Cached(make_base())`) or a subclass defines the method, and then it is unresolved too.
Subscripted bases such as `Repo[int]` or `Generic[T]` resolve to the class they subscript.
## Next
- **TypeScript**, the other language agents spend most of their time in.
## Development
```sh
uv venv && uv pip install -e ".[dev]"
.venv/bin/python -m pytest -q # unit tests, plus the ground truth against its runtime trace
.venv/bin/python -m bench # the benchmark; --write-baseline after a deliberate change
```
MIT licensed. Copyright (c) 2026 Dylan Fodor.
This server cannot be deployed
Maintenance
ActivityMaintained
ResponsivenessNo issues