Skip to main content
Glama

ripple

Know what breaks before you change it. ripple is an MCP server that gives a coding agent (Claude Code, Cursor, anything that speaks MCP) the call graph of a Python repository, so the agent can ask "who calls this, and which tests reach it?" before it edits, and run just those tests after.

Why

A coding agent dropped into unfamiliar code edits blind. It greps for a function name, sees a few hits, and changes the signature. The callers that reach it through an import alias, a self. call in a subclass or a re-export in a package __init__ don't match the grep, and the first sign of trouble is a failing test, or a failure in production. ripple answers the structural question directly and says plainly where static analysis runs out.

Related MCP server: repo-atlas-mcp

Quickstart

Requires Python 3.10+ and uv.

uv tool install git+https://github.com/xtelos/ripple   # puts `ripple` on your PATH

Claude Code, from your project directory:

claude mcp add ripple -- ripple serve --repo "$PWD"

Cursor, in .cursor/mcp.json:

{
  "mcpServers": {
    "ripple": { "command": "ripple", "args": ["serve", "--repo", "${workspaceFolder}"] }
  }
}

Or from a terminal: ripple index ., then ripple impact price_order.

The five tools

Tool

What it answers

Read-only

callers(symbol)

Who calls this directly, with the file:line of each call site

yes

callees(symbol)

What this calls, plus the calls ripple could not resolve and why

yes

impact(symbol, depth=5, limit=100)

Every transitive caller grouped by file, the tests that reach it listed separately as pytest node ids, and same-named call sites ripple could not rule out

yes

path(source, target)

The shortest call chain from one symbol to another

yes

check(tests=None)

Runs the test command (default python -m pytest -q, using the repo's .venv if present), or just the given tests, and returns pass/fail, the summary line and the first failures, never the full log

no, it runs code

A symbol can be a full dotted path (shop.pricing.price_order), a suffix (OrderService.place_order) or a bare name (price_order). An ambiguous name returns the candidates rather than picking one.

The four graph tools advertise readOnlyHint: true, which lets a client run them without prompting. check is marked destructive instead, because a test suite can do anything its tests do.

impact lists at most limit callers and limit tests (100 by default, up to 1000 over MCP). When it found more, it says so: truncated is true and omitted_callers / omitted_tests give the counts (the CLI prints "... and N more callers not listed"). Each test comes with its pytest node id (tests/test_x.py::test_y, or tests/test_x.py::TestZ::test_y for a method), found by pytest's default collection rules: test* functions at module level in test_*.py and *_test.py files, test* methods of Test* classes that have no __init__, and every test* method of a unittest.TestCase subclass whatever its name. A class runs the test methods it inherits, so a method defined on a base is listed once for each class that runs it, under that class's node id (TestChild::test_shared). Pass those node ids to check to run just those tests. Helpers in test files that are not tests themselves are listed with the other callers, since they may need updating.

check(tests=[...]) accepts only test file paths and node ids inside the repository. Anything that starts with - or @ (pytest reads options from an @file) or points outside the repo is refused, so a caller cannot slip in a pytest option (--basetemp=DIR deletes DIR). Output is read from a temporary file and the timeout covers the whole run, so a test that leaves a process running in the background cannot hold up the result. The summary and each failure or output line check returns are cut to 240 characters, ending in [truncated N chars] when something was cut.

Example, trimmed:

$ ripple impact tax_for --repo bench/fixtures/shop
impact of shop.tax.tax_for (depth 5): 6 callers
  shop/pricing.py
    shop.pricing.price_order  line 8  (1 hop)
  shop/service.py
    shop.service.OrderService.place_order  line 17  (2 hops)
    shop.service.OrderService.reprice  line 26  (2 hops)
tests that reach it (3):
  tests/test_pricing.py::test_percent_discount_and_tax
  ...

$ ripple check tests/test_pricing.py::test_percent_discount_and_tax --repo bench/fixtures/shop
{
  "passed": true,
  "summary": "1 passed in 0.01s",
  ...

How it works

Two passes over the standard library ast module, no other parser.

  1. Collect (ripple/indexer.py): each file is read on its own. For every module, class and function it records what names are bound (defs, imports, assignments and what they suggest about a value's class) and every call site.

  2. Resolve (ripple/resolver.py): with the whole repository in view, each call site either becomes an edge, because the source proves its target, or is kept as unresolved with a reason. ripple never guesses. "There is only one method named save, so it must be that one" is exactly the kind of answer that looks right and is wrong in the cases an agent cannot check.

What counts as proof: names defined or imported in scope; import x, from x import y, aliases, relative imports and package re-exports; self.method() through base classes; super(); calling a class (edge to its __init__); and method calls on values whose class is known from a constructor call, a type annotation, a return annotation or a self.attr assigned in the class. When a method is called on self or a typed value, subclass overrides are linked too (class hierarchy analysis).

The index is cached outside your repo (RIPPLE_CACHE_DIR if set, else $XDG_CACHE_HOME/ripple or ~/.cache/ripple) and keyed on every file's mtime and size, so each query re-checks the files and rebuilds only when something changed. On a 105-file package (pydantic) a cold index took 0.8 s and a cached load 0.02 s on my machine.

Benchmark

python -m bench scores ripple against hand-written ground truth: 39 caller, callee and impact queries over three small fixture apps in bench/fixtures (an order and pricing service, a notes CLI, a web handler layer). CI fails if any number drops below bench/baseline.json.

Case

Queries

Exact

Precision

Recall

Imports, aliases, re-exports

11

11

1.000

1.000

Methods, super(), constructors

6

6

1.000

1.000

Typed receivers

5

5

1.000

1.000

Callees

3

2

1.000

0.750

Decorators

4

2

0.667

0.500

Dynamic (getattr, dispatch tables, untyped values)

6

0

1.000

0.125

Impact (transitive)

4

1

1.000

0.459

Overall

39

27

0.985

0.670

How to read this honestly:

  • The truth was written before the code. The fixtures and every expected answer were committed before any indexer code existed (see the first commit), and the dynamic cases were included on purpose to show where static analysis stops.

  • The truth is checked against runtime. bench/trace.py runs each fixture's own test suite under a profiler, records the calls that actually happen, and tests/test_ground_truth.py fails if any expected answer disagrees with that trace. The tracer shares no code with ripple.

  • It is small. 39 queries over three fixtures I wrote, each query aimed at one resolution rule or one known limit. It shows which rules work, not what recall looks like on your codebase, where more of the code is likely to be untyped and dynamic.

  • The one false positive is a design choice. ripple treats a decorated function as called directly, so it reports a test that calls hello() as a caller of hello. At runtime the test calls the decorator's wrapper, which calls hello.

Agent eval

The benchmark checks that ripple's answers are right. The eval checks whether a coding agent does better with them: ten refactoring tasks, each run by Claude Code in headless mode, once with a ripple MCP server and once without, three times each.

  • The tasks (eval/tasks/) are renames and signature changes where grep is noisy or incomplete: a method named total on four classes, a function reached only through a package re-export alias, an override in another package, a getattr lookup that neither grep nor ripple can see. Seven run over eval/fixtures/ledger, a small invoicing app written for the eval, and three reuse the benchmark fixtures. Each task's why field says what makes it hard.

  • Grading. An edit counts as a success when the fixture's own tests pass and still number at least the task's minimum (min_tests), and a hidden test the agent never sees passes too. The hidden test checks that the change was made and that every caller still works. tests/test_eval.py shows each check fails on the untouched fixture and passes on a reference solution, and CI runs it.

  • Same conditions. Both modes get the same prompt, model, tools and permissions (edit files, run pytest). The MCP server is the only difference. Each run starts from a fresh copy under /tmp, in a directory whose name says nothing about the task or mode, and leaves out the user's own settings, hooks and auto-memory. A with-ripple run whose ripple server fails to connect is marked invalid and left out of the numbers.

.venv/bin/python -m eval.run                   # 10 tasks x 2 modes x 3 repeats, into eval/results/
.venv/bin/python -m eval.run --tasks 04-ledger-rename-invoice-total --repeats 1 --out /tmp/smoke.json

Each batch writes eval/results/<date>-<time>-<id>.json (unless --out says otherwise), a .md table and a -diffs/ directory beside it, and refuses to overwrite an earlier batch's files.

It calls the claude CLI on whatever account it is logged in to, so it never runs in CI.

Results

First full run, 2026-09-26, Claude Sonnet 5, 60 runs (eval/results/2026-09-26.json):

Mode

Success

Median input tokens

Median output tokens

Median turns

Median wall time

with ripple

30/30

145,461

3,372

19.5

25.2 s

without

30/30

129,108

3,211

15.0

22.9 s

ripple did not help on these tasks. Both modes solved every task every time, so there is no difference in success to measure, and with ripple the agent spent about 13% more input tokens and four or five more turns, mostly on loading and calling the tools (78 ripple calls across the 30 runs).

The likely reason is the tasks, not the tool: the fixtures are a few hundred lines, small enough that the agent can grep and read everything it needs, which puts both modes at the ceiling. The cases where a call graph should matter (a large unfamiliar codebase, callers the agent never opens) are exactly the ones this first task set does not have. The next version of the eval needs tasks over bigger, real repositories, where the without-ripple agent has a real chance to miss a caller.

In this run each arm's working directory name included its condition label (such as --with-ripple--), which the agent can see; that is fixed now, and the result is unlikely to depend on it since both arms scored 30/30.

Known limits

Static analysis of a dynamic language misses things. ripple does not follow:

  • method calls on values of unknown type: unannotated parameters, loop variables, results of functions without a return annotation (reported as unresolved, with the reason)

  • getattr, dict or registry dispatch, callbacks passed as arguments, functools.partial

  • decorator wrappers (the decorated function is treated as called directly)

  • calls on the result of a call (make().run()), property access, __call__, metaclasses, monkeypatching, importlib, exec

  • code imported under a name that differs from its path in the repo (modules are named by their path from the repo root, with a leading src/ dropped). Two files that get the same name (tests/test_x.py and src/tests/test_x.py) share one set of symbol ids, so a name defined in both is one symbol; ripple index warns when that happens and names the files.

  • directories that hold dependencies or build output: hidden directories, __pycache__, node_modules, site-packages, venv and any directory with a pyvenv.cfg anywhere, and build, dist and env at the repo root only. ripple index lists the skipped directories, except hidden ones and __pycache__.

  • custom pytest collection settings (python_files, python_functions, python_classes): test node ids follow pytest's defaults. ripple sees only the repo's own classes, so a class counts as a unittest.TestCase subclass when it has a base from outside the repo whose name ends in TestCase (unittest.TestCase, django.test.TestCase), and an __init__ inherited from outside the repo is not seen.

Override edges can over-approximate: a subclass override is linked even if that subclass never reaches the call site. Method lookup walks bases depth-first, which matches Python's C3 order except in diamond-shaped hierarchies.

To keep the blind spots visible, callees lists every unresolved call with its reason, and impact lists unresolved call sites that use the same name as the target under possible_missed_callers. A call counts as "code outside the repo" only when it must be: a method missing from a class's repo bases is inherited from an external base such as Exception, unless a base could not be resolved (class Cached(make_base())) or a subclass defines the method, and then it is unresolved too. Subscripted bases such as Repo[int] or Generic[T] resolve to the class they subscript.

Next

  • TypeScript, the other language agents spend most of their time in.

Development

uv venv && uv pip install -e ".[dev]"
.venv/bin/python -m pytest -q      # unit tests, plus the ground truth against its runtime trace
.venv/bin/python -m bench          # the benchmark; --write-baseline after a deliberate change

MIT licensed. Copyright (c) 2026 Dylan Fodor.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    A
    maintenance
    Provides a dependency graph of any local repository with tools for change impact, transitive dependents, health audits, and more, enabling AI coding agents to see structure and refactor safely.
    4,912 npm
    4
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Provides a graph-based code analysis to coding agents, enabling tools to trace call dependencies, assess change impact, and identify hotspots using tree-sitter and Neo4j.
    10
    6 npm
    MIT
  • A
    license
    B
    quality
    B
    maintenance
    Enables AI agents to index local Python repositories and retrieve code intelligence through semantic search, caller lookups, hybrid context expansion, and visualization.
    5
    MIT