Skip to main content
Glama

ripple

A triage tool for broken data tables in DataHub. Give it a table that's broken and it shows you:

  • every table, chart, and dashboard downstream that's affected

  • which of them matter most, and an overall severity (SEV1 to SEV3)

  • who owns them and should be told

Then it records the incident on the table in DataHub.

Try the online demo, no setup needed.

Requirements

  • A running DataHub (a local datahub docker quickstart works)

  • Python 3.10+

  • A DataHub access token

Related MCP server: semley

Install

git clone https://github.com/chakri192/ripple.git
cd ripple
pip install -e .
cp .env.example .env      # add your DataHub token

Try it

Load a sample setup into DataHub (an orders table feeding five tables and three dashboards), then triage it:

python demo/seed_incident_demo.py
ripple triage "urn:li:dataset:(urn:li:dataPlatform:snowflake,prod.raw.orders_raw,PROD)" --no-write-back

Commands

Command

ripple triage URN

Find everything affected, rank it, and record the incident in DataHub

ripple triage URN --no-write-back

Report only, don't change DataHub

ripple triage URN --columns

Also show which columns are affected

ripple triage URN --incident

Also create a DataHub incident

ripple root-cause URN

Look upstream for where bad data probably came from

ripple watch

Watch for tables tagged broken and triage them automatically

ripple web

Open the dashboard at http://localhost:8000

Each triage prints a summary in the terminal and saves a Markdown report in examples/_scratch/.

Severity

  • SEV1: a dashboard or chart is affected (tag a dashboard internal in DataHub if it isn't customer-facing)

  • SEV2: something is affected, but nothing customer-facing

  • SEV3: nothing downstream is affected

What it changes in DataHub

  • Adds an incident tag to the broken table

  • Adds the incident report to the table's description, below anything already written there (running triage again replaces only the report)

  • With --incident (and always in watch mode), creates a DataHub incident

Use --no-write-back if you don't want DataHub changed at all.

Settings

In .env:

Variable

DATAHUB_GMS_URL

DataHub address (default http://localhost:8080)

DATAHUB_GMS_TOKEN

Your access token

ANTHROPIC_API_KEY

Optional. Writes a more readable report using Claude. Without it, a standard template is used

License

Apache-2.0

Contributors

chakri192

Author

aider

AI pair programmer

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    D
    maintenance
    Enables autonomous SRE incident investigation by allowing users to describe incidents in natural language. The agent follows a governed state machine to gather read-only evidence and produce grounded conclusions.
    MIT
  • A
    license
    A
    quality
    C
    maintenance
    Enables agents to resolve data incidents by preparing remediation, inspecting verification results, and staging for human approval, while integrating with DataHub for evidence.
    5
    Apache 2.0
  • A
    license
    A
    quality
    B
    maintenance
    An on-call agent for data incidents, built on a bio-inspired context protocol. It triages data quality and freshness issues using DataHub, computing severity, blast radius, and ownership from the catalogue before generating prose.
    9
    Apache 2.0