Skip to main content
Glama
README.md
# ripple

A triage tool for broken data tables in [DataHub](https://github.com/datahub-project/datahub). Give it a table that's broken and it shows you:

- every table, chart, and dashboard downstream that's affected
- which of them matter most, and an overall severity (SEV1 to SEV3)
- who owns them and should be told

Then it records the incident on the table in DataHub.

**[Try the online demo](https://chakri192.github.io/ripple/)**, no setup needed.

<img src="docs/blast-radius.svg" alt="ripple dashboard" width="860">

## Requirements

- A running DataHub (a local `datahub docker quickstart` works)
- Python 3.10+
- A DataHub access token

## Install

```sh
git clone https://github.com/chakri192/ripple.git
cd ripple
pip install -e .
cp .env.example .env      # add your DataHub token
```

## Try it

Load a sample setup into DataHub (an orders table feeding five tables and three dashboards), then triage it:

```sh
python demo/seed_incident_demo.py
ripple triage "urn:li:dataset:(urn:li:dataPlatform:snowflake,prod.raw.orders_raw,PROD)" --no-write-back
```

## Commands

| Command | |
|---|---|
| `ripple triage URN` | Find everything affected, rank it, and record the incident in DataHub |
| `ripple triage URN --no-write-back` | Report only, don't change DataHub |
| `ripple triage URN --columns` | Also show which columns are affected |
| `ripple triage URN --incident` | Also create a DataHub incident |
| `ripple root-cause URN` | Look upstream for where bad data probably came from |
| `ripple watch` | Watch for tables tagged `broken` and triage them automatically |
| `ripple web` | Open the dashboard at http://localhost:8000 |

Each triage prints a summary in the terminal and saves a Markdown report in `examples/_scratch/`.

## Severity

- **SEV1**: a dashboard or chart is affected (tag a dashboard `internal` in DataHub if it isn't customer-facing)
- **SEV2**: something is affected, but nothing customer-facing
- **SEV3**: nothing downstream is affected

## What it changes in DataHub

- Adds an `incident` tag to the broken table
- Adds the incident report to the table's description, below anything already written there (running triage again replaces only the report)
- With `--incident` (and always in `watch` mode), creates a DataHub incident

Use `--no-write-back` if you don't want DataHub changed at all.

## Settings

In `.env`:

| Variable | |
|---|---|
| `DATAHUB_GMS_URL` | DataHub address (default `http://localhost:8080`) |
| `DATAHUB_GMS_TOKEN` | Your access token |
| `ANTHROPIC_API_KEY` | Optional. Writes a more readable report using Claude. Without it, a standard template is used |

## License

Apache-2.0

## Contributors

| | |
|---|---|
| [chakri192](https://github.com/chakri192) | Author |
| [aider](https://github.com/Aider-AI/aider) | AI pair programmer |