Skip to main content
Glama
README.md
# DataBench — Benchmark Smarter with AI Agents

Automates end-to-end data platform benchmarking: job submission, status tracking, result collection, cost analysis, and report generation — through natural language or a visual dashboard.

## Quick Start

```bash
git clone ssh://git.amazon.com/pkg/DataBench
cd DataBench
pip install -r requirements.txt
```

### Run the Dashboard

```bash
PYTHONPATH=.. streamlit run ui/app.py
```

Opens at http://localhost:8501 with:
- **Dashboard** — ClickBench-inspired ranking of all benchmark runs (GPU vs CPU, EMR vs OSS, Parquet vs Iceberg)
- **Report Builder** — Upload driver logs or CSVs, compare runs, export JSON
- **Submit Jobs** — Pick a template, configure, launch TPC-DS benchmarks
- **Monitor** — Track running EMR jobs in real-time

### Run Tests

```bash
python -m pytest tests/ -v
```

### Use as an MCP Server in Claude Code

DataBench ships an MCP server (`databench.mcpserver.server`) that exposes the
benchmark tools — `benchmark-list-templates`, `benchmark-submit`,
`benchmark-status`, `benchmark-collect`, `benchmark-cost`, `benchmark-compare`,
`benchmark-report`, and more — directly to Claude Code.

First install the package so `databench` is importable:

```bash
pip install -e .
```

Then register the server (`--scope user` makes it available in all your projects):

```bash
claude mcp add databench --scope user \
  -e PYTHONPATH=/path/to/sourcecode/DataBench \
  -- python -m databench.mcpserver.server
```

Use the absolute path to the `python` interpreter where you ran
`pip install -e .` (find it with `which python`) so Claude Code launches the
server with the right environment. Verify it connected:

```bash
claude mcp get databench   # Status: ✔ Connected
```

Restart your Claude Code session and the `databench` tools will be available
(check with `/mcp` inside Claude, or `claude mcp list`). To remove it:

```bash
claude mcp remove databench -s user
```

### Talk to it in plain English

Once the MCP server is connected, you drive the whole workflow — provisioning
nodegroups, running benchmarks, and generating reports — by just asking Claude
Code in natural language. Claude picks the right templates, creates the EKS
managed nodegroups (RAID0 local disk, cluster-autoscaler wiring, pod-template
node pinning), submits the jobs, polls them, and produces the report. Real
examples:

**Provision infrastructure**

- *"Create managed nodegroups in the loadtest-mcp EKS cluster for r7g.4xlarge
  and r8g.4xlarge, scaling from 1 to 8, each with 4×64GB EBS volumes striped as
  RAID0 mounted at /var/data."*
- *"Create r7gd.4xlarge and r8gd.4xlarge nodegroups and mount the local NVMe as
  the Spark spill dir."*
- *"Install the cluster autoscaler so it can scale only these benchmark
  nodegroups up and down — don't let it touch the ops nodegroup or Karpenter."*
- *"Bump all four benchmark nodegroups to max 8 nodes."*
- *"Remove the 8xlarge nodegroups but keep the templates."*

**Run benchmarks**

- *"Run TPC-DS 3TB comparing r7g vs r8g on 8 nodes for EMR on EKS 7.12 Spark
  performance."*
- *"Benchmark TPC-DS 3TB across r7g.4xl, r8g.4xl, r7gd.4xl and r8gd.4xl, then
  compare the results."*
- *"Show me the template setup first, then run r7g.8xlarge vs r8g.8xlarge on 4
  nodes — keep total CPU and memory the same as the 4xlarge run."*
- *"Split the big executor into 6 smaller pods per node without changing total
  CPU or memory, then re-run."*

**Check status & get reports**

- *"What's the status of the r7gd benchmark jobs?"*
- *"Compare all of r7g, r7gd, r8g and r8gd and give me a downloadable report."*
- *"Generate the Spark cost-performance comparison report."*

Claude handles the mechanics behind these — picking `benchmark-list-templates`,
`benchmark-submit`, `benchmark-status`, `benchmark-compare`, and
`benchmark-report`, plus the `eksctl`/`kubectl`/`aws` steps for nodegroup
lifecycle — and asks for confirmation before anything that costs money.

## What's Inside

```
DataBench/
├── models/
│   └── benchmark_result.py    # BenchmarkResult, QueryResult, ComparisonResult
├── tools/
│   ├── collect.py             # Parse driver logs, CSV, JSON → BenchmarkResult
│   ├── compare.py             # Speedup calculator: compare_runs() → ComparisonResult
│   ├── cost.py                # Instance pricing lookup + cost-per-run calculation
│   ├── report.py              # xlsx report generator (Summary, Query Details, Configs)
│   └── status.py              # Cross-platform job status poller
├── adapters/
│   ├── emr_eks.py             # EMR on EKS: submit jobs, check status
│   ├── emr_ec2.py             # EMR on EC2: submit steps
│   └── athena.py              # Athena: submit TPC-DS queries
├── configs/                   # 9 parameterized benchmark templates
├── ui/
│   └── app.py                 # Streamlit dashboard
├── tests/
│   ├── test_collect.py
│   ├── test_compare.py
│   ├── test_compare_real_data.py
│   ├── test_cost.py
│   ├── test_report.py
│   ├── test_athena_submit.py
│   └── sample_benchmark_result.json
├── conftest.py                # pytest import fix (DataBench → databench)
└── README.md
```

## Benchmark Results (baked into dashboard)

| Workstream | Runs | Key Finding |
|------------|------|-------------|
| GPU vs CPU Parquet | 6 | g6 GPU 2.9x faster AND 56% cheaper than CPU |
| EMR vs OSS Spark | 2 | EMR 3.4x faster, 70% cheaper |
| Iceberg GPU vs CPU | 2 | GPU 1.6x faster with split tuning |
| S3 Tables | 2 | GPU 2.1x faster on S3 Tables |
| Velox/Gluten | 3 | Velox 1.6x faster than baseline (10TB) |
| g7 Standalone | 2 | Thread tuning: 683s → 534s (22% faster) |

## Interface Contract

Every tool produces/consumes `BenchmarkResult`:

```python
from databench.models import BenchmarkResult

result = BenchmarkResult.from_json(open("result.json").read())
print(result.total_median())        # 534.0 seconds
print(result.median_time("q1"))     # 4.5 seconds
print(result.query_names())         # ["q1", "q2", ...]
```

## Compare & Report

```python
from databench.tools.compare import compare_runs
from databench.tools.report import generate_report

# Compare GPU vs CPU
comparison = compare_runs([gpu_result, cpu_result], baseline_run_id="cpu-run")
print(comparison.aggregates)  # speedups, wins, cost savings

# Generate xlsx report
report = generate_report([gpu_result, cpu_result], baseline_run_id="cpu-run")
print(report["file_path"])    # databench-report-gpu-vs-cpu.xlsx
```

## Cost Lookup

```python
from databench.tools.cost import calculate_cost

cost = calculate_cost("g6.4xlarge", node_count=8, seconds=534.0, region="us-east-1")
print(cost)  # {"price_per_hour": 1.323, "cluster_cost_per_hour": 10.584, "cost_per_run": 1.57}
```

## Config Templates

```python
from databench.configs import list_templates, load_template

# List all templates
for t in list_templates():
    print(f"{t['template_id']:30s} GPU={t['gpu']}")

# Load with variable resolution
cfg = load_template("gpu-parquet-g6-4xl", bucket="my-bucket", region="us-east-1")
```

## AWS Credentials

The dashboard uses your AWS CLI profile. Set it in the sidebar or:

```bash
ada credentials update --profile aws-emr-bda-admin --provider isengard --once
```

## Team

- Karthik Prabhakar (subbakk) — EMR adapters, orchestration, config templates, UI
- Pathik Shah (pathshah) — Report generation, comparison, cost, Athena adapter

## Links

- [Architecture Doc (Quip)](https://quip-amazon.com/UBYHAuloKUF2)
- [Taskei Epic](https://taskei.amazon.dev/tasks/8e0b732b-ad81-4dbf-948f-f4735459b28b)