dagsmith
README.md
# dagsmith
**Agentic scaffolding for data engineering.** Turn a plain-English description of a new data source into a complete, tested, production-shaped pipeline: dbt models (staging → marts), an Airflow DAG, a Great Expectations suite, and freshness monitors — generated, executed, and self-corrected by an LLM agent that reads your warehouse's live metadata over MCP.
```
┌───────────────────────────────────────────────────────────────────────┐
│ │
│ analyst writes: │
│ "Ingest daily registry events from Postgres. Grain = one event │
│ per registry_id + item_id. Join to customers for LTV rollups." │
│ │
│ │ │
│ ▼ │
│ ┌─────────────────────┐ │
│ │ dagsmith agent │ │
│ │ (Claude Sonnet 5) │ │
│ └──┬──────────────┬───┘ │
│ │ │ │
│ MCP tools ───────▶│ │──── Jinja templates │
│ (Snowflake meta) │ │ + verifier loop │
│ ▼ ▼ │
│ ┌─────────────────┐ ┌─────────────────┐ │
│ │ dbt models │ │ Airflow DAG │ │
│ │ stg_*.sql │ │ registry_ev.py │ │
│ │ fct_*.sql │ │ │ │
│ │ schema.yml │ │ GE checkpoint │ │
│ └─────────────────┘ └─────────────────┘ │
│ │ │ │
│ └──────┬───────┘ │
│ ▼ │
│ ┌─────────────────────┐ │
│ │ verifier │ │
│ │ (real Snowflake + │ │
│ │ sandbox dbt run) │ │
│ └──────┬──────────────┘ │
│ │ │
│ pass ──┴── fail → loop back to agent │
│ │
└───────────────────────────────────────────────────────────────────────┘
```
## Why this exists
Every data engineer builds the same eight files for every new source: a staging model, a mart, a schema.yml, a snapshot config, a DAG, a GE suite, a monitor, a docs entry. Ninety percent of that work is mechanical — but it's not homogeneous enough to templatize with pure code. That gap is what an agent covers well.
`dagsmith` treats pipeline generation as a **compile step**, not a checklist. You describe the intent; the agent generates the artifacts; a verifier runs them against a sandbox schema in real Snowflake; failures loop back into the agent's context until they pass. When you merge, you get real tested code — not a demo.
## Design in one paragraph
The agent is a Claude Sonnet 5 tool-use loop with three tool categories: **discovery** (MCP calls into Snowflake's `INFORMATION_SCHEMA` and `ACCOUNT_USAGE` so the agent can inspect what already exists), **generation** (write files from Jinja templates), and **verification** (run `dbt build --target ci` and Great Expectations checkpoints against an isolated sandbox schema). The verifier's stdout+stderr feed back into the agent's context, and it iterates until green or hits a retry cap.
## What's in the box
| Path | What |
|---|---|
| `src/dagsmith/agent.py` | The tool-use loop. Sonnet 5, ~200 lines. |
| `src/dagsmith/mcp_server.py` | MCP server exposing 12 read-only Snowflake metadata tools. Runs standalone; usable from Claude Desktop, Cursor, or any MCP client. |
| `src/dagsmith/verifier.py` | Sandbox executor — copies templates to a temp dbt project, runs `dbt build`, streams results. |
| `src/dagsmith/templates/*.j2` | Jinja templates the agent fills. Kept intentionally minimal so the agent does real design work. |
| `src/dagsmith/prompts/system.md` | The system prompt with the modeling conventions (Kimball for marts, insert-only for events, `stg_<source>__<table>` naming). |
| `src/dagsmith/cli.py` | `dagsmith generate` and `dagsmith mcp`. |
| `examples/babylist_registry_events/` | Full end-to-end example — natural-language input + everything the agent produced. |
| `docs/architecture.md` | Detailed component design + failure modes. |
| `docs/mcp_tools.md` | Reference for every MCP tool the agent has. |
## Quickstart
```bash
git clone https://github.com/<you>/dagsmith && cd dagsmith
pip install -e ".[dev]"
# 1. Point at your Snowflake sandbox (see docs/setup.md)
export SNOWFLAKE_ACCOUNT=abc12345.us-east-1.aws
export SNOWFLAKE_USER=you@example.com
export SNOWFLAKE_ROLE=DATA_ENG
export SNOWFLAKE_WAREHOUSE=DEV_WH
export ANTHROPIC_API_KEY=sk-ant-...
# 2. Generate a pipeline from a spec
dagsmith generate examples/babylist_registry_events/input/spec.md \
--out examples/babylist_registry_events/generated \
--sandbox-schema DAGSMITH_CI_TAYLOR
# 3. Or run the MCP server for use in Claude Desktop / Cursor / Zed
dagsmith mcp
```
## MCP server (usable on its own)
Even if you don't run the agent, the MCP server is useful. Point any MCP client at it and Claude/Cursor/Zed gains 12 read-only tools:
- `list_databases`, `list_schemas`, `list_tables`
- `describe_table` (columns, types, cluster keys, comments)
- `sample_rows` (limit-capped preview)
- `query_history` (top N recent queries by cost / duration)
- `access_history` (who reads what)
- `column_lineage` (upstream columns for a target column via `ACCESS_HISTORY`)
- `dead_tables` (tables not read in N days)
- `storage_by_schema` and `credits_by_warehouse`
- `check_freshness` (last successful write per table)
MCP config for Claude Desktop:
```json
{
"mcpServers": {
"dagsmith": {
"command": "dagsmith",
"args": ["mcp"],
"env": {
"SNOWFLAKE_ACCOUNT": "abc12345.us-east-1.aws",
"SNOWFLAKE_USER": "you@example.com",
"SNOWFLAKE_ROLE": "READONLY_MC"
}
}
}
}
```
## What the agent will and won't do
**Will:** generate staging + mart dbt models with tests, an Airflow DAG that uses Cosmos, a Great Expectations checkpoint with reasonable defaults (uniqueness on the grain, row-count band, not-null on FKs), a freshness monitor query, and a docs.md entry.
**Won't:** run destructive DDL against production, guess at business logic it can't see, or generate code without verifying it compiles. Every generated file is `dbt parse`-clean before the agent claims done.
## Roadmap
- [ ] `dagsmith audit` — point at an existing dbt project and produce a report on coverage gaps (missing tests, stale models, over-warehouse'd queries)
- [ ] `dagsmith migrate` — port a legacy Fivetran → SQL pipeline into the dbt+Airflow layout the agent generates
- [ ] Evaluation harness — golden-output tests over 20+ source specs so regressions are catchable
- [ ] Backends beyond Snowflake — BigQuery + Redshift adapters share ~80% of the templates
## Why the JD's language fits
This project directly maps to the "harnesses and agentic scaffolding that generate, test, and maintain them" phrasing in the Babylist Senior Data Engineer role:
- **Meta-layer, not just pipelines** — the agent is a pipeline factory, not a pipeline.
- **MCP servers real users depend on** — the metadata server is usable standalone, not just as an agent input.
- **Agentic workflows with tool integrations** — 12+ tools, real verifier feedback loop.
- **Airflow + dbt + Snowflake + Python** — the whole stack the JD names, wired end-to-end.
- **Data monitoring across multi-system journeys** — freshness + GE suites are part of the generated bundle.
## License
MIT. See [LICENSE](LICENSE).
This server cannot be deployed
Maintenance
ActivitySlowing
ResponsivenessNo issues