dataloom
Connects to PostgreSQL databases to introspect schemas and generate synthetic test data directly into existing tables, with optional profiling and transactional inserts.
Generates synthetic datasets into local SQLite databases, including demo scenarios and deterministic replay of generated bundles.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@dataloomGenerate 50 synthetic patients with 3-5 orders each, deterministic replay"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
DataLoom
Alpha · Python 3.11+ · Local CLI, Python library, and MCP server
Connected data. Repeatable scenarios.
DataLoom turns a relational schema into an inspectable schema genome, then
generates synthetic test datasets from a versioned plan. Ask an MCP-connected
agent for a scenario, inspect the saved plan, and replay it in CI without an
LLM. Natural-language generation authors and executes the plan in one operation;
use a separately reviewed plan file when review must precede execution. Or write
the plan yourself and stay offline from the start. A schema on its own is enough to
begin: --auto derives a default plan, saves it for review, and executes it without
a provider.
The practical problem: realistic-looking individual records are easy. Test data with valid relationships, deliberate edge cases, repeatable distributions, and an explanation of what was exercised is harder. DataLoom puts those requirements in one small, extensible Python library.
This is an early Phase 1 implementation, not a claim of universal SQL support or clinical realism. The engine is domain-neutral; healthcare is the first domain pack. All examples are fictional and independently designed.
Try it
Requires Python 3.11+ and an internet connection for the initial dependency installation. From a checkout, run this on Windows, macOS, or Linux:
git clone https://github.com/anzar-ahsan-commits/dataloom.git
cd dataloom
python examples/run_demo.pyThe script creates an isolated .demo-venv, installs the core package, and runs
locally using SQLite. No API key, Docker, account, or database setup is needed.
Installation time depends on your connection; subsequent runs reuse the environment.
make demo and sh examples/run_demo.sh are alternatives on Unix systems.
The demo reflects a fictional patients → orders → lab_results schema, classifies
its columns, creates 50 patients with 3–5 orders each, flags exactly 10% of results
(rounded to the nearest row), inserts everything into SQLite, and verifies both
foreign keys and deterministic replay. It prints the counts, dataset hash, and
coverage suggestions. Each run retains JSON data, its plan/genome, a manifest,
a SQLite database, and sample HL7 messages under .dataloom/demo/run-*.
Related MCP server: Alma Atlas
Install and use
python -m venv .venv
# Unix: source .venv/bin/activate
# PowerShell: .venv\Scripts\Activate.ps1
python -m pip install -e .
dataloom introspect --ddl examples/demo_schema.sql
dataloom classify
dataloom explain
dataloom generate --auto --rows 25 --output out/first
dataloom generate --plan examples/plan.yaml --output out/baseline
dataloom suggest-gapsOptional extras: .[mcp], .[anthropic], .[postgres], .[parquet], or .[all].
For development, install .[dev]. python -m dataloom works wherever the dataloom
executable is not on PATH.
--auto derives a plan from the schema alone: every entity is planned, tables with
no outgoing foreign key receive --rows records, and every child fans out from its
identifying parent. The result is written to .dataloom/authored-plan.json, so you
can read it, edit it, and replay it with --plan. No plan file and no API key are
involved. Schemas whose types, computed columns, or self-references fall outside
Phase 1 are reported in full before anything is written.
Classification recognizes common column names exactly, by token span
(customer_email, billing_city, internal_notes), and in camelCase, then maps
them onto the core generators: personal and full names, companies, job titles,
street addresses, cities, US states, postcodes, countries, emails, usernames, URLs,
phone numbers, IPv4 addresses, currency codes, and sentence-length descriptions.
Identifier-shaped values stay inside ranges reserved for documentation and testing:
example.test, 202-555-01xx, 198.51.100.0/24, and the never-issued 000 SSN
area number. Generic words are deliberately left alone, so product_name and
order_state stay unlabelled rather than borrowing a person or place meaning.
Unlabelled text receives obviously synthetic test-... values; give those columns an
explicit rule. Values longer than a declared length are truncated to fit.
Output directories must be new. Choose another path to run again; DataLoom does not replace existing datasets. A bundle contains:
manifest.json: row counts, hashes, versions, and table-to-file mapping.genome.jsonandplan.json: the exact inputs for replay.One data file per table. Table names are UTF-8 hex-encoded to avoid path and filename collisions; use the manifest to find them.
Replay a bundle:
dataloom generate --genome out/baseline/genome.json \
--plan out/baseline/plan.json --output out/replayCompare receipt.data_hash in the two manifests. See
the reproducibility contract.
Plans are the product
Business fields can now depend on each other: calculate totals from quantity and price, choose discounts or flags from other values, and generate shipping or resolution dates after creation. Dependencies are resolved automatically; unknown references and cycles fail before row generation. All rules stay deterministic and work through the existing CLI and MCP tools.
Try the complete business-consistency demo:
python examples/run_demo.py --business-rulesIt generates 100 fictional fulfillments and independently checks their arithmetic, discount policy, and shipping dates before exporting a replayable bundle. See the plan and derived-field reference.
format_version: 1
seed: 42
reference_date: '2025-01-01'
entities:
patients:
rows: 50
orders:
fanout:
parent: patients
foreign_key: [patient_id]
minimum: 3
maximum: 5
rules:
priority:
choices: [routine, urgent]Include every required parent. Each entity specifies either rows or fanout.
Rules support semantic generators, choices, numeric bounds, null quotas, and
exact value proportions. See the complete example and
plan reference. Unknown fields and conflicting rules fail validation.
To generate without an existing schema, use a domain template:
dataloom list-domain-packs
# Supply a plan with a patients entity and row count:
dataloom generate --template healthcare:patients --plan patients.yaml --output out/patientsFor natural language, install .[anthropic], set ANTHROPIC_API_KEY and
DATALOOM_MODEL to a model available in your account, then run:
dataloom generate --request "50 patients, 3 to 5 orders each, 10 percent abnormal lab results" \
--output out/scenarioThis authors and saves .dataloom/authored-plan.json, then executes it. LLM-authored
plans receive the same validation as hand-written plans. Providers only receive
schema metadata; profiling values are not included in the prompts. For a review
step before execution, a client can author a plan through the library's
author_plan function, inspect/save it, then invoke offline generation.
Live databases and exports
Put the connection URL in an environment variable, for example SOURCE_DATABASE_URL
containing postgresql+psycopg://user:password@localhost/database.
dataloom introspect --database-env SOURCE_DATABASE_URL --profile-rows 1000
dataloom classify
dataloom generate --plan examples/plan.yaml --format parquet --output out/parquet
dataloom generate --plan examples/plan.yaml --format csv --output out/csv
dataloom generate --plan examples/plan.yaml --format postgres --target-database-env TARGET_DATABASE_URLProfiling is optional and bounded per table. It records sample cardinality, null rates, extrema, common values, numeric quantiles, and consistent string shapes. Numeric generation can interpolate observed quantiles; null quotas can follow observed rates. It does not learn cross-column correlations or copy observed string values into generated rows. Profiled genomes can contain real sample values: review them before committing or sharing.
Database output inserts into existing tables in one transaction. It neither creates schemas nor clears existing rows. Key conflicts cause rollback. Target sequences are not advanced to match explicitly inserted synthetic IDs; use a dedicated test database and manage subsequent application inserts accordingly.
MCP
Install .[mcp]. DataLoom uses the official Python SDK's FastMCP v1 API with a
<2 dependency bound. Configure a stdio server in Claude Desktop using absolute
paths (on Windows, JSON paths require doubled backslashes):
{
"mcpServers": {
"dataloom": {
"command": "/absolute/path/to/dataloom/.venv/bin/python",
"args": ["-m", "dataloom.server"],
"env": {"DATALOOM_WORKSPACE": "/absolute/path/to/dataloom"}
}
}
}The same stdio command can be registered with Claude Code or another MCP client.
The server exposes introspect_schema, classify_columns, generate_dataset,
explain_schema_genome, list_domain_packs, and suggest_gaps. Every tool accepts
a typed args object and returns structured output. genome://current exposes
the default .dataloom/genome.json artifact.
Artifact paths are constrained to DATALOOM_WORKSPACE, including resolved symlinks.
Installed domain entry points execute trusted Python code. The server is local
stdio only; do not treat it as a hosted multi-user service.
Architecture
flowchart LR
DDL[DDL file] --> Genome[Versioned schema genome]
DB[SQLAlchemy reflection] --> Genome
Profile[Optional bounded profiling] --> Genome
Classify[Heuristics + optional provider] --> Genome
Request[Natural language] --> Provider[Provider interface]
Provider --> Plan[Versioned plan]
YAML[Hand-authored YAML / JSON] --> Plan
Genome --> Engine[Seeded relational engine]
Plan --> Engine
Packs[Domain generators / templates] --> Engine
Engine --> Validate[Constraint validation + receipt]
Validate --> Outputs[CSV / JSON / Parquet / PostgreSQL]
Validate --> Coverage[Local coverage observations]service.py coordinates these components. CLI and MCP are thin adapters. Domain
packs register semantic generators and entity templates; output connectors consume
validated datasets. See CONTRIBUTING for extension examples.
Scope and checks
Phase 1 supports scalar relational schemas, composite foreign keys, acyclic parent
ordering, PK/unique constraints, and a documented SQL CHECK subset. Unsupported
types or checks fail explicitly. Cycles, overlapping foreign keys, computed columns,
advanced PostgreSQL types, arbitrary SQL expressions, masking, and subsetting are
outside the generation scope. DDL import accepts CREATE TABLE, not migration
scripts. See support boundaries.
Healthcare includes CMS-checksum-valid NPIs, small ICD-10-CM and LOINC identifier subsets, and minimal HL7 v2.5 ADT/ORU builders. These are test primitives, not a clinical simulation or registry of real providers. Standards attribution and terms are in THIRD_PARTY_NOTICES.
python -m pytest
python -m ruff check src tests examples
python -m ruff format --check src tests examples
python -m mypy src/dataloomCI runs these checks on Python 3.11–3.13 on Windows and Linux, plus a PostgreSQL 16
integration job. Local PostgreSQL tests require DATALOOM_TEST_POSTGRES pointing
to a disposable database with permission to create/drop test schemas. The standard
suite needs no live LLM. Milestone evidence.
MIT licensed. Public, fictional examples only. Contributions from other domains are welcome through the documented plugin interfaces.
Project resources
Contributing and changelog.
Security policy: report vulnerabilities privately.
The development version is installed from this repository. These instructions do not assume that a DataLoom package has been published to PyPI.
This server cannot be deployed
Maintenance
Related MCP Connectors
- SchemaOAuthai.schemalabs
The AI that understands raw data: Schema over your tables and databases, as MCP tools.
Governed data discovery, exact queries, decisions, simulations, and runtime utilities over MCP.
Data-ontology maps of your business systems, served to AI agents over MCP.
Hosted MCP server for AI-driven data ops. Create apps, manage schemas, and CRUD structured data.
Related MCP Servers
FlicenseNot gradedqualityAmaintenanceVisual no-code generator that turns any database into multiple scoped MCP servers — one per access group, with PII masking and fail-closed query scoping built in.05-- AlicenseNot gradedqualityCmaintenanceEnables AI agents to query live schema, lineage, and query-context across data warehouses, dbt projects, orchestration systems, and BI tools via MCP tools.Apache 2.0
- AlicenseNot gradedqualityBmaintenanceProvides deterministic tools for understanding, transforming, and verifying structured data via MCP, enabling rule inference from examples and verification of transformed records.MIT
- AlicenseNot gradedqualityCmaintenanceEnterprise-grade MCP server for Microsoft SQL Server, enabling semantic schema discovery, table profiling, safe data operations with preview/confirm, and multi-environment support.33MIT