data-governance-mcp
Generates production-ready dbt schema.yml files with tests (not_null, unique, value tests) from audited datasets, enabling automated data quality checks in dbt projects.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@data-governance-mcpsanitize PII in customer_data.csv and generate a synthetic twin"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
The Vibe Coding Problem vs The Solution
When developers feed raw datasets into AI coding assistants (Cursor, Claude Desktop, Antigravity, Windsurf), three critical bottlenecks arise:
Dimension | Raw Prompting (The Naive Way) | With | Enterprise Impact |
Privacy & GDPR | Uploads customer PII, internal IDs, and secrets into cloud LLMs. | Synthesizes Differential-Privacy Digital Twins with zero real records. | 100% Compliance / Zero Leakage |
Context Consumption | 12,000โ50,000 tokens per prompt turn ($0.15โ$0.60 / message). | Condenses datasets into statistical fingerprints of ~680 tokens. | -94.5% Token Expenditure |
Model Attention | Context bloat triggers hallucinations and degrades reasoning. | High-density 5-number moments and edge-case exemplars. | Higher LLM Code Precision |
Quality Baseline | Blind scripting without knowing null rates, outliers, or skewness. | Certified DAMA-DMBOK 6-dimension scorecards (0โ100 index). | Predictable Reliability |
Production Path | Ephemeral, throwaway Python scripts discarded after prompting. | Auto-generates production | Instant CI/CD Readiness |
Related MCP server: DataHub MCP Server
Live AI Assistant Showcase
Here is how an AI assistant running data-governance-mcp handles proprietary enterprise data inside Cursor or Claude Desktop:
> User in Cursor:
"@data-governance I need to write a predictive anomaly detection pipeline for furnace_telemetry.csv,
but corporate compliance prohibits uploading internal telemetry and customer data to external LLMs."
> Claude (via Data Governance MCP):
"I have audited `furnace_telemetry.csv` and secured your development session:
1. ๐จ PII Sanitization: Flagged unmasked email addresses and API keys in column `operator_email`.
2. ๐งช Synthetic Digital Twin: Generated `synthetic_twin.csv` (500 rows) matching exact statistical
moments (ฮผ=1,245.8ยฐC, ฯ=42.1ยฐC), categorical distributions, and null ratios with differential privacy (ฮต=1.0).
3. โก Token Optimization: Compressed 14,250 tokens down to 680 tokens (-95.2% prompt reduction).
4. ๐ Production Artifacts:
- Generated `schema.yml` with dbt tests (`not_null`, `unique`, and Tukey outlier range tests).
- Exported interactive executive audit report to `reports/audit_dashboard.html`.
You can now develop and test your predictive model in Cursor using the synthetic twin with zero compliance risk."System Architecture
The architecture comprises four decoupled operational stages and four concrete enterprise deliverables:
(A) Raw Data Ingestion & PII Sanitization
Schema Sniffer: Delimiter and format auto-detection (CSV/TSV, Parquet/Arrow, SQLite, JSON, memory streams), strict type inference, encoding detection, and header validation.
PII Sanitization: Pattern-based regex & heuristic interception for emails, phone numbers, tax IDs (DNI/NIE), credit cards, and API secrets/tokens before data enters the LLM prompt.
(B) Synthetic Digital Twin Generation
Statistical Moment Profiler: Computes empirical moments ($\mu$, $\sigma$, min, max, skewness, kurtosis), categorical frequency distributions, and correlation structures.
Differential Privacy Engine: Injects calibrated Laplacian/Gaussian noise $(\varepsilon, \delta)$ to generate statistically faithful mock datasets with identical column types and null dynamics without exposing a single real row.
(C) DAMA-DMBOK Quality Audit & Token Optimization
DAMA-DMBOK 6 Dimensions: Rigorous audit of Completeness, Uniqueness, Validity, Consistency, Timeliness, and Accuracy.
Tukey IQR Accuracy Engine: Evaluates outlier fences ($2.5 \times \text{IQR}$) and Z-score distributions across numeric domains.
Token Compressor: Condenses tabular datasets into dense statistical fingerprints, achieving 90% to 95% prompt token reduction without loss of schema semantics.
(D) Model Context Protocol (MCP) & AI Integration
FastMCP Server: Standardized JSON-RPC stdio protocol exposing 8 audit and transformation tools directly into Cursor IDE, Claude Desktop, Antigravity, and Cline.
Enterprise Deliverables:
Synthetic dataset (privacy-safe): Zero-leakage drop-in replacement for code generation and test execution.
Quality audit report (DAMA): Multi-dimensional scorecards with radar diagrams and prioritized remediation steps.
dbt schema (auto-generated): Production-ready
schema.ymlwithnot_null,unique, and value tests.Optimized context for LLMs: High-density token representations (e.g. 12,450 tokens $\rightarrow$ 680 tokens, -94.5% compression).
Interactive HTML Dashboard
Generate zero-dependency, self-contained executive audit reports containing interactive SVG radar charts and dimension health gauges:
data-governance-mcp dashboard telemetry.csv --out audit_report.htmlTerminal CLI Experience
When working directly in your terminal, data-governance-mcp delivers a rich, color-coded diagnostic dashboard powered by rich:
data-governance-mcp audit examples/sample_datasets/industrial_furnace_telemetry.csvโโโโโโโโโโโโโโโโโโโโ DATA GOVERNANCE & PRIVACY TWIN AUDIT โโโโโโโโโโโโโโโโโโโโโ
โ Dataset: Delimited file (industrial_furnace_telemetry.csv) โ
โ Records: 13 | DAMA-DMBOK Score: 94.6 / 100 (EXCELLENT) โ
โ Security & PII Risk: HIGH (2 findings flagged) โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
DAMA-DMBOK 6 Core Quality Dimensions
โโโโโโโโโโโโโโโโฌโโโโโโโโโฌโโโโโโโโโฌโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ Dimension โ Weight โ Score โ Status โ Diagnostics โ
โโโโโโโโโโโโโโโโผโโโโโโโโโผโโโโโโโโโผโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ Completeness โ 22% โ 97.8% โ [ PASS ] โ 2 null values detected (2.2% โ
โ โ โ โ โ missingness). โ
โ Uniqueness โ 18% โ 84.6% โ [ FAIL ] โ 1 exact duplicate row found. โ
โ Validity โ 22% โ 100.0% โ [ PASS ] โ All columns conform to types.โ
โ Accuracy โ 16% โ 89.5% โ [ WARNING ] โ 1 statistical outlier (IQR). โ
โ Consistency โ 12% โ 100.0% โ [ PASS ] โ No logical contradictions. โ
โ Timeliness โ 10% โ 95.0% โ [ PASS ] โ Evaluated on 'timestamp'. โ
โโโโโโโโโโโโโโโโดโโโโโโโโโดโโโโโโโโโดโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ Action Plan โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ Recommended Remediation Workflow: โ
โ 1. Generate privacy twin: data-governance-mcp twin telemetry.csv โ
โ 2. Export dbt tests: data-governance-mcp dbt telemetry.csv โ
โ 3. Generate HTML report: data-governance-mcp dashboard telemetry.csv โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโDAMA-DMBOK Quality Dimensions Matrix
Dimension | Weight | Detection Criteria | Remediation Strategy |
Completeness | 22% | Missing values ( | Targeted imputation or automated filtering. |
Uniqueness | 18% | Duplicate rows and natural primary key candidate viability. | Deduplication rules and surrogate key creation. |
Validity | 22% | Type conformance and mixed non-numeric values in numeric columns. | Robust schema casting and string sanitization. |
Accuracy | 16% | Tukey IQR fences (2.5x) and Z-score outlier detection. | Domain boundary enforcement and telemetry capping. |
Consistency | 12% | Cross-column logic (e.g. chronology inversions: | Relational sanity checks and constraint rules. |
Timeliness | 10% | Temporal freshness, date parsing validation, and cadence continuity. | ISO 8601 formatting and drift tracking. |
Quickstart
Installation
# Using uv (Recommended)
uv tool install data-governance-mcp
# Or via standard pip
pip install data-governance-mcpConfigure in Claude Desktop
Add to your claude_desktop_config.json:
{
"mcpServers": {
"data-governance": {
"command": "python",
"args": ["-m", "data_governance_mcp.server"]
}
}
}Configure in Cursor IDE / Antigravity / Cline
Add via standard stdio transport under Features > MCP Servers:
Name:
data-governanceType:
commandCommand:
python -m data_governance_mcp.server
MCP Tools Reference
Tool Name | Parameters | Return Format | Purpose |
|
| Markdown | Full DAMA scorecard, PII risk level, and prioritized remediation plan. |
|
| JSON | Anonymized statistical twin for safe vibe coding without data leakage. |
|
| JSON | High-density statistical summary saving 90-95% of prompt tokens. |
|
| YAML | Ready-to-commit |
|
| String (Path) | Standalone interactive HTML report with offline compatibility. |
|
| JSON | Explicit pattern matching for emails, cards, phones, and API secrets. |
|
| JSON | Granular dimension scores (0-100) and analytical diagnostics. |
|
| Markdown | Concrete Python and SQL scripts tailored to repair detected issues. |
Verification & Testing
Every release is verified across Python 3.10, 3.11, and 3.12:
pytest -vtests/test_dimensions.py::test_completeness_perfect PASSED [ 5%]
tests/test_dimensions.py::test_completeness_with_nulls_and_blanks PASSED [ 10%]
tests/test_dimensions.py::test_uniqueness_with_duplicates PASSED [ 15%]
tests/test_dimensions.py::test_accuracy_outliers PASSED [ 21%]
tests/test_dimensions.py::test_consistency_chronology_inversion PASSED [ 26%]
tests/test_dimensions.py::test_evaluate_all_summary PASSED [ 31%]
tests/test_loader.py::test_load_inline_csv PASSED [ 36%]
tests/test_loader.py::test_load_sqlite PASSED [ 42%]
tests/test_pii_detector.py::test_pii_detection_clean PASSED [ 47%]
tests/test_pii_detector.py::test_pii_detection_email_and_secrets PASSED [ 52%]
tests/test_server.py::test_audit_dataset_tool PASSED [ 57%]
tests/test_server.py::test_profile_schema_tool PASSED [ 63%]
tests/test_server.py::test_detect_pii_tool PASSED [ 68%]
tests/test_server.py::test_evaluate_dama_dimensions_tool PASSED [ 73%]
tests/test_server.py::test_suggest_remediations_tool PASSED [ 78%]
tests/test_wow_features.py::test_synthetic_twin_generator PASSED [ 84%]
tests/test_wow_features.py::test_token_compressor PASSED [ 89%]
tests/test_wow_features.py::test_dbt_exporter PASSED [ 94%]
tests/test_wow_features.py::test_dashboard_exporter PASSED [100%]
============================= 19 passed in 2.19s ==============================Author & Governance Credentials
Developed by Ioseba Alonso
Certified Data Management Professional (CDMP) by DAMA International & Industrial AI Practitioner.
This server cannot be deployed
Maintenance
Related MCP Connectors
Governance copilot for AI-assisted coding. 72 packs, 532 rules, proof bundles.
The grounded data layer for any LLM: governed SQL, metrics, lineage and catalog over your data.
Paid deterministic data-quality and execution-verification tools for AI agents.
Classify data safety before storing or sharing. GDPR, HIPAA, PCI-DSS, CCPA. AI-powered.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to understand and query your database safely by providing a semantic layer of metadata, with tools to search, explain, validate, and generate safe SQL.2MIT
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to search, explore data lineage, understand business context, and generate SQL queries across an organization's data ecosystem.Apache 2.0
- AlicenseNot gradedqualityFmaintenanceEnables data engineers and BI professionals to perform data pipeline development, quality assurance, visualization, and API integration locally with AI assistance.5MIT
- AlicenseAqualityAmaintenanceLets AI assistants work on sensitive files by reasoning over schemas and masked output while deterministic local code handles raw data, never exposing real values to the model.547 PyPIMIT