Skip to main content
Glama
README.md
# ๐Ÿš€ mcpkit-data

> **The ultimate MCP server for data engineers who want to stop context-switching and start shipping** โœจ

A comprehensive [Model Context Protocol](https://modelcontextprotocol.io) server that gives your AI assistant superpowers to work with Kafka, databases, AWS, data pipelines, and more. Stop jumping between terminals, dashboards, and docs. Just ask your AI to do it.

[![Python 3.10+](https://img.shields.io/badge/python-3.10+-blue.svg)](https://www.python.org/downloads/)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)

---

## โš ๏ธ Early Days Disclaimer

**Hey there! ๐Ÿ‘‹** Before you dive in, here's the deal:

This is a **brand new project** fresh out of the oven ๐Ÿž. We're talking:
- ๐Ÿ†• **First release** - Like, literally just born
- ๐Ÿงช **Not production-ready** - More like "works on my machine" ready
- ๐Ÿ‘ฅ **Limited testing** - It's mostly been tested by... well, me
- ๐Ÿ› ๏ธ **Actively evolving** - Things might break, change, or surprise you

**What this means:**
- โœ… Great for experimentation and learning
- โœ… Perfect for side projects and personal use
- โœ… Awesome if you want to contribute and shape the future
- โŒ Probably not ideal for critical production systems (yet!)
- โŒ Might have bugs, quirks, or "interesting" behaviors
- โŒ Documentation might be incomplete or confusing

**But here's the cool part:** You can help make it better! ๐Ÿš€ Found a bug? Report it! Want a feature? Ask for it! Want to contribute? We'd love that!

*TL;DR: This is the "early access" version. Use at your own risk, have fun, and help us make it awesome!* ๐ŸŽ‰

---

## ๐ŸŽฏ What is this?

**mcpkit-data** is your AI assistant's Swiss Army knife for data engineering. It exposes **66+ tools** through the MCP protocol, letting your AI:

- ๐Ÿ” **Debug Kafka pipelines** without leaving your editor
- ๐Ÿ“Š **Query databases** (JDBC, Athena) and analyze results
- ๐Ÿผ **Process data** with pandas, polars, and DuckDB
- โ˜๏ธ **Work with AWS** (Athena, S3, Glue) seamlessly
- ๐Ÿ—๏ธ **Inspect infrastructure** (Nomad, Consul) in real-time
- ๐Ÿ“ˆ **Generate charts** and evidence bundles automatically
- โœ… **Validate data quality** with Great Expectations

**All without you writing a single line of code.** Just chat with your AI. ๐ŸŽ‰

---

## โšก Quick Start

### Installation

```bash
pip install -e ".[dev]"
```

### Run the Server

```bash
python -m mcpkit.server
```

### Configure Cursor (or your MCP client)

Add to your Cursor MCP settings:

```json
{
  "mcpServers": {
    "mcpkit-data": {
      "command": "python",
      "args": ["-m", "mcpkit.server"],
      "env": {
        "MCPKIT_KAFKA_BOOTSTRAP": "localhost:9092",
        "MCPKIT_SCHEMA_REGISTRY_URL": "http://localhost:8081",
        "MCPKIT_JDBC_DRIVER_CLASS": "org.postgresql.Driver",
        "MCPKIT_JDBC_URL": "jdbc:postgresql://localhost/test",
        "MCPKIT_JDBC_JARS": "/path/to/postgresql.jar",
        "MCPKIT_ROOTS": "/allowed/path1:/allowed/path2"
      }
    }
  }
}
```

**That's it!** Your AI can now use all 66+ tools. ๐ŸŽŠ

---

## ๐Ÿ› ๏ธ Tools Overview

### ๐Ÿ“ฆ Dataset Registry (4 tools)

Your data's home base. Store, query, and manage datasets as Parquet files.

| Tool | What it does | When to use |
|------|-------------|-------------|
| `dataset_list` | ๐Ÿ“‹ List all datasets | See what data you have |
| `dataset_info` | โ„น๏ธ Get dataset metadata | Check schema, row count |
| `dataset_put_rows` | ๐Ÿ’พ Store rows as Parquet | Save query results |
| `dataset_delete` | ๐Ÿ—‘๏ธ Delete a dataset | Clean up old data |

---

### ๐ŸŽง Kafka Tools (8 tools)

Debug, consume, and analyze Kafka topics like a pro.

| Tool | What it does | When to use |
|------|-------------|-------------|
| `kafka_list_topics` | ๐Ÿ“ List all topics | Discover what's available |
| `kafka_offsets` | ๐Ÿ“Š Get partition offsets | Check lag, find positions |
| `kafka_consume_batch` | ๐Ÿ“ฅ Consume messages โ†’ dataset | Extract data for analysis |
| `kafka_consume_tail` | ๐Ÿ”š Get last N messages | Debug recent events |
| `kafka_filter` | ๐Ÿ” Filter with JMESPath | Find specific records |
| `kafka_flatten` | ๐Ÿ“ Flatten nested records | Normalize JSON structures |
| `kafka_groupby_key` | ๐Ÿ”‘ Group by extracted key | Aggregate by message key |
| `kafka_describe_topic` | ๐Ÿ“‹ Get topic config/partitions | Inspect topic metadata |

---

### ๐Ÿ“‹ Schema Registry & Decoding (4 tools)

Work with Avro, Protobuf, and schema discovery.

| Tool | What it does | When to use |
|------|-------------|-------------|
| `schema_registry_get` | ๐Ÿ“„ Get schema by ID/subject | Fetch schema definitions |
| `schema_registry_list_subjects` | ๐Ÿ“š List all subjects | Discover available schemas |
| `avro_decode` | ๐Ÿ”“ Decode Avro messages | Parse binary Avro data |
| `protobuf_decode` | ๐Ÿ”“ Decode Protobuf messages | Parse binary Protobuf data |

---

### ๐Ÿ—„๏ธ Database Tools (2 tools)

Query databases safely with read-only enforcement.

| Tool | What it does | When to use |
|------|-------------|-------------|
| `db_query_ro` | ๐Ÿ” Execute read-only SQL | Query any database |
| `db_introspect` | ๐Ÿ”Ž Introspect schema | Discover tables/columns |

**Safety**: Automatically blocks `DROP`, `INSERT`, `UPDATE`, `DELETE`, etc. โœ…

---

### โ˜๏ธ AWS Tools (9 tools)

Query Athena, list S3, manage Glue partitions.

| Tool | What it does | When to use |
|------|-------------|-------------|
| `athena_start_query` | ๐Ÿš€ Start SQL query | Run analytics queries |
| `athena_poll_query` | โณ Check query status | Monitor execution |
| `athena_get_results` | ๐Ÿ“Š Get query results | Fetch data |
| `athena_explain` | ๐Ÿ“– Explain query plan | Optimize queries |
| `athena_partitions_list` | ๐Ÿ“‚ List partitions | Check table structure |
| `athena_repair_table` | ๐Ÿ”ง Repair table (MSCK) | Fix partition metadata |
| `athena_ctas_export` | ๐Ÿ“ค CTAS export | Export to S3 |
| `s3_list_prefix` | ๐Ÿ“ List S3 objects | Browse buckets |

**Safety**: SQL validation blocks dangerous operations by default. โœ…

---

### ๐Ÿผ Pandas Tools (12 tools)

The data scientist's best friend, now in your AI assistant.

| Tool | What it does | When to use |
|------|-------------|-------------|
| `pandas_from_rows` | ๐Ÿ“Š Create DataFrame | Build datasets |
| `pandas_describe` | ๐Ÿ“ˆ Get statistics | Understand your data |
| `pandas_groupby` | ๐Ÿ”ข Group + aggregate | Summarize by categories |
| `pandas_join` | ๐Ÿ”— Join datasets | Combine data sources |
| `pandas_filter_query` | ๐Ÿ” Filter rows | Subset your data |
| `pandas_filter_time_range` | โฐ Filter by time | Time-series analysis |
| `pandas_diff_frames` | ๐Ÿ”„ Compare datasets | Find differences |
| `pandas_schema_check` | โœ… Validate schema | Ensure data quality |
| `pandas_sample_stratified` | ๐ŸŽฒ Stratified sample | Balanced sampling |
| `pandas_sample_random` | ๐ŸŽฒ Random sample | Quick data preview |
| `pandas_count_distinct` | ๐Ÿ”ข Count unique values | Cardinality analysis |
| `pandas_export` | ๐Ÿ’พ Export to CSV/Parquet/Excel | Save results |
| `dataset_head_tail` | ๐Ÿ“„ Get first/last N rows | Quick data preview |

---

### โšก Polars Tools (3 tools)

Lightning-fast data processing with Polars.

| Tool | What it does | When to use |
|------|-------------|-------------|
| `polars_from_rows` | ๐Ÿ“Š Create DataFrame | Build datasets |
| `polars_groupby` | ๐Ÿ”ข Group + aggregate | Fast aggregations |
| `polars_export` | ๐Ÿ’พ Export to CSV/Parquet/JSON | Save results |

---

### ๐Ÿฆ† DuckDB Tools (1 tool)

SQL on local data sources. No server needed.

| Tool | What it does | When to use |
|------|-------------|-------------|
| `duckdb_query_local` | ๐Ÿ” SQL on local data | Query Parquet/CSV files |

---

### ๐ŸŽจ JSON & Data Quality (5 tools)

Transform, validate, and correlate events.

| Tool | What it does | When to use |
|------|-------------|-------------|
| `jq_transform` | ๐Ÿ”„ JMESPath transform | Extract/transform JSON |
| `event_validate` | โœ… JSONSchema validation | Validate data structure |
| `event_fingerprint` | ๐Ÿ”‘ Generate fingerprint | Create unique IDs |
| `dedupe_by_id` | ๐Ÿงน Deduplicate records | Remove duplicates |
| `event_correlate` | ๐Ÿ”— Correlate events | Join across batches |

---

### ๐Ÿ“ File & Format Tools (4 tools)

Work with files and convert formats.

| Tool | What it does | When to use |
|------|-------------|-------------|
| `parquet_inspect` | ๐Ÿ” Inspect Parquet schema | Understand file structure |
| `arrow_convert` | ๐Ÿ”„ Convert formats | Parquet โ†” CSV โ†” JSON |
| `fs_read` | ๐Ÿ“– Read file (with offset) | Read large files efficiently |
| `fs_list_dir` | ๐Ÿ“‚ List directory | Browse filesystem |
| `rg_search` | ๐Ÿ”Ž Search with ripgrep | Find patterns in code/data |

---

### ๐Ÿ“Š Visualization & Evidence (2 tools)

Generate charts and evidence bundles automatically.

| Tool | What it does | When to use |
|------|-------------|-------------|
| `dataset_to_chart` | ๐Ÿ“ˆ Auto-generate charts | Visualize data instantly |
| `evidence_bundle_plus` | ๐Ÿ“ฆ Generate evidence bundle | Export stats + samples |

---

### ๐Ÿ—๏ธ Infrastructure Tools (8 tools)

Monitor Nomad jobs and Consul services.

| Tool | What it does | When to use |
|------|-------------|-------------|
| `nomad_list_jobs` | ๐Ÿ“‹ List Nomad jobs (fuzzy search) | Find running services |
| `nomad_get_job_status` | ๐Ÿ“Š Get full job status | Inspect job details |
| `nomad_get_allocation_logs` | ๐Ÿ“œ Read allocation logs | Debug failing jobs |
| `nomad_get_allocation_events` | ๐Ÿ“… Get lifecycle events | Track job history |
| `nomad_get_node_status` | ๐Ÿ–ฅ๏ธ Get node status | Check node health |
| `consul_get_service_ips` | ๐ŸŒ Get service IPs (fuzzy) | Find service endpoints |
| `consul_list_services` | ๐Ÿ“š List all services | Discover service catalog |
| `consul_get_service_health` | โค๏ธ Get service health | Check service status |

**Fuzzy search**: Find services even with partial names! ๐ŸŽฏ

---

### ๐ŸŒ HTTP Tools (1 tool)

Make HTTP requests safely.

| Tool | What it does | When to use |
|------|-------------|-------------|
| `http_request` | ๐ŸŒ HTTP GET/POST | Call APIs, test endpoints |

**Safety**: POST disabled by default. Opt-in required. โœ…

---

### โœ… Great Expectations (1 tool)

Data quality validation with GE.

| Tool | What it does | When to use |
|------|-------------|-------------|
| `great_expectations_check` | โœ… Run GE validation | Validate data quality |

---

### ๐Ÿ”„ Reconciliation (1 tool)

Compare datasets and find differences.

| Tool | What it does | When to use |
|------|-------------|-------------|
| `reconcile_counts` | ๐Ÿ” Reconcile record counts | Compare datasets |

---

## ๐Ÿ”’ Security & Guardrails

### ๐Ÿ›ก๏ธ Built-in Safety

- **Filesystem Allowlist**: All file operations restricted to `MCPKIT_ROOTS`
- **SQL Validation**: JDBC and Athena queries validated (read-only by default)
- **HTTP Safety**: POST requests disabled by default
- **Output Limits**: Automatic caps on rows, records, and bytes
- **Path Traversal Protection**: Blocks `..` and unsafe paths

### ๐Ÿ“ Configurable Limits

| Variable | Default | What it does |
|----------|---------|--------------|
| `MCPKIT_TIMEOUT_SECS` | `15` | Operation timeout |
| `MCPKIT_MAX_OUTPUT_BYTES` | `1000000` | Max output size |
| `MCPKIT_MAX_ROWS` | `500` | Max rows returned |
| `MCPKIT_MAX_RECORDS` | `500` | Max records returned |

---

## โš™๏ธ Configuration

### ๐Ÿ” AWS Credentials

AWS uses the standard boto3 credential chain:

1. **Environment Variables** (recommended):
   ```bash
   export AWS_ACCESS_KEY_ID=your-key
   export AWS_SECRET_ACCESS_KEY=your-secret
   export AWS_REGION=eu-central-1
   ```

2. **AWS Credentials File** (`~/.aws/credentials`):
   ```ini
   [default]
   aws_access_key_id = your-key
   aws_secret_access_key = your-secret
   ```

3. **IAM Roles** (auto-detected on EC2/ECS/Lambda)

4. **AWS SSO** (use `aws sso login` first)

### ๐Ÿ“ Filesystem Security

```bash
export MCPKIT_ROOTS="/allowed/path1:/allowed/path2"
```

### ๐ŸŽง Kafka Configuration

```bash
export MCPKIT_KAFKA_BOOTSTRAP="localhost:9092"
export MCPKIT_KAFKA_SECURITY_PROTOCOL="SASL_SSL"  # optional
export MCPKIT_KAFKA_SASL_MECHANISM="PLAIN"  # optional
export MCPKIT_KAFKA_SASL_USERNAME="user"  # optional
export MCPKIT_KAFKA_SASL_PASSWORD="pass"  # optional
```

### ๐Ÿ“‹ Schema Registry

```bash
export MCPKIT_SCHEMA_REGISTRY_URL="http://localhost:8081"
export MCPKIT_SCHEMA_REGISTRY_BASIC_AUTH="user:pass"  # optional
```

### ๐Ÿ—„๏ธ JDBC

```bash
export MCPKIT_JDBC_DRIVER_CLASS="org.postgresql.Driver"
export MCPKIT_JDBC_URL="jdbc:postgresql://localhost/test"
export MCPKIT_JDBC_JARS="/path/to/postgresql.jar"
export MCPKIT_JDBC_USER="user"  # optional
export MCPKIT_JDBC_PASSWORD="pass"  # optional
```

---

## ๐ŸŽฏ Example Workflows

### ๐Ÿ” End-to-End Kafka Pipeline Debugging

**You:** "Debug the 'order-events' topic: consume recent messages, decode Avro, filter errors, analyze patterns, and create a chart"

**AI Response:**
```
I'll debug the order-events pipeline end-to-end.

[Step 1: kafka_consume_tail] Consuming last 50 messages...
โœ“ Consumed 50 messages from 'order-events'

[Step 2: schema_registry_get] Fetching Avro schema...
โœ“ Schema retrieved: order-event v2.3 (ID: 145)

[Step 3: avro_decode] Decoding Avro messages...
โœ“ Decoded 50 messages successfully

[Step 4: kafka_filter] Filtering error events...
โœ“ Found 8 error events matching criteria

[Step 5: pandas_groupby] Analyzing error patterns...
Error breakdown:
- Payment failures: 4 (50%)
- Validation errors: 2 (25%)
- Timeout errors: 2 (25%)

[Step 6: dataset_to_chart] Creating visualization...
โœ“ Chart saved: .artifacts/order_errors_analysis.png

[Step 7: pandas_export] Exporting error details...
โœ“ CSV exported: .artifacts/order_errors_20251221.csv

Summary: 8 errors out of 50 messages (16% error rate). Payment failures are the primary issue.
```

**Tools used:** `kafka_consume_tail` โ†’ `schema_registry_get` โ†’ `avro_decode` โ†’ `kafka_filter` โ†’ `pandas_groupby` โ†’ `dataset_to_chart` โ†’ `pandas_export` (7 tools)

---

### ๐Ÿ“Š Multi-Source Data Reconciliation

**You:** "Compare sales data between Athena and Postgres: query both, join on order_id, find discrepancies, and export differences"

**AI Response:**
```
I'll reconcile sales data between Athena and Postgres.

[Step 1: athena_start_query] Querying Athena sales table...
โœ“ Query started: execution_id abc123

[Step 2: athena_poll_query] Waiting for completion...
โœ“ Query completed in 3.2s (scanned 1.2 GB)

[Step 3: athena_get_results] Fetching results...
โœ“ Retrieved 5,234 rows from Athena

[Step 4: pandas_from_rows] Storing Athena data...
โœ“ Dataset created: athena_sales_20251221

[Step 5: db_query_ro] Querying Postgres sales table...
โœ“ Retrieved 5,189 rows from Postgres

[Step 6: pandas_from_rows] Storing Postgres data...
โœ“ Dataset created: postgres_sales_20251221

[Step 7: pandas_join] Joining datasets on order_id...
โœ“ Joined datasets: 5,180 matching orders

[Step 8: pandas_diff_frames] Finding discrepancies...
Discrepancies found:
- Missing in Postgres: 54 orders
- Missing in Athena: 9 orders
- Amount mismatches: 12 orders
- Total discrepancy value: $3,456.78

[Step 9: pandas_export] Exporting differences...
โœ“ Differences exported: .artifacts/sales_reconciliation_diff.csv

[Step 10: reconcile_counts] Generating reconciliation report...
โœ“ Report saved: .artifacts/sales_reconciliation_report.json
```

**Tools used:** `athena_start_query` โ†’ `athena_poll_query` โ†’ `athena_get_results` โ†’ `pandas_from_rows` โ†’ `db_query_ro` โ†’ `pandas_from_rows` โ†’ `pandas_join` โ†’ `pandas_diff_frames` โ†’ `pandas_export` โ†’ `reconcile_counts` (10 tools)

---

### ๐Ÿ—๏ธ Service Health Investigation

**You:** "Investigate the 'payment-service': find it in Nomad, check Consul health, get logs, filter errors, correlate with events, and export report"

**AI Response:**
```
I'll investigate the payment-service health comprehensively.

[Step 1: nomad_list_jobs] Finding payment-service...
โœ“ Found: platform-payment-service-app (status: running)

[Step 2: nomad_get_job_status] Getting job details...
โœ“ Job running on node: node-abc-123
โœ“ Allocation: alloc-def-456

[Step 3: consul_get_service_health] Checking Consul health...
โœ“ Service health: passing (3/3 checks passing)
โœ“ Service IPs: 10.0.1.45:8080, 10.0.1.46:8080

[Step 4: nomad_get_allocation_logs] Retrieving logs...
โœ“ Retrieved last 200 log lines

[Step 5: pandas_from_rows] Storing logs as dataset...
โœ“ Dataset created: payment_service_logs

[Step 6: pandas_filter_query] Filtering error logs...
โœ“ Found 15 error entries in last hour

[Step 7: kafka_consume_batch] Consuming related events...
โœ“ Consumed 100 payment events from Kafka

[Step 8: event_correlate] Correlating logs with events...
โœ“ Correlated 12 error logs with payment events

[Step 9: pandas_groupby] Analyzing error patterns...
Error analysis:
- Database timeouts: 8 (53%)
- External API failures: 4 (27%)
- Validation errors: 3 (20%)

[Step 10: evidence_bundle_plus] Generating evidence bundle...
โœ“ Bundle created:
  - .artifacts/payment_service_investigation_stats.json
  - .artifacts/payment_service_error_logs.csv
  - .artifacts/payment_service_correlated_events.csv

Summary: Service is healthy but experiencing intermittent database timeouts affecting 8 transactions.
```

**Tools used:** `nomad_list_jobs` โ†’ `nomad_get_job_status` โ†’ `consul_get_service_health` โ†’ `nomad_get_allocation_logs` โ†’ `pandas_from_rows` โ†’ `pandas_filter_query` โ†’ `kafka_consume_batch` โ†’ `event_correlate` โ†’ `pandas_groupby` โ†’ `evidence_bundle_plus` (10 tools)

---

### โœ… Data Quality Validation Pipeline

**You:** "Validate Kafka events: consume batch, flatten nested JSON, validate against schema, check data quality, and generate report"

**AI Response:**
```
I'll run a complete data quality validation on Kafka events.

[Step 1: kafka_consume_batch] Consuming event batch...
โœ“ Consumed 500 messages from 'user-actions' topic

[Step 2: kafka_flatten] Flattening nested JSON structures...
โœ“ Flattened 500 records into 45 columns
โœ“ Dataset created: user_actions_flattened

[Step 3: schema_registry_get] Fetching validation schema...
โœ“ Retrieved JSONSchema: user-action-event v1.2

[Step 4: pandas_sample_random] Sampling for validation...
โœ“ Sampled 50 random records

[Step 5: event_validate] Validating against schema...
Validation results:
- Valid records: 47 (94%)
- Invalid records: 3 (6%)
  - Missing required field 'timestamp': 2
  - Invalid enum value 'action_type': 1

[Step 6: pandas_schema_check] Checking data types and constraints...
โœ“ Schema check passed: all required columns present
โš ๏ธ Warning: 12% null values in 'metadata' column

[Step 7: pandas_count_distinct] Analyzing data quality...
Distinct value counts:
- user_id: 234 unique (expected: ~250)
- action_type: 8 distinct (expected: 10)
- session_id: 198 unique

[Step 8: pandas_filter_time_range] Filtering recent events...
โœ“ Filtered to last 24 hours: 487 events

[Step 9: great_expectations_check] Running GE validation suite...
โœ“ Expectations passed: 8/10
โš ๏ธ Failed: 
  - expect_column_values_to_be_unique (user_id)
  - expect_column_values_to_not_be_null (metadata)

[Step 10: evidence_bundle_plus] Generating quality report...
โœ“ Quality report: .artifacts/data_quality_report_20251221.zip
```

**Tools used:** `kafka_consume_batch` โ†’ `kafka_flatten` โ†’ `schema_registry_get` โ†’ `pandas_sample_random` โ†’ `event_validate` โ†’ `pandas_schema_check` โ†’ `pandas_count_distinct` โ†’ `pandas_filter_time_range` โ†’ `great_expectations_check` โ†’ `evidence_bundle_plus` (10 tools)

---

### ๐Ÿ“ˆ Time-Series Analysis & Forecasting

**You:** "Analyze sales trends: query Athena for last 6 months, filter by date range, group by week, sample for quality, create chart, and export"

**AI Response:**
```
I'll analyze sales trends over the last 6 months.

[Step 1: athena_start_query] Starting sales query...
โœ“ Query: SELECT * FROM sales WHERE date >= '2025-06-01'

[Step 2: athena_poll_query] Monitoring execution...
โœ“ Query running... (scanned 2.1 GB so far)

[Step 3: athena_get_results] Fetching results...
โœ“ Retrieved 45,678 sales records

[Step 4: pandas_from_rows] Storing in dataset...
โœ“ Dataset created: sales_6months

[Step 5: pandas_filter_time_range] Filtering date range...
โœ“ Filtered: 2025-06-01 to 2025-12-21 (45,234 records)

[Step 6: pandas_groupby] Grouping by week and region...
โœ“ Aggregated: 26 weeks ร— 4 regions = 104 groups

[Step 7: pandas_sample_stratified] Stratified sampling...
โœ“ Sampled 20 records per region for validation

[Step 8: pandas_describe] Statistical analysis...
Statistics:
- Total sales: $12,345,678
- Average per week: $474,833
- Trend: +5.2% week-over-week
- Peak week: Week 24 ($567,890)

[Step 9: dataset_to_chart] Creating time-series chart...
โœ“ Chart: .artifacts/sales_trends_6months.png

[Step 10: pandas_export] Exporting aggregated data...
โœ“ Exported: .artifacts/sales_by_week_region.csv (104 rows)
โœ“ Exported: .artifacts/sales_statistics.json
```

**Tools used:** `athena_start_query` โ†’ `athena_poll_query` โ†’ `athena_get_results` โ†’ `pandas_from_rows` โ†’ `pandas_filter_time_range` โ†’ `pandas_groupby` โ†’ `pandas_sample_stratified` โ†’ `pandas_describe` โ†’ `dataset_to_chart` โ†’ `pandas_export` (10 tools)

---

### ๐Ÿ”„ Cross-System Event Correlation

**You:** "Correlate user events across systems: get Kafka events, query database for user profiles, join datasets, deduplicate, correlate timestamps, and export"

**AI Response:**
```
I'll correlate user events across Kafka and database systems.

[Step 1: kafka_consume_batch] Consuming user events...
โœ“ Consumed 1,000 events from 'user-events' topic

[Step 2: kafka_flatten] Flattening event structure...
โœ“ Flattened to 1,000 records with 32 columns

[Step 3: db_query_ro] Querying user profiles...
โœ“ Retrieved 850 user profiles from database

[Step 4: pandas_from_rows] Storing profiles...
โœ“ Dataset created: user_profiles

[Step 5: pandas_join] Joining events with profiles...
โœ“ Joined on user_id: 987 matched records

[Step 6: dedupe_by_id] Removing duplicate events...
โœ“ Deduplicated: 23 duplicates removed (964 unique)

[Step 7: event_correlate] Correlating by timestamp...
โœ“ Correlated events into 234 user sessions

[Step 8: pandas_groupby] Analyzing session patterns...
Session analysis:
- Average session duration: 12.5 minutes
- Events per session: 4.1
- Most active users: 12 users with 10+ events

[Step 9: pandas_filter_query] Filtering high-value sessions...
โœ“ Found 45 sessions with purchase events

[Step 10: pandas_export] Exporting correlated data...
โœ“ Exported: .artifacts/correlated_user_sessions.csv
โœ“ Exported: .artifacts/session_analysis.json
```

**Tools used:** `kafka_consume_batch` โ†’ `kafka_flatten` โ†’ `db_query_ro` โ†’ `pandas_from_rows` โ†’ `pandas_join` โ†’ `dedupe_by_id` โ†’ `event_correlate` โ†’ `pandas_groupby` โ†’ `pandas_filter_query` โ†’ `pandas_export` (10 tools)

---

## ๐Ÿงช Testing

### Quick Start

```bash
# Run all tests
pytest

# Run unit tests only (no Docker required)
make test-unit

# Run integration tests (requires Docker)
make test-integration

# Run with coverage
make test-coverage
```

### Coverage

```bash
# Run tests with coverage
make test-coverage

# Generate HTML report
make coverage-html
open htmlcov/index.html
```

**Target**: 95% coverage for core modules. See [`tests/COVERAGE.md`](tests/COVERAGE.md) for details.

All tests use Docker containers for external services (Kafka, PostgreSQL, LocalStack, Consul, Nomad). No real AWS credentials needed! โœ…

### ๐Ÿ”„ Reloading MCP Server After Code Changes

When you modify MCP tool code (e.g., `mcpkit/core/*.py` or `mcpkit/server.py`), Cursor needs to reload the MCP server to pick up your changes.

**Quick reload:**
```bash
./reload_mcp_server.sh
```

This script modifies your Cursor MCP config file to trigger an automatic reload. If automatic reload doesn't work, restart Cursor manually.

**When to reload:**
- โœ… After editing tool implementations in `mcpkit/core/*.py`
- โœ… After adding/modifying tools in `mcpkit/server.py`
- โœ… After changing tool documentation or parameters
- โŒ Not needed for config changes (environment variables)

**Manual reload alternative:**
1. Open Cursor settings โ†’ MCP
2. Temporarily disable the `mcpkit-data` server
3. Re-enable it
4. Cursor will reload the server

---

## ๐Ÿค Contributing

We โค๏ธ contributions! Here's how to help:

1. **๐Ÿ› Found a bug?** Open an issue
2. **๐Ÿ’ก Have an idea?** Propose a new tool or feature
3. **๐Ÿ”ง Want to code?** Check out `TOOLS_ANALYSIS.md` for gaps
4. **๐Ÿ“ Docs unclear?** Improve them!

**Design Principles:**
- ๐ŸŽฏ **KISS**: One tool, one job
- ๐Ÿ”— **Composable**: Tools work together seamlessly
- ๐Ÿ”’ **Safe**: Read-only by default, explicit opt-ins
- โšก **Fast**: Efficient operations, smart limits

---

## ๐Ÿ“š Learn More

- [Model Context Protocol](https://modelcontextprotocol.io) - The protocol we use
- [FastMCP](https://github.com/jlowin/fastmcp) - The framework we build on
- [Cursor](https://cursor.sh) - The IDE that makes this shine โœจ

---

## ๐Ÿ“„ License

MIT License - Use it, fork it, make it better! ๐Ÿš€

---

**Made with โค๏ธ for data engineers who want to ship faster**

*Stop context-switching. Start shipping.* ๐ŸŽฏ