Skip to main content
Glama
README.md
# LeakageLens MCP

Dataset forensics for AI coding agents. LeakageLens catches target leakage, cross-split
entities, time leakage, PII, identifier features, class imbalance, and preprocessing-before-
split errors **before** an agent trains a misleading model.

It is not a “chat with CSV” server. The dataset remains local; the MCP client receives bounded,
structured evidence and an explicit `pass`, `review`, or `block` verdict.

## Why MCP?

Any MCP-compatible coding agent can discover the same audit tools and guidance without a custom
integration. Resources provide experiment policy, prompts enforce a review workflow, and tools
return typed evidence. The core engine works without an LLM; an optional Gemini check reasons
about whether features exist at prediction time.

## Quick start

```bash
python -m venv .venv
source .venv/bin/activate
pip install -e '.[dev]'

leakagelens examples/leaky_churn.csv \
  --target churned --split split --entity customer_id --time-column signup_date

pytest -q
python evals/benchmark.py
```

Expected demo verdict: `block`, with direct-target, entity-overlap, PII, identifier, and semantic
leakage findings.

## Connect an MCP client

Example local configuration:

```json
{
  "mcpServers": {
    "leakagelens": {
      "command": "/absolute/path/to/.venv/bin/leakagelens-mcp",
      "env": {
        "LEAKAGELENS_DATA_ROOT": "/absolute/path/to/leakagelens-mcp"
      }
    }
  }
}
```

Then ask your agent:

> Review `examples/leaky_churn.csv` before training a churn model. The target is `churned`, the
> entity is `customer_id`, `signup_date` is the time column, and `split` defines train/test.

### Remote transport

```bash
MCP_TRANSPORT=streamable-http LEAKAGELENS_DATA_ROOT="$PWD" leakagelens-mcp
```

Or run the container with a read-only data mount:

```bash
docker build -t leakagelens-mcp .
docker run --rm -p 8000:8000 -v "$PWD/examples:/data:ro" leakagelens-mcp
```

## MCP surface

| Primitive | Name | Purpose |
|---|---|---|
| Tool | `profile_dataset` | Bounded schema, missingness, uniqueness, and samples |
| Tool | `audit_dataset` | Leakage, PII, split, identifier, and metric checks |
| Tool | `audit_training_code` | AST audit of split/preprocessing order |
| Tool | `review_feature_availability` | Optional Gemini semantic review |
| Resource | `guidance://experiment-contract` | Minimum valid experiment contract |
| Prompt | `review_before_training` | Reusable pre-training workflow |

All file tools are restricted to `LEAKAGELENS_DATA_ROOT` to avoid arbitrary host-file access.

## Detection design

- **Direct leakage:** equality, near-perfect numeric association, deterministic categorical maps
- **Semantic leakage:** outcome-like and target-derived column names
- **Split leakage:** entity overlap and invalid chronological boundaries
- **Privacy:** PII column names and value-pattern scans
- **Pipeline leakage:** Python AST detects fitting or resampling before the split
- **Evaluation risk:** imbalance-aware metric recommendations

Every finding includes a stable code, severity, evidence, affected columns, and remediation. The
risk score is deterministic: critical `30`, high `15`, medium `7`, low `3`, capped at `100`.

## Optional Gemini review

```bash
pip install -e '.[gemini]'
export GEMINI_API_KEY='...'
export GEMINI_MODEL='gemini-2.5-flash'
```

Only column names/types and the user-provided prediction moment are sent. Raw rows are not sent.
Security decisions never depend solely on the model.

## Benchmark

`evals/benchmark.py` generates 35 reproducible scenarios: 25 seeded leakage/privacy/evaluation
cases and 10 clean negative controls. It reports scenario recall and false-block rate. Add cases
before adding heuristics; this prevents a growing collection of unmeasured rules.

Current synthetic benchmark result: **35/35 scenarios passed, 100% scenario recall, 0% false-block
rate**. These figures validate the included seeded cases; they are not estimates of performance on
arbitrary real-world datasets.

## Current boundaries

- Statistical association is a warning, not proof of leakage.
- Semantic rules cannot know feature availability without a prediction-time contract.
- The AST audit recognizes common scikit-learn patterns, not arbitrary dynamic Python.
- This release audits supplied splits; it does not mutate the user’s dataset.

## Architecture

```text
MCP client → FastMCP tools → path boundary → audit engine → typed findings
                                      ├── dataframe checks
                                      ├── split validation
                                      ├── Python AST audit
                                      └── optional Gemini review
```

## Development

```bash
ruff check .
pytest -q
python evals/benchmark.py
```

See [`HELPER_GUIDE.md`](HELPER_GUIDE.md) for a concise code walkthrough.

TDQS

B3.3/5.0

Scored across 4 tools

Disambiguation4/5

Each tool targets a distinct stage of leakage detection: profiling, dataset auditing, code auditing, and feature availability review. However, profile_dataset and audit_dataset could be confused by an agent since both operate on datasets, though their descriptions help separate general profiling from leakage-specific auditing.

Naming Consistency5/5

All tool names consistently follow a verb_noun pattern using lowercase snake_case: profile_dataset, audit_dataset, audit_training_code, and review_feature_availability. This makes the tool set predictable and easy to navigate.

Tool Count5/5

Four tools is well-scoped for a specialized leakage-detection server. Each tool has a clear, non-redundant role, and the count feels appropriate rather than thin or bloated.

Completeness4/5

The tool surface covers the core leakage-detection workflow: profile data, audit data for leakage, audit training code, and verify feature availability. Minor gaps exist around remediation or actionable reporting after an audit, but the main detection loop is complete.

Maintenance

ActivitySlowing
ResponsivenessNo issues