DriftScope
# š¬ DriftScope ā Autonomous MLOps & Statistical Diagnostics MCP Server
[](https://python.org)
[](https://modelcontextprotocol.io/)
[](https://driftscope-mcp.onrender.com)
[](https://pola.rs)
> An open-source, cloud-deployed **Model Context Protocol (MCP)** server that equips Large Language Models (LLMs) with deterministic statistical computing engines to audit data drift and model degradation in production tabular pipelines.
---
## š” The Problem
LLMs are exceptional at high-level reasoning, code generation, and root-cause analysis, but notoriously unreliable at **precise statistical math**. When monitoring machine learning pipelines, asking an LLM to evaluate distribution shift directly leads to severe numerical hallucinations.
**DriftScope** solves this by bridging the LLM to a dedicated Python analytical engine via the open **Model Context Protocol**:
- **The LLM** handles orchestration, triage, hypothesis generation, and incident reporting.
- **DriftScope** executes deterministic, vector-accelerated hypothesis testing (**Two-sample Kolmogorov-Smirnov**) and **Population Stability Index (PSI)** calculations using **Polars** and **SciPy**.
---
## šļø Architecture
```
āāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāā
ā LLM Host ā
ā (Claude Desktop / Cursor / Custom Agent) ā
āāāāāāāāāāāāāāāāāāāāāāāāāāāāā¬āāāāāāāāāāāāāāāāāāāāāāāāāāāāā
ā JSON-RPC 2.0
ā¼ (mcp-remote bridge)
āāāāāāāāāāāāāāāāāāāāāāāāāā
ā Internet / HTTPS ā
āāāāāāāāāāāāāā¬āāāāāāāāāāāā
ā Server-Sent Events (SSE)
ā¼
āāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāā
ā DriftScope MCP Server (Render) ā
ā āāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāā ā
ā ā FastMCP Router ā ā
ā āāāāāāāāāā¬āāāāāāāāāāāāāāāāāāāāāāāā¬āāāāāāāāāāāāāāāā¬āā ā
ā ā ā ā ā
ā ā¼ ā¼ ā¼ ā
ā [ Tools Engine ] [ Resources Hub ] [ Prompts ]ā
ā ⢠check_drift ⢠standards:// ⢠audit_ ā
ā ⢠compute_psi drift-policy feature ā
ā ⢠generate_mock ā
ā ā ā
ā ā¼ ā
ā [ Data Layer: Polars + SciPy + NumPy ] ā
āāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāāā
```
---
## ⨠Features & MCP Primitives
### 1. š ļø Tools (Callable Actions)
- `check_feature_drift(baseline_csv, current_csv, feature_column, significance_level)`: Executes two-sample **Kolmogorov-Smirnov (KS)** tests to detect continuous covariate shift with rigorous $p$-values.
- `compute_psi(baseline_csv, current_csv, feature_column, bins)`: Computes the **Population Stability Index (PSI)** with equal-frequency baseline binning and Laplace smoothing to classify shift severity (`Stable`, `Moderate`, or `Critical`).
- `generate_mock_datasets()`: Synthesizes baseline and drifted production samples on the fly for pipeline verification.
*Supports both local cloud file paths and remote public HTTP/HTTPS URLs (AWS S3, GitHub raw, etc.).*
### 2. š Resources (Passive Knowledge)
- `standards://drift-policy`: Exposes organization-wide MLOps threshold standards (e.g., $p < 0.05$ rejection criteria, PSI warning zones) directly into the agent's context.
### 3. š Prompts (Reusable Workflows)
- `audit_feature`: Pre-engineered diagnostic prompt guiding the model through end-to-end drift triage, impact evaluation, and retraining recommendations.
---
## š Live Cloud Deployment
DriftScope is deployed as a live cloud service on **Render** utilizing the **Server-Sent Events (SSE)** transport.
- **Live SSE Endpoint**: `https://driftscope-mcp.onrender.com/sse`
---
## š Quickstart: Connect to Claude Desktop
You can connect your local Claude Desktop to the live cloud deployment in seconds:
1. Open your Claude Desktop configuration file:
- **Windows**: `%APPDATA%\Claude\claude_desktop_config.json`
- **macOS**: `~/Library/Application Support/Claude/claude_desktop_config.json`
2. Add `driftscope-cloud` to your `mcpServers` object:
```json
{
"mcpServers": {
"driftscope-cloud": {
"command": "npx",
"args": [
"-y",
"mcp-remote",
"https://driftscope-mcp.onrender.com/sse"
]
}
}
}
```
3. Restart Claude Desktop. The šØ **hammer icon** will appear in the chat interface showing your active tools!
---
## š Sample Interaction & Output
### Prompt:
> *"Audit both the 'income' and 'age' features for drift using DriftScope and give me an MLOps summary."*
### Output:
```text
## MLOps Drift Summary
| Feature | KS Stat | p-value | PSI Score | Verdict |
|---|---|---|---|---|
| income | 0.3960 | 0.00000 | 0.7060 | š“ Critical shift |
| age | 0.0840 | 0.05910 | 0.0230 | š¢ Stable |
Attention needed: income
- Both the KS test and PSI confirm income has drifted severely (p-value ā 0, PSI = 0.706 > 0.20 threshold).
- Recommendation: Upstream investigation required; trigger model retraining fallback pipeline.
```
---
## š» Local Development
Clone the repository and run locally using `uv`:
```bash
git clone https://github.com/Aymenrahmanii/driftscope-mcp.git
cd driftscope-mcp
# Install dependencies
uv sync
# Launch the interactive MCP Inspector UI
uv run mcp dev server.py
```
---
## š¦ Tech Stack
- **Protocol**: Model Context Protocol (MCP) Python SDK
- **Framework**: FastMCP (Starlette, Uvicorn, SSE)
- **Data & Math**: Polars, NumPy, SciPy, Scikit-learn, HTTPX
- **Infrastructure**: Render Web Services, GitHub
TDQS
Scored across 4 tools
check_feature_drift and compute_psi both assess feature drift but with distinct statistical methods (KS test vs. PSI) and different outputs (shift detection vs. severity). This is a clear enough separation, though an agent might still pause to choose between them for a generic drift-check request. generate_mock_datasets and get_monitoring_policy are unambiguous.
All tools use a consistent verb_noun snake_case pattern: generate_mock_datasets, check_feature_drift, get_monitoring_policy, compute_psi. Minor abbreviation in compute_psi doesn't break the convention. The set is highly predictable.
Four tools are well-scoped for a focused drift analysis server: one data generator, two drift metrics, and one reference tool. Each tool clearly earns its place without redundancy. The count fits comfortably within the ideal 3-15 range.
The surface covers continuous feature drift via KS and PSI, plus data generation and policy guidelines. However, it lacks categorical drift tests, batch/multi-feature analysis, and any policy update or model drift tools. These are notable gaps for a server named DriftScope.