Skip to main content
Glama

šŸ”¬ DriftScope — Autonomous MLOps & Statistical Diagnostics MCP Server

Python Protocol Deployment Engine

An open-source, cloud-deployed Model Context Protocol (MCP) server that equips Large Language Models (LLMs) with deterministic statistical computing engines to audit data drift and model degradation in production tabular pipelines.


šŸ’” The Problem

LLMs are exceptional at high-level reasoning, code generation, and root-cause analysis, but notoriously unreliable at precise statistical math. When monitoring machine learning pipelines, asking an LLM to evaluate distribution shift directly leads to severe numerical hallucinations.

DriftScope solves this by bridging the LLM to a dedicated Python analytical engine via the open Model Context Protocol:

  • The LLM handles orchestration, triage, hypothesis generation, and incident reporting.

  • DriftScope executes deterministic, vector-accelerated hypothesis testing (Two-sample Kolmogorov-Smirnov) and Population Stability Index (PSI) calculations using Polars and SciPy.


Related MCP server: CI-1T Prediction Stability Engine

šŸ—ļø Architecture

ā”Œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”
│                      LLM Host                          │
│         (Claude Desktop / Cursor / Custom Agent)       │
ā””ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”¬ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”˜
                            │ JSON-RPC 2.0
                            ā–¼ (mcp-remote bridge)
               ā”Œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”
               │    Internet / HTTPS    │
               ā””ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”¬ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”˜
                            │ Server-Sent Events (SSE)
                            ā–¼
ā”Œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”
│             DriftScope MCP Server (Render)             │
│  ā”Œā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”  │
│  │                  FastMCP Router                  │  │
│  ā””ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”¬ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”¬ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”¬ā”€ā”˜  │
│           │                       │               │    │
│           ā–¼                       ā–¼               ā–¼    │
│    [ Tools Engine ]      [ Resources Hub ]   [ Prompts ]│
│    • check_drift         • standards://      • audit_   │
│    • compute_psi           drift-policy        feature │
│    • generate_mock                                     │
│           │                                            │
│           ā–¼                                            │
│    [ Data Layer: Polars + SciPy + NumPy ]              │
ā””ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”€ā”˜

✨ Features & MCP Primitives

1. šŸ› ļø Tools (Callable Actions)

  • check_feature_drift(baseline_csv, current_csv, feature_column, significance_level): Executes two-sample Kolmogorov-Smirnov (KS) tests to detect continuous covariate shift with rigorous $p$-values.

  • compute_psi(baseline_csv, current_csv, feature_column, bins): Computes the Population Stability Index (PSI) with equal-frequency baseline binning and Laplace smoothing to classify shift severity (Stable, Moderate, or Critical).

  • generate_mock_datasets(): Synthesizes baseline and drifted production samples on the fly for pipeline verification.

Supports both local cloud file paths and remote public HTTP/HTTPS URLs (AWS S3, GitHub raw, etc.).

2. šŸ“‚ Resources (Passive Knowledge)

  • standards://drift-policy: Exposes organization-wide MLOps threshold standards (e.g., $p < 0.05$ rejection criteria, PSI warning zones) directly into the agent's context.

3. šŸ“ Prompts (Reusable Workflows)

  • audit_feature: Pre-engineered diagnostic prompt guiding the model through end-to-end drift triage, impact evaluation, and retraining recommendations.


šŸš€ Live Cloud Deployment

DriftScope is deployed as a live cloud service on Render utilizing the Server-Sent Events (SSE) transport.

  • Live SSE Endpoint: https://driftscope-mcp.onrender.com/sse


šŸ”Œ Quickstart: Connect to Claude Desktop

You can connect your local Claude Desktop to the live cloud deployment in seconds:

  1. Open your Claude Desktop configuration file:

    • Windows: %APPDATA%\Claude\claude_desktop_config.json

    • macOS: ~/Library/Application Support/Claude/claude_desktop_config.json

  2. Add driftscope-cloud to your mcpServers object:

{
  "mcpServers": {
    "driftscope-cloud": {
      "command": "npx",
      "args": [
        "-y",
        "mcp-remote",
        "https://driftscope-mcp.onrender.com/sse"
      ]
    }
  }
}
  1. Restart Claude Desktop. The šŸ”Ø hammer icon will appear in the chat interface showing your active tools!


šŸ“Š Sample Interaction & Output

Prompt:

"Audit both the 'income' and 'age' features for drift using DriftScope and give me an MLOps summary."

Output:

## MLOps Drift Summary

| Feature | KS Stat | p-value | PSI Score | Verdict |
|---|---|---|---|---|
| income  | 0.3960  | 0.00000 | 0.7060    | šŸ”“ Critical shift |
| age     | 0.0840  | 0.05910 | 0.0230    | 🟢 Stable |

Attention needed: income
- Both the KS test and PSI confirm income has drifted severely (p-value ā‰ˆ 0, PSI = 0.706 > 0.20 threshold).
- Recommendation: Upstream investigation required; trigger model retraining fallback pipeline.

šŸ’» Local Development

Clone the repository and run locally using uv:

git clone https://github.com/Aymenrahmanii/driftscope-mcp.git
cd driftscope-mcp

# Install dependencies
uv sync

# Launch the interactive MCP Inspector UI
uv run mcp dev server.py

šŸ“¦ Tech Stack

  • Protocol: Model Context Protocol (MCP) Python SDK

  • Framework: FastMCP (Starlette, Uvicorn, SSE)

  • Data & Math: Polars, NumPy, SciPy, Scikit-learn, HTTPX

  • Infrastructure: Render Web Services, GitHub

Available Tools

4 tools
check_feature_driftB

Evaluates covariate shift for a continuous feature between baseline and production datasets using the Kolmogorov-Smirnov (KS) test. Accepts both local file paths and public HTTP/HTTPS URLs.

ParametersJSON Schema
NameRequiredDescriptionDefault
current_csvYes
baseline_csvYes
feature_columnYes
significance_levelNo

TDQS

B3.2/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It usefully discloses that inputs may be public HTTP/HTTPS URLs (network access) as well as local paths, but says nothing about the return value (p-value, statistic, verdict), sampling behavior, or failure modes for non-numeric features.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences, front-loaded with the operation and method, followed by the input-source caveat. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no annotations, no output schema, and 0% schema description coverage, the description should explain what a caller receives (test statistic, p-value, pass/fail) and the role of significance_level. Neither appears, leaving the tool under-specified for its complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the schema alone documents nothing. The description partially compensates by explaining that baseline_csv/current_csv can be paths or URLs and that the feature must be continuous, but significance_level and the exact meaning of feature_column go unaddressed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (evaluates), a precise statistical target (covariate shift for a continuous feature), the datasets involved, and the exact test method (Kolmogorov-Smirnov). That is far more specific than the sibling names alone, though it never explicitly distinguishes itself from compute_psi, which is another drift metric.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no statement of when to prefer this over compute_psi or the monitoring-policy tools, and no mention of prerequisites such as minimum sample size or the data types the KS test requires. Usage context is only inferable from the word 'drift'.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

compute_psiB

Computes Population Stability Index (PSI) to evaluate feature shift severity. Accepts both local file paths and public HTTP/HTTPS URLs.

ParametersJSON Schema
NameRequiredDescriptionDefault
binsNo
current_csvYes
baseline_csvYes
feature_columnYes

TDQS

B3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses that inputs may be local paths or public HTTP/HTTPS URLs, which is real context beyond the schema, but says nothing about return values, error behavior for unreachable URLs, or computational cost.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two short sentences, front-loaded with the core purpose and with no filler. Every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

A 4-parameter computation tool with no annotations, no output schema, and 0% schema description coverage needs far more: what the inputs must contain, what bins controls, and what the result looks like. The description leaves most of this unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0% across 4 parameters, so the description must compensate. The file-path/URL hint partially explains the two CSV inputs, but baseline_csv, current_csv, feature_column, and especially bins (default 10) are never given meaning.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Computes Population Stability Index') and its intent ('evaluate feature shift severity'). However, it never distinguishes itself from the close sibling check_feature_drift, which plausibly covers similar ground, so an agent cannot route between them from the description alone.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is a soft usage cue ('evaluate feature shift severity') and input-source coverage, but no explicit when-to-use, when-not-to-use, or comparison against check_feature_drift/generate_mock_datasets. The agent is left to infer which drift tool applies.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

generate_mock_datasetsB

Generates synthetic baseline and production CSVs in the local cloud storage.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations exist, so the description carries the full burden. It does hint at a write side effect by naming a storage destination, but it never states whether existing files are overwritten, what permissions are needed, how large the generated data is, or whether the operation is repeatable. For a mutating generator this is a significant gap.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single front-loaded sentence with no filler; every word (synthetic, baseline, production, CSVs, storage) contributes to the meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema and no annotations, the description should explain what gets written and what is returned, but 'local cloud storage' remains vague and the difference between the baseline and production outputs is never characterized. Adequate but leaves real ambiguity for a side-effecting tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has zero parameters, so there is nothing for the description to disambiguate. Baseline 4 applies per the rubric.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb ('Generates') plus a concrete resource ('synthetic baseline and production CSVs') with a stated destination. It is unambiguous what the tool produces, though it does not explicitly contrast itself with the sibling monitoring tools (check_feature_drift, compute_psi).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no indication of when to use this tool, no prerequisites, and no conditions that would select it over the sibling tools. An agent must infer that this is a fixture/setup step from the word 'mock' alone.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

get_monitoring_policyA

Returns organizational reference guidelines for drift and model monitoring thresholds.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A3.7/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. 'Returns' implies a read-only fetch, but it does not state permission requirements, whether the policy is cached, or how it relates to the drift-checking siblings. For a simple zero-parameter getter this is adequate, but it adds little beyond the verb.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single, front-loaded sentence with no filler. The resource and scope are stated immediately and nothing is repeated from the schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be explained in prose, and the tool is a trivial zero-param read. The description covers what is needed; only the relationship to the drift-evaluation siblings is unaddressed.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters, so there is no parameter semantics to document and the baseline of 4 applies. Nothing in the description misleads about inputs.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description gives a specific verb ('Returns') and a well-defined resource ('organizational reference guidelines for drift and model monitoring thresholds'), so an agent can tell what comes back. It never differentiates itself from siblings like check_feature_drift or compute_psi, which is the only thing keeping it from a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There is no explicit when-to-use, when-not-to-use, or alternative routing. Usage is only implied: this looks like the tool that supplies reference thresholds before drift evaluation. That inference is reasonable but not stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 4 tool updatesv0.1.0
    • First observedcheck_feature_drift
    • First observedcompute_psi
    • First observedgenerate_mock_datasets
    • First observedget_monitoring_policy

TDQS

A3.5/5.0

Scored across 4 tools

Disambiguation4/5

check_feature_drift and compute_psi both assess feature drift but with distinct statistical methods (KS test vs. PSI) and different outputs (shift detection vs. severity). This is a clear enough separation, though an agent might still pause to choose between them for a generic drift-check request. generate_mock_datasets and get_monitoring_policy are unambiguous.

Naming Consistency5/5

All tools use a consistent verb_noun snake_case pattern: generate_mock_datasets, check_feature_drift, get_monitoring_policy, compute_psi. Minor abbreviation in compute_psi doesn't break the convention. The set is highly predictable.

Tool Count5/5

Four tools are well-scoped for a focused drift analysis server: one data generator, two drift metrics, and one reference tool. Each tool clearly earns its place without redundancy. The count fits comfortably within the ideal 3-15 range.

Completeness3/5

The surface covers continuous feature drift via KS and PSI, plus data generation and policy guidelines. However, it lacks categorical drift tests, batch/multi-feature analysis, and any policy update or model drift tools. These are notable gaps for a server named DriftScope.

Maintenance

ActivityMaintained
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    Provides advanced evaluation tools for assessing AI safety, alignment, and performance of LLM outputs. Enables programmatic evaluation of quality, safety metrics like toxicity and PII detection, and operational metrics including carbon footprint and cost estimation.
    4
    Apache 2.0
  • A
    license
    Not graded
    quality
    C
    maintenance
    Enables language models to run data-quality checks and profiling on local files, using dbt-style assertions like not_null, unique, relationships, and accepted_values.
    MIT