MCP Stats Tools Server
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@MCP Stats Tools ServerRun a paired significance test on these two samples and report the p-values"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
mcp-stats-tools
An MCP server with three statistical tools that an agent can call:
survival_analysis: Kaplan-Meier curves, plus a log-rank test when exactly two groups are givencompare_two_samples: paired t-test, Wilcoxon signed-rank test and a bootstrap CI on the mean differencecheck_distribution_drift: Kolmogorov-Smirnov test and Population Stability Index between a reference and a current sample
No API keys or paid services. The tools are plain statistics, not model calls.
Running
pip install -r requirements.txt
make test # 18 tests
make evaluate # quality, latency and cost report
make run # server over stdioTo poke at it with the MCP Inspector:
npx @modelcontextprotocol/inspector python -m app.serverRelated MCP server: MCP ML Monitor
Layout
app/tools.py the statistics, no MCP imports
app/server.py MCP tool definitions
app/evaluate.py evaluation report
tests/ unit tests for tools.py and in-process tests through the serverThe logic in tools.py raises ValueError on bad input. server.py converts that
to ToolError, because the SDK treats any other exception as an unexpected crash
and the client never sees the message.
Tool functions return dict[str, Any] rather than a bare dict. With a bare
dict the SDK can't build an output schema and results only come back as text,
not structured_content.
Evaluation
python -m app.evaluate prints three sections.
Quality. Inputs with a known answer: two groups with different hazards should
give a significant log-rank test and two identical ones should not, a 10-unit
shift should be recovered by compare_two_samples, and drift should be flagged for
a shifted sample but not a stable one.
Latency. 50 calls per tool through server.call_tool, so argument validation
and serialisation are included. On my machine drift takes about 1 ms, the paired
comparison about 9 ms and survival analysis about 40 ms (it bootstraps).
Cost. The tools run locally, so there's no bill to read from. The report
estimates monthly compute cost from mean latency, a daily call volume and an
assumed per-vCPU-hour rate set at the top of evaluate.py.
Limitations
Only stdio transport is used. HTTP isn't exercised.
The tests call the server in-process. Nothing here connects a real agent to it.
The cost figure is an estimate, not a measured bill.
This server cannot be deployed
Maintenance
Related MCP Connectors
DriftOracle - 15 tools for model/data drift monitoring: PSI, KS-test, alerts, evidence packs.
Valid and reliable data engineering and statistical analysis without hallucinations.
The statistical analyst in your AI chat — validated, citable, re-runnable analysis of your data.
Data observability tools for engineering teams: alerts, freshness, schema drift, lineage, quality.
Related MCP Servers
AlicenseNot gradedqualityCmaintenanceProvides advanced evaluation tools for assessing AI safety, alignment, and performance of LLM outputs. Enables programmatic evaluation of quality, safety metrics like toxicity and PII detection, and operational metrics including carbon footprint and cost estimation.4Apache 2.0- FlicenseNot gradedqualityDmaintenanceMonitors ML models in production for data drift and performance degradation, providing automated alerts and retraining recommendations.-
- AlicenseNot gradedqualityBmaintenanceEnables LLM evaluation and observability by uploading documents, building test sets, running RAG pipelines, and automatically scoring answers for groundedness, hallucination risk, retrieval quality, latency, and cost, with tools exposed to MCP-compatible clients.1MIT
- FlicenseAqualityBmaintenanceEnables LLMs to audit data drift and model degradation in tabular ML pipelines through deterministic statistical tests such as Kolmogorov-Smirnov and Population Stability Index, plus reusable prompts and standards resources.4-