Data Analytics MCP Toolkit
# Data Analytics MCP Toolkit
An MCP (Model Context Protocol) server that exposes data visualization and simple machine learning tools. When an external LLM calls the toolkit, it can use the high-level **run_analytics** tool to describe intent and data; the server selects and runs the appropriate pipeline (visualization or ML) and returns charts or metrics.
## Features
- **Data**: `load_data` (CSV/JSON string or URL), `clean_data` (drop NA, optional normalize)
- **Visualization**: `plot_bar`, `plot_line`, `plot_scatter`, `plot_histogram`, `plot_box`, `plot_heatmap` (return base64 PNG)
- **ML**: `train_test_split`, `train_linear_regression`, `train_logistic_regression`, `train_kmeans`, plus `evaluate_regression`, `evaluate_classification`, `evaluate_clustering`
- **Pipeline**: `run_analytics(intent, data_source)` — intent-based routing to the right pipeline
## Install
```bash
cd /path/to/trying_IBM_MCP
pip install -e .
# or
pip install -r requirements.txt
```
From the project root, ensure `src` is on `PYTHONPATH` when running the server (or install in editable mode).
## Run the MCP server
**stdio (for Cursor / IDE):**
```bash
# From project root, with src on path
PYTHONPATH=src python -m data_analytics_mcp.server
```
Or with uv:
```bash
uv run --project . python -m data_analytics_mcp.server
```
(If using a `pyproject.toml` that sets `packages` under `src`, install first with `pip install -e .` then run `python -m data_analytics_mcp.server` from the repo root.)
## Cursor MCP configuration
Add the server to Cursor (e.g. in Cursor Settings → MCP, or project `.cursor/mcp.json`):
```json
{
"mcpServers": {
"data-analytics": {
"command": "python",
"args": ["-m", "data_analytics_mcp.server"],
"cwd": "/path/to/trying_IBM_MCP",
"env": { "PYTHONPATH": "src" }
}
}
}
```
Use the full path for `cwd`. If you installed the package (`pip install -e .`), you can use:
```json
{
"mcpServers": {
"data-analytics": {
"command": "python",
"args": ["-m", "data_analytics_mcp.server"],
"cwd": "/Users/jerrychen/projects/trying_IBM_MCP"
}
}
}
```
## Usage
- **One-shot**: Call `run_analytics` with a natural-language intent (e.g. "show distribution of sales", "predict price from square_feet", "cluster into 4 groups") and the data as CSV/JSON string or URL. The server returns either a chart (base64 image) or ML metrics and a short model summary.
- **Step-by-step**: Use `load_data` → get `data_id` → then call `clean_data`, `plot_*`, or `train_test_split` → `train_*` → `evaluate_*` as needed. Use resources `analytics://pipelines` and `analytics://pipelines/visualization` (etc.) to see pipeline descriptions.
## Project layout
```
src/data_analytics_mcp/
server.py # MCP app, tools, resources
pipeline.py # Intent → pipeline; execute_pipeline
data.py # load_data, clean_data
viz.py # Plot functions → base64 PNG
ml.py # Train/evaluate regression, classification, clustering
store.py # In-memory session store
```
TDQS
Scored across 16 tools
Most tools have distinct purposes, with clear separation between data loading/cleaning, plotting, model training, and evaluation functions. However, 'run_analytics' overlaps significantly with the specialized tools, as it can perform many of the same functions through a single interface, which could cause confusion about when to use it versus the specific tools.
Tool names follow a consistent verb_noun pattern throughout, with clear and predictable naming conventions. All tools use snake_case, and verbs like 'clean', 'evaluate', 'load', 'plot', 'run', 'train', and 'train_test_split' are applied consistently to their respective nouns, making the set highly readable and predictable.
With 16 tools, the count is slightly high but reasonable for a comprehensive data analytics toolkit. It covers a broad range of functions from data ingestion to visualization and machine learning, though it might be borderline heavy for some use cases. Each tool appears to earn its place without obvious redundancy, except for the overlap with 'run_analytics'.
The tool set provides complete coverage for a data analytics workflow, including data loading, cleaning, splitting, multiple types of plots, training for classification, regression, and clustering models, and corresponding evaluation metrics. The inclusion of 'run_analytics' as a high-level tool further ensures no gaps, offering a flexible alternative for common tasks.