Skip to main content
Glama
Lesvitadsouza

tracker-mcp

README.md
# Experiment Tracker MCP

An MCP (Model Context Protocol) server that lets you **query your machine-learning experiments in natural language**. Train models, log every run to a local store, then ask an AI assistant — Claude Desktop, VS Code Copilot, or any MCP client — questions like *"which run had the best recall?"* and have it call your tools to answer.

The project is a small, self-contained demonstration of three things: a clean experiment-tracking data model, a backend-agnostic storage layer, and an intent-oriented MCP server on top of it.

---

## What it does

1. **`train.py`** trains scikit-learn models on a dataset and logs each run — hyperparameters, metrics, and the saved model file — through a tracker.
2. **The store** persists runs in SQLite behind a swappable interface.
3. **`server.py`** exposes the store as MCP tools, so an LLM can list, inspect, rank, and compare your experiments on request.

```
train models ──▶ log to store (SQLite) ──▶ MCP server ──▶ "which run is best?" (any MCP client)
```

---

## Architecture

The codebase is layered by responsibility, and **every dependency points inward toward an interface** rather than toward a concrete implementation:

```
server.py ─┐
           ├─▶ store/base.py  ◀── store/sqlite_store.py ──▶ db/connection.py ──▶ schema.sql
train.py ──┘   (Tracker: the        (SQLiteTracker:          (connection +        (the DDL)
                interface)           an implementation)       initialization)
```

`server.py` and `train.py` — the application edges — depend only on the abstract `Tracker` interface, never on SQLite. SQLite is an adapter plugged in behind the interface. This is the Dependency Inversion Principle in practice: swapping the backend (e.g. for MLflow) means adding one class in `store/`, and nothing above the interface changes.

---

## Project structure

```
tracker-mcp/
├── pyproject.toml          # dependencies + package metadata
├── .gitignore              # keeps tracker.db, models, __pycache__ out of git
├── README.md
├── .vscode/
│   └── mcp.json            # VS Code MCP server config
├── tracker_mcp/            # the package
│   ├── __init__.py         # re-exports Tracker, SQLiteTracker
│   ├── config.py           # single source of truth for filesystem paths
│   ├── db/                 # ── infrastructure layer ──
│   │   ├── schema.sql      #    the database structure, as real SQL
│   │   └── connection.py   #    connection config + schema initialization
│   ├── store/              # ── repository layer ──
│   │   ├── base.py         #    the Tracker interface (abstract)
│   │   └── sqlite_store.py #    the SQLite implementation
│   ├── train.py            # ── application: trains models, logs runs ──
│   └── server.py           # ── application: the FastMCP server ──
├── models/                 # saved .joblib model artifacts (gitignored)
└── tracker.db              # the SQLite store (gitignored)
```

---

## Setup

Requires **Python 3.12+**.

```bash
# from the project root
python -m venv .venv
source .venv/bin/activate          # Windows: .venv\Scripts\activate

pip install fastmcp scikit-learn joblib
pip install -e .                   # editable install: makes `python -m tracker_mcp.*` work anywhere
```

The editable install is what lets the MCP server be launched by name (`-m tracker_mcp.server`) from any working directory — important because MCP clients spawn the server from an unpredictable location.

---

## Usage

### 1. Train some models

```bash
python -m tracker_mcp.train
```

This trains four models on the scikit-learn breast-cancer dataset (a shallow decision tree as a weak baseline, a default and a tuned random forest, and gradient boosting), logs each run to `tracker.db`, and saves each model to `models/`.

### 2. Run the MCP server

```bash
python -m tracker_mcp.server      # stdio transport — what MCP clients expect
```

Or debug it interactively with the MCP Inspector:

```bash
fastmcp dev tracker_mcp/server.py
```

### 3. Connect it to a client

**Claude Desktop** — add to `claude_desktop_config.json`:

```json
{
  "mcpServers": {
    "experiment-tracker": {
      "command": "/absolute/path/to/tracker-mcp/.venv/bin/python",
      "args": ["-m", "tracker_mcp.server"]
    }
  }
}
```

**VS Code (GitHub Copilot, agent mode)** — `.vscode/mcp.json`. Note VS Code uses the `servers` key, not `mcpServers`:

```json
{
  "servers": {
    "experiment-tracker": {
      "type": "stdio",
      "command": "/absolute/path/to/tracker-mcp/.venv/bin/python",
      "args": ["-m", "tracker_mcp.server"]
    }
  }
}
```

Then ask, in natural language:

- *"List my experiment runs."*
- *"I'm building a cancer screening model — which run should I use?"*
- *"Compare the tuned random forest and gradient boosting."*
- *"Show me the details of run `910ced952bb6`."*

---

## The tools

The server exposes tools designed around **what a user asks**, not a 1:1 mirror of the store's methods:

| Tool | Purpose |
| --- | --- |
| `list_experiments(limit)` | Recent runs, newest first (summary). |
| `get_experiment(run_id)` | Full detail of one run, including all metrics. |
| `find_best_model(metric, mode)` | The run with the best value for a metric (`max`/`min`). |
| `compare_experiments(run_ids)` | Several runs side by side, by params and metrics. |

It also exposes each run as a read-only **resource** at `experiment://{run_id}`, and ships a server-level **`instructions`** block — domain context (what a run is, which metric fits which goal) delivered to the model during the MCP handshake, so it reasons about experiments correctly without the user re-explaining each turn.

---

## Design decisions

These are the choices worth understanding, and the reasoning behind them.

**Two tables, split by access pattern.** Runs get one table; metrics get their own. Hyperparameters are heterogeneous across model types and rarely queried individually, so they're stored as a single JSON column. Metrics are the opposite — you constantly rank and compare across them, and they can form time series over training steps — so they live in a normalized table where `MAX(value) WHERE key = 'accuracy'` is trivial. This *params-as-JSON, metrics-as-rows* asymmetry is the core of the data model.

**Backend-agnostic storage.** The server and training code depend on the abstract `Tracker` interface, so the SQLite backend is a detail, not an assumption. A different backend (MLflow, Postgres) would be a new class behind the same interface.

**Intent-oriented tools.** MCP tools are shaped around user questions (`find_best_model`, `compare_experiments`) built on top of the store's primitives, rather than mechanically exposing `get_run`/`list_runs`. Clear, distinct tool descriptions are what let the model pick the right tool.

**Single source of truth for paths.** `config.py` anchors the database and model directory to the project root via `__file__`, so the training script (writer) and the server (reader) always resolve to the same files regardless of the working directory a command is launched from. Env vars (`TRACKER_DB`, `TRACKER_MODELS`) override when needed.

**Reproducibility as a logged value.** The random seed lives inside each run's logged `params`, so a stored run can be recreated exactly — the same dict both builds the model and gets recorded.

**Failures are recorded.** A run is created before training starts and marked `failed` if training raises, so the experiment history reflects reality rather than only successes.

---

## Tech stack

- **[FastMCP](https://gofastmcp.com)** (3.x) — the MCP server framework
- **scikit-learn** — model training
- **SQLite** (stdlib `sqlite3`) — the tracking store
- **joblib** — model serialization

---

## Possible extensions

- An `MLflowTracker` implementation behind the same `Tracker` interface.
- A test suite under `tests/` (the in-memory `fastmcp.Client` makes end-to-end tool tests straightforward).
- Per-step metric logging to support training curves, and a `best_run` that ranks by each run's final step.
- A `delete_run` / archival tool, and filtering `list_experiments` by model type or dataset.