Skip to main content
Glama

ml-party in 40 seconds: connect an agent over MCP, let it track a sweep, watch the dashboard live, ask questions weeks later

An agent-native platform for ML experiment tracking and lineage/knowledge.

Every existing tracker is human-dashboard-first: great at scalars-over-time, silent about why a run exists and how it relates to the others. ml-party treats a run's abstract, typed relationships, and reproducibility contract as first-class data — a durable, queryable lab notebook that agents write (over MCP) and query, and that humans read live (CLI + web viewer). It is also a full local tracker: metrics, artifacts, checkpoints, all stored locally; mirroring to W&B/MLflow is a future optional adapter, not the product.

Core ideas:

  • The write contract is the product. run_start pre-registers intent (title, purpose, hypothesis, full parameters) and auto-captures the repro tuple (source snapshot, env lock, invocation, hardware, outer-git provenance). run_finalize refuses without method + result/verdict + reproduce — refusals are machine-readable {missing, invalid}. Failures are knowledge (run_fail); silent deaths become abandoned.

  • Internal git per experiment. A run snapshots the actually-running source — the directory you name, bounded by what git tracks there, never a guess — as a commit in a bare internal repo. A run references a commit — a launch-arg sweep is many runs on one commit, distinguished by mandatory parameters. Declared derives_from lineage becomes the commit's parent; every run gets its own ref (no races).

  • A knowledge graph over it. Typed nodes (project / experiment / run / note), a controlled edge vocabulary (derives-from, supersedes, refutes, …), hybrid BM25(+optional embeddings)+graph retrieval that downranks superseded/refuted beliefs. Knowledge is append-only; correction is an edge.

  • Boards: agents author whole views. An agent logs a self-contained HTML page as an artifact and the UI renders it sandboxed — comparison dashboards, demo galleries, live reports that fetch current data from the read-only API at view time (boards guide).

  • Run control through registered templates. Users register shell templates with typed placeholders — the allowlist; agents invoke them with validated, shell-quoted values (never commands), and every invocation is recorded in the graph, edged to the run it controlled. Restart is never a mutation: a new run, derives-from the old (run-control guide).

  • Local-first, remote-ready. Writers always write a local spool store; spool-and-flush sync ships runs to a served store over three idempotent streams (remote-tracking guide). Multi-user auth (roles, per-user tokens, UI login) activates with the first mlp user add (deployment guide).

  • Runs that execute elsewhere stay findable. Launching from a laptop onto a cluster, a queue or a cloud box? run_set_compute records the job system, its id and the link back to the job — and a run that declares it up front stops recording your laptop's hardware as its own. A pointer, not an integration — ml-party submits nothing and polls nothing, so it works with any job system, including your in-house one (tracking guide).

Quickstart

Requirements: Linux or macOS (Windows via WSL — the store relies on POSIX file locking) and Python ≥ 3.11. The web UI ships prebuilt in the wheel; nothing to compile, no Node required.

Option 1 — let your agent install it

Open your agent in the project you want to track and give it this prompt:

Look at https://github.com/paul-krug/ml-party — install ml-party and set up the MCP server for this project.

It will find the install skill in this repo and follow it: install the package, reuse your store if you already have one (rather than starting a second, disconnected notebook), and register the server with your client.

Then restart your agent session. A newly registered MCP server cannot load into the session that registered it — the skill tells your agent to say this, but restart even if it forgets. On the next start, approve ml-party when prompted; /mcp should then list it.

Related MCP server: Selvedge

Option 2 — install it yourself

You will need to follow three small steps.

1) Run the install. One personal store holds all your projects.

pip install mlparty
mlp init --root ~/.mlparty --no-mcp
export ML_PARTY_STORE=~/.mlparty      # put this in your shell profile

2) Make the MCP server visible to your agent.

Register the store once, globally, so it is there in every session.

Claude Code users paste:

claude mcp add -s user ml-party -- "$(which mlp)" serve-mcp --root ~/.mlparty

For any other agent/harness, print the snippet and paste it into that client's own MCP configuration:

mlp mcp-config

It prints a standard mcpServers entry; its home differs per client (Cursor's .cursor/mcp.json, and so on — a few use a different top-level key, so check your client's docs).

Either way ml-party is then available in every session, in any directory, and each codebase you work in becomes a project inside that one graph.

Prefer a store that lives inside one repository instead — because the notebook should travel with it, or a team shares it? That variant, and when it is worth the trade-off, is in the MCP setup guide.

3) Restart the session. A newly registered MCP server only loads on the next start. Approve ml-party when prompted; /mcp lists what is active, and claude mcp list shows what was loaded if it is missing.

See it running

mlp demo &                                  # a real run: contract + live metrics
mlp ui                                      # → http://127.0.0.1:7327

Open the browser: the demo run is streaming its loss curve live. It went through the full lifecycle a real training does — pre-registered with purpose/hypothesis/parameters, metrics streamed, then finalized with a verdict. Click into it: Overview | Metrics | Artifacts | Code.

Then just tell your agent

The server is self-teaching: "use ml-party for this run" is all an agent needs to hear. Connecting hands it a short brief pointing at a workflow_guide tool that returns the full operating manual — a tool rather than the connection text, because MCP clients truncate server instructions (Claude Code at 2048 characters, silently), and a tool is the one channel an agent can pull on its own initiative.

You can also just ask it about ml-party. The guides ship inside the wheel, and a help tool serves them — so "how do I run the demo?", "how do I instrument my training script?", "how do I share this store with my team?" are answered from the installed package, with no checkout and no browser. help() also reports every mlp command, the installed version, which store it is serving, and where to file a bug.

Before the first run it will ask which directory holds your code. That directory (source_root) is what gets snapshotted into the store, so the agent shows you the file list before anything is written, and never captures what your .gitignore excludes. Name no directory and the run simply records no code — nothing is ever swept up silently.

Other setups: mlp connect --root <store> --project <dir> registers an existing store for one more project.

Instrument a training script

The agent brackets the run over MCP and launches your script with ML_PARTY_STORE/ML_PARTY_RUN set; the script attaches as the second writer:

import mlparty

h = mlparty.attach()                        # env-var handshake (inert without ML_PARTY_RUN)
h.log_metric("loss", loss, step=step)       # streams to the live view
h.log_artifact("ckpt/best.pt")
h.finalize(method=..., result={"summary": ..., "verdict": "confirmed",
                               "metrics": {"wer": 0.048}},
           reproduce="python train.py --lr 1e-3")
# unhandled exceptions auto-fail the run with the traceback

CLI mirror: mlp status / runs / show <ref> / tail <run> / query "…" / diff <a> <b> / janitor / rebuild-index.

Web viewer

mlp ui serves a read-only SPA: experiments → runs → tabbed run pages with live SSE metric dashboards, a finder-style artifact browser (image/audio/ video viewers, an .npy/.npz tensor slicer), agent-authored boards, the lineage graph, search, and diff. Remote box → tunnel like TensorBoard: ssh -L 7327:localhost:7327 <box>. For a shared server with logins and sync ingest, see the deployment guide and the remote-tracking guide.

From source

For development, or to run an unreleased revision — needs Node ≥ 20, since the UI bundle is built rather than downloaded:

git clone https://github.com/paul-krug/ml-party && cd ml-party
python -m venv .venv && .venv/bin/pip install -e .
(cd ui && npm install && npm run build)     # web UI bundle, once
.venv/bin/python -m pytest tests/ -q

Contributions: CONTRIBUTING.md for the flow and the maintainer-only paths, AGENTS.md for conventions, SECURITY.md for the trust model.

Design

Layering: pydantic ontology → journal-first store (journal.jsonl is the source of truth; SQLite/FTS5 is a rebuildable index) → dulwich internal-git engine → MlParty core API → thin frontends (MCP server, mlp CLI, in-process client, read-only HTTP+SSE for the viewer).

Status

Beta (0.x): APIs may still move between minor versions; the store format is journal-first and rebuildable, and every release migrates it forward. Licensed Apache-2.0.

Tests: .venv/bin/python -m pytest tests/ -q

Related MCP Connectors

Related MCP Servers