ml-party
Planned optional adapter for mirroring locally tracked experiment runs and data to MLflow.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@ml-partyTrack my latest training run and compare it to previous runs."
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.

An agent-native platform for ML experiment tracking and lineage/knowledge.
Every existing tracker is human-dashboard-first: great at scalars-over-time, silent about why a run exists and how it relates to the others. ml-party treats a run's abstract, typed relationships, and reproducibility contract as first-class data — a durable, queryable lab notebook that agents write (over MCP) and query, and that humans read live (CLI + web viewer). It is also a full local tracker: metrics, artifacts, checkpoints, all stored locally; mirroring to W&B/MLflow is a future optional adapter, not the product.
Core ideas:
The write contract is the product.
run_startpre-registers intent (title, purpose, hypothesis, full parameters) and auto-captures the repro tuple (source snapshot, env lock, invocation, hardware, outer-git provenance).run_finalizerefuses without method + result/verdict + reproduce — refusals are machine-readable{missing, invalid}. Failures are knowledge (run_fail); silent deaths becomeabandoned.Internal git per experiment. A run snapshots the actually-running source — the directory you name, bounded by what git tracks there, never a guess — as a commit in a bare internal repo. A run references a commit — a launch-arg sweep is many runs on one commit, distinguished by mandatory
parameters. Declaredderives_fromlineage becomes the commit's parent; every run gets its own ref (no races).A knowledge graph over it. Typed nodes (
project / experiment / run / note), a controlled edge vocabulary (derives-from,supersedes,refutes, …), hybrid BM25(+optional embeddings)+graph retrieval that downranks superseded/refuted beliefs. Knowledge is append-only; correction is an edge.Boards: agents author whole views. An agent logs a self-contained HTML page as an artifact and the UI renders it sandboxed — comparison dashboards, demo galleries, live reports that fetch current data from the read-only API at view time (boards guide).
Run control through registered templates. Users register shell templates with typed placeholders — the allowlist; agents invoke them with validated, shell-quoted values (never commands), and every invocation is recorded in the graph, edged to the run it controlled. Restart is never a mutation: a new run,
derives-fromthe old (run-control guide).Local-first, remote-ready. Writers always write a local spool store; spool-and-flush sync ships runs to a served store over three idempotent streams (remote-tracking guide). Multi-user auth (roles, per-user tokens, UI login) activates with the first
mlp user add(deployment guide).Runs that execute elsewhere stay findable. Launching from a laptop onto a cluster, a queue or a cloud box?
run_set_computerecords the job system, its id and the link back to the job — and a run that declares it up front stops recording your laptop's hardware as its own. A pointer, not an integration — ml-party submits nothing and polls nothing, so it works with any job system, including your in-house one (tracking guide).
Quickstart
Requirements: Linux or macOS (Windows via WSL — the store relies on POSIX file locking) and Python ≥ 3.11. The web UI ships prebuilt in the wheel; nothing to compile, no Node required.
Option 1 — let your agent install it
Open your agent in the project you want to track and give it this prompt:
Look at https://github.com/paul-krug/ml-party — install ml-party and set up the MCP server for this project.It will find the install skill in this repo and follow it: install the package, reuse your store if you already have one (rather than starting a second, disconnected notebook), and register the server with your client.
Then restart your agent session. A newly registered MCP server cannot
load into the session that registered it — the skill tells your agent to say
this, but restart even if it forgets. On the next start, approve ml-party
when prompted; /mcp should then list it.
Related MCP server: Selvedge
Option 2 — install it yourself
You will need to follow three small steps.
1) Run the install. One personal store holds all your projects.
pip install mlparty
mlp init --root ~/.mlparty --no-mcp
export ML_PARTY_STORE=~/.mlparty # put this in your shell profile2) Make the MCP server visible to your agent.
Register the store once, globally, so it is there in every session.
Claude Code users paste:
claude mcp add -s user ml-party -- "$(which mlp)" serve-mcp --root ~/.mlpartyFor any other agent/harness, print the snippet and paste it into that client's own MCP configuration:
mlp mcp-configIt prints a standard mcpServers entry; its home differs per client
(Cursor's .cursor/mcp.json, and so on — a few use a different top-level
key, so check your client's docs).
Either way ml-party is then available in every session, in any directory,
and each codebase you work in becomes a project inside that one graph.
Prefer a store that lives inside one repository instead — because the notebook should travel with it, or a team shares it? That variant, and when it is worth the trade-off, is in the MCP setup guide.
3) Restart the session. A newly registered MCP server only loads on the
next start. Approve ml-party when prompted; /mcp lists what is active,
and claude mcp list shows what was loaded if it is missing.
See it running
mlp demo & # a real run: contract + live metrics
mlp ui # → http://127.0.0.1:7327Open the browser: the demo run is streaming its loss curve live. It went through the full lifecycle a real training does — pre-registered with purpose/hypothesis/parameters, metrics streamed, then finalized with a verdict. Click into it: Overview | Metrics | Artifacts | Code.
Then just tell your agent
The server is self-teaching: "use ml-party for this run" is all an agent
needs to hear. Connecting hands it a short brief pointing at a
workflow_guide tool that returns the full operating manual — a tool rather
than the connection text, because MCP clients truncate server instructions
(Claude Code at 2048 characters, silently), and a tool is the one channel an
agent can pull on its own initiative.
You can also just ask it about ml-party. The guides ship inside the
wheel, and a help tool serves them — so "how do I run the demo?", "how do I
instrument my training script?", "how do I share this store with my team?" are
answered from the installed package, with no checkout and no browser. help()
also reports every mlp command, the installed version, which store it is
serving, and where to file a bug.
Before the first run it will ask which directory holds your code. That
directory (source_root) is what gets snapshotted into the store, so the
agent shows you the file list before anything is written, and never captures
what your .gitignore excludes. Name no directory and the run simply
records no code — nothing is ever swept up silently.
Other setups: mlp connect --root <store> --project <dir> registers an
existing store for one more project.
Instrument a training script
The agent brackets the run over MCP and launches your script with
ML_PARTY_STORE/ML_PARTY_RUN set; the script attaches as the second
writer:
import mlparty
h = mlparty.attach() # env-var handshake (inert without ML_PARTY_RUN)
h.log_metric("loss", loss, step=step) # streams to the live view
h.log_artifact("ckpt/best.pt")
h.finalize(method=..., result={"summary": ..., "verdict": "confirmed",
"metrics": {"wer": 0.048}},
reproduce="python train.py --lr 1e-3")
# unhandled exceptions auto-fail the run with the tracebackCLI mirror: mlp status / runs / show <ref> / tail <run> / query "…" / diff <a> <b> / janitor / rebuild-index.
Web viewer
mlp ui serves a read-only SPA: experiments → runs → tabbed run pages with
live SSE metric dashboards, a finder-style artifact browser (image/audio/
video viewers, an .npy/.npz tensor slicer), agent-authored boards, the
lineage graph, search, and diff. Remote box → tunnel like TensorBoard:
ssh -L 7327:localhost:7327 <box>. For a shared server with logins and
sync ingest, see the deployment guide
and the remote-tracking guide.
From source
For development, or to run an unreleased revision — needs Node ≥ 20, since the UI bundle is built rather than downloaded:
git clone https://github.com/paul-krug/ml-party && cd ml-party
python -m venv .venv && .venv/bin/pip install -e .
(cd ui && npm install && npm run build) # web UI bundle, once
.venv/bin/python -m pytest tests/ -qContributions: CONTRIBUTING.md for the flow and the maintainer-only paths, AGENTS.md for conventions, SECURITY.md for the trust model.
Design
User guide (rendered from docs/): tracking runs, the web UI, boards, run control, MCP setup & tools, remote tracking, deployment & auth.
DESIGN.md — the living design document (ontology, contract, internal-git model, storage, surfaces, remote mode, forward design).
AGENTS.md — conventions for agents/contributors working in this repo (incl. the doc-sync rule). Planning lives in the repo's GitHub Project, not in tracked files.
Layering: pydantic ontology → journal-first store (journal.jsonl is the
source of truth; SQLite/FTS5 is a rebuildable index) → dulwich internal-git
engine → MlParty core API → thin frontends (MCP server, mlp CLI,
in-process client, read-only HTTP+SSE for the viewer).
Status
Beta (0.x): APIs may still move between minor versions; the store format is journal-first and rebuildable, and every release migrates it forward. Licensed Apache-2.0.
Tests: .venv/bin/python -m pytest tests/ -q
This server cannot be deployed
Maintenance
Related MCP Connectors
Cross-agent artifact workspace with provenance across Claude Code, Codex, Cursor, LangGraph.
- memnodeOAuthdev.memnode
Persistent, inspectable memory for AI agents with lineage, correction, and a hosted MCP endpoint.
Persistent knowledge graph for AI-augmented teams. Store decisions, findings, and standing rules across agent sessions with semantic search and typed connections. Includes cross-session memory, audit trail, workspace isolation, and secret detection. Built for teams running agents that need to remember. Free until launch with team tier as default, anon trial available.
Persistent memory and knowledge graphs for AI agents. Hybrid search, context checkpoints, and more.
Related MCP Servers
- FlicenseNot gradedqualityDmaintenanceProvides observability for multi-agent workflows by tracking hierarchical task structure, architectural decisions, reasoning, encountered problems, code modifications with Git diffs, and temporal metrics.57 npm-
- AlicenseAqualityAmaintenanceChange tracking for AI-era codebases. AI agents call it to log structured change events (entity + diff + reasoning) before the session ends, then query history with diff, blame, history, changeset, and search. Captures the intent that would otherwise evaporate.8348 PyPI23MIT
- AlicenseNot gradedqualityDmaintenanceEnables AI agents to read and write a local-first knowledge base of plain markdown files in git, with governance gates for safe, hash-anchored edits.1Apache 2.0
- AlicenseAqualityDmaintenanceProvides a shared, persistent workspace with versioned files, semantic search, run logging, and cross-agent provenance, allowing agents to maintain context across sessions and tools.208 npmApache 2.0