mdl-train-mcp
Allows monitoring and managing training jobs on Modal, including listing apps, browsing logs with summary/window/grep modes, and stopping running apps.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@mdl-train-mcpwhat training jobs are currently running?"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
mdl-train-mcp
An MCP server for monitoring and managing training jobs on Modal. Built for LLMs that need to check on long-running GPU training without drowning in log output.
Why?
Training logs on Modal can be tens of thousands of lines — weight loading bars, Omniverse init spam, 8 ranks of identical output. Dumping all of that into an LLM context is wasteful and often hits resource limits.
This server gives you browsable logs: start with a summary, then drill into what matters.
Tool | What it does |
| List running, deployed, and recent apps with filtering |
| Browse logs with summary/window/grep modes |
| Stop a running app |
Related MCP server: MCP Task
The get_logs workflow
Instead of returning a giant blob, get_logs has three modes:
1. Summary (default) — returns line count, first/last 10 lines, and any errors with line numbers. Small response, always works.
get_logs(app_id="ap-xxx")
→ {total_lines: 30000, errors: [{line: 847, text: "CUDA error: ..."}], head: [...], tail: [...]}2. Window — read a specific range. Like scrolling through a file.
get_logs(app_id="ap-xxx", window_start=840, window_size=30)
→ 30 lines around the error3. Grep — search with regex and context lines. Like grep -C.
get_logs(app_id="ap-xxx", grep="Error|Traceback", grep_context=15)
→ all errors with 15 lines of surrounding contextLandmarks — pass landmark_patterns in summary mode to get a table of contents:
get_logs(app_id="ap-xxx", landmark_patterns=["Iteration \\d+", "success_rate", "checkpoint"])
→ landmarks: [{line: 200, text: "Iteration 1/3000"}, {line: 5000, text: "success_rate: 0.95"}, ...]Landmark sampling is fair across patterns — one pattern won't dominate.
Features
Progress bar collapsing — tqdm bars, HF weight loading, and downloads are collapsed to their latest update (50 progress lines → 1 showing current state)
Auto-retry on resource limits — if Modal's API rejects a large
tail, automatically retries with smaller values and tells you what happenedError deduplication — 10,000 identical
[Error]lines become a handful of unique entriesCase-sensitive error detection — won't false-positive on metric names like
rot_align_error
Setup
1. Install
# Using uv (recommended)
uv pip install mdl-train-mcp
# Or from source
git clone https://github.com/JoshuaSP/mdl-train-mcp
cd mdl-train-mcp
uv venv && uv pip install -e .2. Configure Modal
Make sure you have the Modal CLI installed and authenticated:
pip install modal
modal setup3. Add to Claude Code
Add to your .mcp.json:
{
"mcpServers": {
"mdl": {
"command": "mdl-train-mcp",
"env": {
"MODAL_PROFILE": "your-profile"
}
}
}
}Or from source:
{
"mcpServers": {
"mdl": {
"command": "uv",
"args": ["--directory", "/path/to/mdl-train-mcp", "run", "mdl-train-mcp"],
"env": {
"MODAL_BIN": "/path/to/modal",
"MODAL_PROFILE": "your-profile"
}
}
}
}Environment variables
Variable | Description | Default |
| Path to modal CLI binary |
|
| Modal profile to use | (default profile) |
Tools reference
list_apps
list_apps(state?: string, name_contains?: string)Filter by state ("running", "deployed", "stopped", "ephemeral") or name substring.
get_logs
get_logs(
app_id: string,
tail?: number, # log entries to fetch (default 500, max 5000)
since?: string, # "1h", "30m", "2d", or ISO datetime
until?: string,
source?: string, # "stdout", "stderr", "system"
window_start?: number, # line number for window mode
window_size?: number, # lines to return (default 50, max 200)
grep?: string, # regex search (case-insensitive)
grep_context?: number, # context lines around matches (max 30)
landmark_patterns?: string[] # regex patterns for summary landmarks
)stop_app
stop_app(app_id: string)Irreversible — terminates the app and all its containers.
License
MIT
This server cannot be deployed
Maintenance
Related MCP Connectors
Hosted MCP server for LLM cost estimation, model comparison, and budget-aware routing.
MCP server for building and testing AI agents with multi-model experimentation and insights.
Remote MCP server for supportsheep: run AI interviews and manage support content for your blog.
Related MCP Servers
- AlicenseAqualityDmaintenanceA Model Context Protocol server that enables LLMs and AI assistants to create, manage, and interact with isolated cloud-based Python environments with GPU support on Modal.com.111MIT
- AlicenseAqualityBmaintenanceAsync MCP server for running long-running AI tasks with real-time progress monitoring, enabling users to start, monitor, and manage complex AI workflows across multiple models.658 npm5MIT
- AlicenseAqualityDmaintenanceMCP server for log file analysis. Gives LLMs the ability to efficiently analyze large log files without loading them into context.7100MIT
- AlicenseAqualityAmaintenanceAn MCP server for managing Modal — apps, containers, volumes, and secrets — and for deploying & running Modal apps directly from Claude Code and other MCP clients.1289 PyPI2MIT