jobd
The server is a self-hostable, GPU-aware job broker with an MCP interface for submitting, monitoring, and managing jobs across a fleet of workers.
Submit jobs asynchronously or synchronously (
jobd_submit), with options for GPU/VRAM requirements, CPU/memory profiles, host pins, tags, environment variables, dependencies, retries, timeouts, preemption/checkpoint settings, and job arrays viacountor parametersweep.Dry-run submissions to preview routing and validation without queueing.
Check job status (
jobd_status) for full records: state, exit code, timings, host, scheduling internals, and routing diagnostics; optionally block until terminal.Read job logs (
jobd_logs) for running or finished jobs, with tail sizing up to 1 MiB.Cancel jobs (
jobd_cancel) with an optional reason, affecting queued or running jobs.Preempt running/assigned preemptible jobs (
jobd_preempt) with a grace window for checkpointing.List jobs (
jobd_list) filtered by state/project, with per-state counts and active-set defaults.Inspect events (
jobd_events) to understand why jobs were blocked, skipped, dispatched, preempted, or failed, with filters by time, type, job, project, or source.View fleet health (
jobd_workers) — worker states, live VRAM/RAM/CPU capacity, capability tags, slot usage, and overall health.Remove workers (
jobd_worker_delete) from the broker registry once they are offline/stale.
jobd
A self-hostable, GPU-aware job broker for your own machines — with native MCP/agent integration.
Like task-spooler or pueue, but across all your machines — and VRAM-aware.
You have a couple of boxes with GPUs — a workstation, a server, maybe a laptop — wired together over Tailscale or a LAN. You want to fire off training runs, data pipelines, and long batch jobs from anywhere, have them land on whichever machine actually has the VRAM free, survive across sessions, and get preempted cleanly when something more important shows up. You don't have a cloud, a Kubernetes cluster, or a Slurm install, and you don't want one.
jobd is that missing piece: a lightweight, single-process broker that turns a handful of personal machines into a single queue — and an LLM agent can drive it directly.
# from any machine on your tailnet:
job submit --project myproj --gpu --vram-required 16 --wait -- python train.py
# → routed to whichever worker has ≥16 GB VRAM free, streamed back to your terminalVRAM routing tracks one GPU per host: the worker reports free memory on GPU index 0 only.
Why it exists
Most schedulers assume a datacenter. The lightweight ones that don't (a bare nohup, a tmux session, an ssh-and-pray script) give you nothing: no queue, no VRAM-aware routing, no preemption, no record of what ran where. jobd fills the gap between "ssh in and run it" and "stand up Slurm":
VRAM-fit routing. The broker matches each job against live worker capacity (free VRAM / RAM / CPUs, capability tags, arch/OS) and dispatches to a worker that actually fits — instead of you guessing which box is free. One GPU per host is tracked (GPU index 0).
Preempt + checkpoint window. A higher-priority job can preempt a running one: the worker sends
SIGTERM, the workload gets a grace window, thenSIGKILL. jobd gives each job a per-jobJOBD_CHECKPOINT_DIRto write into during that window; saving the checkpoint is the workload's job. A preempted job ends in the terminalpreemptedstate and is not re-run automatically — to resume, you submit a new job pointed at the old checkpoint. (See docs/preemption.md.)Survives sessions. Submit, close your laptop, check back tomorrow. Jobs live in the broker, not your shell.
Agent-native. Ships a first-class MCP server so an LLM agent (Claude Code, etc.) can submit, monitor, and babysit jobs as tool calls.
Yours. One broker process you run on a machine you own. No accounts, no telemetry, no per-GPU-hour billing; the broker and worker talk only to each other and to your clients. The optional self-update scripts do go online to fetch releases:
scripts/update-worker.sh(whichjob fleet addinstalls on a timer) installs from PyPI, andscripts/deploy-broker.shqueries the GitHub API and pulls the image from GHCR.Loopback by default, tailnet-only beyond it. The broker binds
127.0.0.1unless you setJOBD_HOST. A request from any address that is neither loopback nor in Tailscale's CGNAT range (100.64.0.0/10) gets a403(JOBD_DISABLE_TAILNET_ACL=1turns that check off).
Related MCP server: jungle-grid-mcp-server
Why not just use…?
Tool | What it gives you | Why jobd instead |
| Runs a command on one box | No queue, no VRAM-aware routing, no preemption, no record of what ran where |
A real job queue — on a single machine | jobd queues across all your machines and routes by live VRAM/CPU fit | |
A mature single-machine command queue daemon | Pueue's own README declares distributed execution out of scope — jobd is that missing layer, plus GPU awareness | |
Multi-machine task scheduling with HPC roots, single binary | HQ counts GPUs but doesn't track VRAM, and has no preemption/checkpoint contract or agent interface | |
Slurm | Datacenter-grade scheduling | Heavy to stand up and operate for 2–3 personal boxes; jobd is one process + a poller per host |
SkyPilot / dstack | Provision and run on clouds + your own machines | SkyPilot's "existing machines" mode installs a Kubernetes cluster (k3s) on your boxes; dstack wants Docker + passwordless sudo on every host. jobd is one process + a poller — no containers, no sudo, no K8s |
Modal | Serverless GPU compute on Modal's cloud | Cloud-only: it runs on Modal's machines, not yours |
Ray | A distributed-compute framework; Ray Jobs also runs any shell command on a Ray cluster | You first stand up and run a Ray cluster; jobd is one broker process + a poller per host, with live VRAM-fit routing and a checkpoint window on preemption |
Closest in spirit are Pueue and task-spooler (single-machine by design) and HyperQueue (multi-machine, HPC-shaped). jobd's niche is the 2–5-GPU homelab: multi-machine live VRAM-fit routing (one GPU per host tracked) + preempt/checkpoint window + a native MCP interface — a combination we haven't found in the tools above — with nothing heavier than a Python process per host.
Architecture
flowchart TD
CLI["job CLI"]:::client --> B
MCP["jobd-mcp<br/>MCP tools"]:::client --> B
API["HTTP · SSE"]:::client --> B
B["<b>jobd broker</b> — FastAPI<br/>queue · matcher · priorities · SQLite"]:::broker
B <-->|poll · dispatch| WA["worker A<br/>24 GB GPU"]:::worker
B <-->|poll · dispatch| WB["worker B<br/>8 GB GPU"]:::worker
B <-->|poll · dispatch| WC["worker C<br/>CPU-only"]:::worker
classDef client fill:#1f2937,stroke:#4b5563,color:#e5e7eb;
classDef broker fill:#0e7490,stroke:#155e75,color:#ecfeff;
classDef worker fill:#14532d,stroke:#166534,color:#dcfce7;Workers poll the broker (pull model — no inbound connection to a worker); the broker matches each job against live capacity and hands it back on the poll. One broker process, one poller per host.
Broker — a FastAPI + SQLite service. Holds the queue, runs the matcher, resolves per-project priorities and defaults, exposes a small HTTP API and an SSE stream. Single source of truth.
Workers — lightweight polling agents, one per host. Each advertises live capacity via heartbeat, claims jobs it can run, executes them (
shell=False: the argv you submit is run as-is, not through a shell — exceptjob submit --stdinand the MCPjobd_submittool, which take a command string and run it asbash -c <string>), streams logs back, and honors preemption signals.Clients — the
jobCLI, thejobd-mcpMCP server, or anything that speaks the HTTP API.
Install
pip install jobd # broker + CLI
pip install "jobd[mcp]" # adds the MCP server
pip install "jobd[worker]" # adds the worker daemon (jobd-worker)Requires Python ≥ 3.11. Everything ships in the one jobd package: the broker (jobd), the CLI (job), the MCP server (jobd-mcp), and the worker (jobd-worker). The worker's extra runtime deps (psutil, nvidia-ml-py) live behind the [worker] extra since they're only needed on machines that actually run jobs. scripts/install-worker.sh sets a worker up under ~/jobd-worker with its own venv and a generated config.
Quickstart (single host)
# 1. start the broker (binds 127.0.0.1:8765 by default)
JOBD_ALLOW_NO_AUTH=1 jobd # no-auth is fine for a loopback-only broker
# 2. in another shell, install + start a worker pointed at it
pip install "jobd[worker]"
JOBD_URL=http://127.0.0.1:8765 JOBD_WORKER_HOST=local jobd-worker
# 3. submit a job and wait for it
job submit --project demo --wait -- echo hello
job list
job logs <id>For a real multi-host deployment (Docker broker + systemd workers, Tailscale binding, shared auth token), see docs/security.md and the templates in docker-compose.yml and scripts/. Adding a worker to a running fleet is one command:
job fleet add user@newbox # ssh in, install pinned to the broker's version,
# wire systemd units + the self-update timer,
# verify it registers. `job fleet status` shows drift.Day-2 operations (health, draining a worker, upgrades, token rotation, backups) are in docs/runbook.md.
Supported platforms
Python 3.11+ everywhere.
Component | Linux | macOS | Windows |
Broker ( | ✅ | ☑️ | ☑️ (WSL recommended) |
CLI ( | ✅ | ☑️ | ☑️ |
Worker ( | ✅ full | ⚠️ degraded | untested |
✅ = CI-tested on Linux (Ubuntu), with limits: CI has no GPU runner, so the NVIDIA/VRAM paths are tested only against mocks, and the tests that need a systemd --user scope skip on GitHub's runners. ☑️ = pure-Python and expected to work, but not exercised by CI — please file an issue if something is broken there.
The worker runs its best on Linux with a systemd user instance: memory caps, process reaping, and preemption use systemd-run --user scopes and cgroups. On non-systemd hosts the worker still executes jobs, but silently drops those guarantees — fine for a single trusted box, not for hard resource isolation. GPU features need NVIDIA + nvidia-ml-py. The broker, CLI, and MCP server are pure-Python and portable.
CLI
job submit -p PROJ [--gpu] [--vram-required N] [--needs TAG]... [--count N | --sweep K=v1,v2]... [--wait] -- CMD...
job list [--state STATE] [--project P] [--array A<id>] # queue + recent jobs
job status ID | A<id> [--watch] # one job, or an array's aggregate
job logs ID [-n BYTES] # tail captured output
job wait ID # block until terminal
job cancel ID / job preempt ID # stop a job
job adopt --pid N -p PROJ [--gpu] # register a process you already started (Linux)
job workers # fleet snapshot + health
job projects list | set NAME PRI | nudge NAME DELTA
job audit [--project P] [--since 24h] # event historyjob adopt makes a process started outside jobd (nohup, tmux) visible to the broker: it holds a slot and its VRAM until it exits, and nothing is launched. Its exit code is unknowable, so it ends orphaned, never completed — see docs/adoption.md.
job submit --explain dry-runs the resolution (priority, profile, project defaults, host pin) and prints the effective config without enqueuing anything.
Job arrays
Submit N jobs from one template with --count N. Each member is a normal job — it routes, runs, preempts, and checkpoints independently — and {i} in the command is replaced by the member's 0-based index:
job submit -p train --count 8 -- python train.py --fold {i}
# → Submitted array A42: 8 jobs (ids 42..49)
job list --array A42 # the members, with their index annotations
job status A42 # aggregate: state tally + per-member rollupThe array is identified as A<id> (the first member's job id). job status A42 exits non-zero if any member ended in a non-completed terminal state, so it composes with shell &&.
For a grid search, use --sweep KEY=v1,v2,v3 (repeatable) instead of --count. The broker fans out the cartesian product of all axes, substituting {KEY} per member; {i} (the flat member index) is also available:
job submit -p train --sweep lr=0.1,0.01 --sweep seed=1,2,3 \
-- python train.py --lr {lr} --seed {seed} --out run-{i}
# → Submitted array A50: 6 jobs (ids 50..55) # 2 × 3 = 6 members--sweep and --count are mutually exclusive, the product is capped at 1000 members, and i is reserved as an axis key. Substitution is a literal {key} replace (not str.format), so JSON literals and shell braces in the command pass through untouched.
Coming from pueue or task-spooler?
Most verbs map directly — what changes is that the queue spans every machine you own. The last row is only a rough equivalent: jobd has no named groups with their own parallelism limit.
You ran… | With jobd |
|
|
|
|
|
|
|
|
commands piped to |
|
|
|
| approximately: projects + priorities ( |
What you gain on top: jobs route to whichever machine actually has the VRAM/CPU free, live in the broker rather than in one machine's shell, can be preempted with a checkpoint window instead of killed, and are drivable by an LLM agent over MCP. What you lose: there is no pause/resume, stash, or edit of a queued job (cancel and resubmit instead), and if a worker dies mid-job, a running job not marked idempotent ends orphaned rather than being re-run. A one-machine deployment (broker + one worker on the same host) otherwise behaves like a network-reachable pueue.
MCP / agent integration
jobd ships an MCP server (jobd-mcp) exposing the queue as nine tools — jobd_submit, jobd_status, jobd_logs, jobd_list, jobd_cancel, jobd_preempt, jobd_events, jobd_workers, jobd_worker_delete. docs/agent-cookbook.md is the worked tour: fire-and-babysit polling, surviving preemption with checkpoints, sweeps, and asking the broker why a job won't schedule.
One-liner for Claude Code:
claude mcp add jobd --env JOBD_URL=http://127.0.0.1:8765 --env JOBD_API_TOKEN=<your-token> -- jobd-mcpOr point any other MCP client at it:
{
"mcpServers": {
"jobd": {
"command": "jobd-mcp",
"env": {
"JOBD_URL": "http://127.0.0.1:8765",
"JOBD_API_TOKEN": "<your-token>"
}
}
}
}JOBD_API_TOKEN must match the broker's token, or every call returns 401. Omit it only when the broker runs with JOBD_ALLOW_NO_AUTH=1.
Now an agent can "run this overnight," check on it next session, and route GPU work through the broker instead of colliding on a shared card. The examples/claude-code-hooks/ directory has optional Claude Code hooks that nudge (or hard-block) an agent toward submitting heavy commands through jobd — including a VRAM-aware GPU guard with # NO_GPU / # CONCURRENT_OK / # VRAM=NGB override markers.
Configuration
Three optional YAML files under JOBD_CONFIG_DIR (default /app/config, the Docker image's path — set it when you run from pip). Example files live in this repo's config/ directory; they are not included in the pip package, and docker-compose.yml mounts them into the container:
projects.yaml— per-project base priority and submit defaults (preemptibility, wall/idle timeouts, host pins, capability requirements). Entries may also declareroots:so a job typed with an unregistered run label is priced by the project whose directory it runs in. See docs/projects-yaml.md for the full resolution model and docs/events.md for the event catalog.profiles.yaml— named resource bundles (--profile gpu-train-large) the matcher uses to size a job.classifier.yaml— rules that auto-suggest a profile from the command string.
All three are optional; with none present, every job runs at the global default priority.
Everything else is environment variables — the complete JOBD_* catalog (broker, worker, CLI/MCP, and the vars provided to workloads) lives in docs/configuration.md, and a CI test keeps it in lockstep with the source in both directions.
Concurrency (multislotting)
By default each worker runs one job at a time (JOBD_WORKER_MAX_CONCURRENT_JOBS=1). Raise it to let a worker bin-pack several jobs that fit side by side:
JOBD_WORKER_MAX_CONCURRENT_JOBS=3 jobd-workerThe matcher is resource-aware, so this is not blind N-up oversubscription. Each in-flight job reserves its vram_gb / ram_gb / cpus footprint, and the worker's heartbeat advertises only what's left. For VRAM, NVML's free figure already counts memory a running job has allocated, so only the not-yet-allocated part of the reservations is subtracted: free_vram = nvml_free − max(0, Σ in-flight vram_gb − VRAM held by this worker's jobs), where nvml_free is GPU index 0 only. RAM and CPUs subtract the full reservations. The broker won't place a job that doesn't fit the remaining headroom. The practical payoff: a CPU-only job and a GPU job run at the same time — the CPU job reserves 0 VRAM, so it never blocks the GPU slot, and vice-versa. Two GPU jobs co-run only if both fit live VRAM: before starting a job it was handed, the worker re-reads free VRAM and refuses the job if it no longer fits. A job whose command contains the literal marker # CONCURRENT_OK skips that last check — use it when you know the requested VRAM is overstated.
job workers reports each worker's slot usage — running jobs out of max_concurrent — alongside the live resource ad:
// job workers
{ "host": "desktop", "state": "online", "running": 2, "max_concurrent": 3,
"free_vram_gb": 9.1, "idle_cpus": 6, ... }Set the limit per worker from its environment (systemd unit, shell, or worker.yaml env) — it's a worker-local knob, not a broker setting.
Retention
By default jobd keeps every job record and .log file forever — history is never lost. On a long-running broker, opt into pruning:
JOBD_JOB_RETENTION_DAYS=30 jobd # delete terminal jobs + their logs after 30 daysThe sweeper deletes jobs in a terminal state whose finished_at is older than the horizon, unlinks their per-job .log, and emits a jobs_pruned event. Freed SQLite pages are reused under WAL, so the DB file stays bounded without a global-locking VACUUM. The default (0) keeps everything; pruning old terminal parents is safe for any still-pending dependents.
Security
The broker has no TCP-layer auth beyond a shared bearer token, so it is meant to run on a trusted network (loopback or a Tailscale tailnet), never on a public interface. Three stacked controls:
Interface binding — set
JOBD_HOSTto127.0.0.1(the default) or a Tailscale CGNAT address (100.64.0.0/10), not0.0.0.0. The broker itself does not enforce this: it starts on any bind address, and only warns (or refuses) when a non-loopback bind is combined with no-auth. The only check on the bind value is a CI lint (tests/test_deploy_lint.py) on the shipped Docker deployment; control 2 is what holds at runtime.Source-IP check (runtime) — whatever the bind, the broker answers
403to any request whose source address is neither loopback nor in100.64.0.0/10(src/jobd/auth.py).JOBD_DISABLE_TAILNET_ACL=1turns this off.Bearer token — set
JOBD_API_TOKEN(≥32 random bytes) on every broker/worker/CLI/MCP host. The broker refuses to start without it unless you explicitly setJOBD_ALLOW_NO_AUTH=1.JOBD_ALLOW_NO_AUTH=1is for a loopback-only broker (JOBD_HOST=127.0.0.1) — for local dev/tests. Combined with a non-loopbackJOBD_HOSTit exposes an unauthenticated RCE endpoint to your whole tailnet; the broker logs a startup warning if you do this. Don't.
Three endpoints are exempt from the source-IP check and the token — /livez, /readyz and /metrics answer with no bearer token and no source-IP check, because a generic HTTP monitor cannot send a token. /metrics is the one that matters: it publishes the broker version, job counts by state, and every worker's hostname and version. No commands, cwd, env or project names — but it does fingerprint the fleet. That is why, for these three, the JOBD_HOST bind above is load-bearing rather than defence-in-depth: port-forward the broker and you publish that inventory. Full table: Unauthenticated surface.
Full threat model, env-var reference, and token rotation: docs/security.md.
License
MIT — see LICENSE.
Available Tools
9 toolsjobd_cancelB
Cancel a job (queued → cancelled; running → SIGTERM via worker signal poll, ~2s).
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | ||
| reason | No | Why — recorded on the broker's job_cancelled event (see jobd_events). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral transparency burden on its own. It discloses meaningful state-dependent behavior: queued jobs are cancelled, and running jobs receive SIGTERM via a worker signal poll in about 2 seconds. It does not cover idempotency, auth, or failure conditions, but the core cancellation semantics are clearly exposed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single dense sentence that front-loads the primary action and immediately clarifies state-dependent behavior. There is no filler, and the parenthetical adds crucial implementation detail without bloating the text.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
The description is adequate for a simple cancellation operation, but gaps remain. It does not explain the response/acknowledgement behavior, what happens if the job is already in a terminal state, or how cancellation relates to the sibling jobd_preempt. Given that there is no output schema to compensate, these missing contextual details keep it below a 4.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 50%, and the description adds little parameter-level meaning. 'Cancel a job' implicitly identifies job_id as the target job, but the description does not elaborate on job_id or reason. The reason parameter is already described in the schema, so the description itself contributes minimal semantic value beyond the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb and resource: 'Cancel a job'. It further clarifies the two main job states and what happens to each, which is more informative than a generic 'cancel job'. However, it does not explicitly distinguish itself from the sibling jobd_preempt, so it does not earn a 5.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives no explicit guidance on when to use this tool versus alternatives such as jobd_preempt or jobd_submit. It implies usage through its name and purpose, but the agent is left to infer the appropriate context, and no exclusions or alternative routing are provided.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jobd_eventsA
The broker's event stream — the surface that explains WHY, not just what. /jobs says a job is queued; only this says it has been skipped 400 times because no worker advertises cuda-32gb, or that its dependency was cancelled, or that a watchdog killed it. Filter by since (2h/3d/1w), event type, job_id, project, or source (broker|worker). Use when a job is not doing what you expect and jobd_status alone does not explain it.
| Name | Required | Description | Default |
|---|---|---|---|
| event | No | Filter to one event type. Known types: admission_blocked, auto_preempt, checkpoint_complete, cwd_identity_applied, cwd_refused, cwd_route_warning, dispatch_skip, env_scrubbed, gpu_contention_warning, job_adopted, job_cancelled, job_completed, job_dispatched, job_orphaned, job_resurrected, job_retry_scheduled, job_started, job_submitted, job_uncancelled, jobs_pruned, logs_pruned, preflight_warning, reclaim_suppressed, scheduling_timeout, serialization_warning, stale_scope_sweep, submit_warning, sweep_warning, unknown_project, version_drift, watchdog_fired, worker_offline, worker_registered, worker_shutdown, worker_stale. Hook-ingested events may carry custom names beyond these. | |
| limit | No | Max rows, newest-last. Broker clamps to 10000. | |
| since | No | Relative window (2h, 3d, 1w) or an ISO-8601 timestamp. Default: all retained. | |
| job_id | No | Only events for this job. | |
| source | No | Which side emitted the event. | |
| project | No | Only events for this project. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
There are no annotations, so the description carries the full burden of behavioral disclosure. It does explain the semantic role of the stream and its diagnostic value, but it never explicitly states that this is a read-only operation, nor does it describe response shape, ordering, retention, or error behavior. The schema supplies some behavioral details (limit default and broker clamp), so this is not a complete gap, but the description alone is only minimally transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is front-loaded with a strong purpose statement and every sentence contributes something: reason to use it, concrete illustrative examples, filter options, and a selection rule. It is slightly more expansive than necessary, but not bloated or repetitive.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a filtered event stream with six optional parameters, the description covers the key selection context and when to use it, while the schema fills in defaults, enums, and the known event-type list. The lack of annotations and an output schema means returned payload shape is not fully specified, and the source mismatch is an inconsistency, but the tool is still reasonably invokable and well contextualized.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline for parameter semantics is 3. However, the description's filter summary says source (broker|worker, while the schema enumerates four sources: broker, worker, hook, and mcp. This narrowing contradicts the schema and could lead an agent to ignore hook or mcp sources; it also omits limit from the summary. The description adds conceptual examples but actually detracts from parameter clarity.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description names a specific resource—the broker's event stream—and states its distinguishing purpose: explaining WHY, not just what. It differentiates itself from /jobs and jobd_status with concrete examples (skipped 400 times, dependency cancelled, watchdog killed), so an agent can tell it apart from sibling tools without opening the schema.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description gives an explicit when-to-use rule: 'Use when a job is not doing what you expect and jobd_status alone does not explain it.' It also contrasts with /jobs and mentions the alternative surface (jobd_status), providing an agent with a clear decision condition.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jobd_listA
List jobs on the broker with per-state counts. Defaults to the active set (queued/assigned/running); filter by state (e.g. ['failed']) or project to find past runs. Each row is a compact summary: job_id, project, state, host, exit_code, queued_at, started_at — call jobd_status for a job's full record. Use to answer 'what is running / queued right now?' or to locate a job id you've lost.
| Name | Required | Description | Default |
|---|---|---|---|
| limit | No | Max jobs returned (newest first; clamped to [1,200]). `counts` still covers every job matching the filters; a `truncated` field reports how many were cut. | |
| state | No | States to include — any of: queued, assigned, running, completed, failed, cancelled, preempted, orphaned, scheduling_timeout. Omit for the active set (queued/assigned/running); pass [] for all states. | |
| project | No | Restrict to one project's jobs (the --project value used at submit). |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the behavioral burden. It discloses that the tool defaults to the active set, returns per-state counts, and emits compact summary rows with specified fields. It also clarifies that jobd_status is needed for full records. It doesn't explicitly state that the operation is read-only or side-effect-free, but the verb 'list' implies it, and the behavior described is accurate and useful.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three sentences, each earning its place: purpose plus default scope, filtering guidance, and output format with a pointer to jobd_status. The most important operational details are front-loaded, with no filler or repetition of schema content.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
No output schema exists, so the description compensates by enumerating the row fields and indicating that counts accompany the list. It also explains defaults, filter options, and how to get full records. This is complete for a list-style tool of moderate complexity, and the schema handles the remaining input details.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the baseline is 3. The description adds a little usage context ('filter by state (e.g. ["failed"]) or project to find past runs') but does not materially extend what the schema already documents. It earns the baseline but no more.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with a specific verb and resource: 'List jobs on the broker with per-state counts.' It clearly distinguishes itself from sibling jobd_status by stating that each row is a compact summary and that jobd_status provides the full record. An agent can immediately tell what this tool does and how it differs from related tools.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It states when to use the tool: answer 'what is running / queued right now?' or locate a lost job id, and it explains the default active set and how to filter by state or project for past runs. It doesn't explicitly say 'use this instead of X' beyond pointing to jobd_status for full records, but the use case is clearly scoped.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jobd_logsA
Tail the captured stdout/stderr of a job (workers stream output to the broker's per-job log as it runs). Returns log_tail (last tail_bytes, default 8 KiB, max 1 MiB) plus size_bytes/returned_bytes/truncated — works for running AND finished jobs. Use to check progress mid-run, diagnose a failure's traceback, or grab a job's final output.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes | Numeric job id whose captured output to read. | |
| tail_bytes | No | How many bytes from the END of the log to return (server caps reads at 1 MiB). Raise for context, lower for a quick liveness peek. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the full burden. It discloses stream mechanics ('workers stream output to the broker's per-job log as it runs'), lifecycle support ('running AND finished'), and exact return semantics including tail_bytes default/max and truncation metadata. It does not explicitly state permissions, rate limits, or side-effect absence, but those are less critical for a tail/read operation.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Three dense sentences front-load the core action, follow with return format and lifecycle nuance, and end with concrete use cases. No filler, no repetition of schema fields that isn't functional.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite lacking an output schema, the description fully specifies the return shape (log_tail plus size_bytes/returned_bytes/truncated), the parameter semantics, and the supported job states. For a 2-parameter read tool, nothing essential is missing for an agent to select and invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds practical guidance beyond the schema: 'Raise for context, lower for a quick liveness peek' for tail_bytes, and it reiterates the default and max. This is modest but genuine added semantic value.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Specific verb+resource: 'Tail the captured stdout/stderr of a job.' The description clearly distinguishes from siblings like jobd_status by focusing on output logs, and clarifies it returns a tail of the log. It names the return payload and available metadata, making the tool's function unambiguous.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit use cases are given: 'check progress mid-run, diagnose a failure's traceback, or grab a job's final output.' It also notes the tool works for 'running AND finished jobs,' which helps an agent decide when to call it. It does not name alternatives or when not to use it, so it stops shy of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jobd_preemptA
Preempt a running/assigned preemptible job (worker SIGTERMs with grace; final state 'preempted'). Refused if not preemptible or not running.
| Name | Required | Description | Default |
|---|---|---|---|
| job_id | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, this description carries the behavioral burden and does well: it discloses the worker side effect (SIGTERM with grace), the terminal state ('preempted'), and refusal conditions. There is a minor ambiguity between 'assigned' being preemptable and 'refused if not running', but overall the key behaviors are transparent.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The entire description is one dense sentence with no filler. It front-loads the core action and immediately follows with behavioral consequences and refusal criteria; every phrase earns its place.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter tool with no output schema, the description supplies the essential invocation context: what happens, what the final state is, and when the call fails. It does not describe the return payload, but given no output schema and simple side effects, the description is reasonably complete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 0%, but the sole parameter job_id is a simple integer whose meaning is obvious from the description. The description adds semantic constraints on which job_ids are valid (preemptible, running/assigned), which is useful context beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Preempt') and a precise resource ('a running/assigned preemptible job'), and it names the distinguishing final state ('preempted'). It also implicitly differentiates from jobd_cancel by restricting preemption to preemptible jobs, so an agent can tell which sibling applies.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It clearly states when the tool is valid ('running/assigned preemptible job') and when it will be refused ('if not preemptible or not running'). It does not explicitly name an alternative such as jobd_cancel for non-preemptible jobs, which keeps it just short of a 5.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jobd_statusA
The full JobInfo record for one job_id: state, exit_code, timings, host, and the scheduling internals — depends_on + cascade policy (depends_on_any_exit), pending cancel/preempt signal, resolved profile, requires (gpu/tags/idempotent), host pin, fast_path, timeouts, termination_reason. Use for a quick state check AND for debugging why a job routed/failed/stalled. Pass wait=true to block until terminal or wait_timeout_s.
| Name | Required | Description | Default |
|---|---|---|---|
| wait | No | ||
| job_id | Yes | ||
| wait_timeout_s | No | Server clamps to 270. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations present, the description carries the behavioral disclosure burden. It discloses the full return surface, the wait=true blocking behavior, and the timeout parameter, which is substantial context. It does not mention error behavior or side effects, but for a status tool this is reasonably complete.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is dense but every clause earns its place: return contents first, use cases second, wait behavior last. It avoids fluff and is well-structured for scanning.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the lack of an output schema, the description compensates by enumerating the returned fields and giving practical use cases. It could additionally explain how this tool differs from jobd_events or jobd_logs, and clarify whether a timeout returns partial data or errors, but it is largely sufficient.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 33%, but the description explains wait semantics directly ('block until terminal or wait_timeout_s') and makes job_id's role implicitly clear. The timeout clamp is already in the schema. This adds meaningful meaning beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description precisely identifies the resource ('full JobInfo record for one job_id') and enumerates the returned fields, from state and exit_code to scheduling internals. It also distinguishes itself from sibling tools like jobd_list by emphasizing a single job and deep debugging detail.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description clearly states when to use the tool: for a quick state check and for debugging why a job routed, failed, or stalled. It does not explicitly name alternatives or exclusions, so it stops short of full guidance, but the usage context is unambiguous.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jobd_submitA
Submit a job to the jobd broker. Default async; pass wait=true to block up to wait_timeout_s (server clamps to 270).
| Name | Required | Description | Default |
|---|---|---|---|
| cwd | Yes | Absolute path; broker validates against worker mount_roots. | |
| gpu | No | true = pin to a GPU-capable worker. false or omitted = no GPU preference (any worker, GPU or not) — the same as leaving off the CLI's --gpu. Forbidding GPU workers (CLI --no-gpu) is not offered here. | |
| host | No | Host alias pin (laptop, desktop-vm). | |
| wait | No | Sync mode: block until terminal or timeout. For an array submit (count/sweep), waits on every member under one shared deadline and returns an aggregate {array_id, count, job_ids, states, all_completed, members:[{job_id, state, exit_code}]}. | |
| extra | No | Escape hatch: idempotent (bool), depends_on (int[]), depends_on_any_exit (bool), priority (int delta), max_wall_s (int), idle_timeout_s (int), max_retries (int, default 0: re-run up to N times on a plain non-zero exit), retry_delay_s (int), scheduling_timeout_s (int 1..604800 — give up and terminate the job as 'scheduling_timeout' if it is still QUEUED after N seconds; omit to wait indefinitely for a capable worker), checkpoint_grace_s (int 1..300), vram_gb (float — explicit GPU VRAM the job needs at dispatch; falls back to cuda-Ngb tier-tag max, then to 2 GB floor for --gpu jobs), count (int 1..1000 — submit a job array of N members, with `{i}` in the command replaced by the 0-based index; response is {array_id, count, job_ids, warnings} instead of a single job), sweep (list of {key, values[]} — parameter-sweep axes; broker fans out the cartesian product, substituting `{key}` per member plus `{i}`; mutually exclusive with count; product capped at 1000), profile (str), env (dict), preemptible (bool), session_id (str), arch (str — pin to a worker CPU arch), os (str — pin to a worker OS). | |
| needs | No | Tool tags (R, python3, cuda). | |
| command | Yes | Shell command run by the worker shell. | |
| dry_run | No | Preview mode: run full validation + routing decision (profile, project defaults, cwd, depends_on, preflight, gpu_contention) and return the would-be plan WITHOUT queueing. Response has state='dry-run', would_route_to (list[host]), would_use_worker (host or null), validation (resolved fields + warnings). Per dry-run convention 2026-05-18. | |
| project | Yes | Scheduling identity. A registered projects.yaml name (matched case- and -/_-insensitively) prices at its priority; an unregistered name is priced by the project whose roots: contain cwd, else by _default. The result's project_label carries the name as typed when the two differ. | |
| wait_timeout_s | No | Seconds; permissive — server clamps to 270. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries the burden, and it does disclose key behavioral traits: submission is asynchronous by default, wait=true switches to blocking, and wait_timeout_s is clamped to 270. It does not describe success responses or error conventions, though the side effect of submission is clear from the verb.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences with no filler: the primary behavior is front-loaded, and the timeout caveat is stated compactly. Nothing in the description repeats schema text unnecessarily.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 10-parameter submission tool with no output schema, the description is relatively thin: it explains sync versus async but never states what a successful response contains (e.g., job_id or array_id), which the agent needs to hand off to jobd_status or jobd_logs. The rich parameter schema compensates for parameter details, but response expectations are left to inference.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the parameter baseline is 3. The description adds a useful restatement of wait/wait_timeout_s semantics, but the remaining parameters are already fully documented in the schema, so no further compensation is needed.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a direct verb–resource pair ('Submit a job to the jobd broker') that clearly states the operation. It is set apart from the sibling tools, which are all post-submission queries or mutations such as jobd_status, jobd_logs, and jobd_cancel.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It gives an explicit usage default ('Default async') and the exact condition for the blocking alternative ('pass wait=true'), so an agent knows how to invoke it. It does not explicitly name sibling tools as alternatives, but the sibling list is dominated by post-submission operations, making the correct context fairly clear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jobd_worker_deleteA
Remove a worker from the broker registry. The broker refuses (409) if the worker is still online — caller stops the worker process or waits for the heartbeat sweeper first.
| Name | Required | Description | Default |
|---|---|---|---|
| host | Yes | Worker host identifier. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals the 409 conflict behavior, the online/offline prerequisite, and the need to stop the worker or wait for the sweeper. This is substantive context beyond what the schema alone provides.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Two sentences, each earning its place: the first states the action, the second explains the key precondition and failure mode. No filler or repetition.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a single-parameter, no-output-schema tool, the description covers the action, prerequisite, and error condition. It could optionally mention the success response shape or idempotency, but nothing essential to invoking it correctly is missing.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so the parameter 'host' is already documented as 'Worker host identifier.' The tool description adds no further meaning about the host parameter, so a baseline score of 3 is appropriate.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description uses a specific verb ('Remove') with a specific resource ('worker from the broker registry'), making the action unmistakable. It clearly differentiates from siblings like jobd_workers or jobd_submit without needing to compare schemas.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides clear context: the worker must not be online, and the broker will return 409 otherwise. It also tells the caller how to proceed ('stops the worker process or waits for the heartbeat sweeper first'), though it does not explicitly mention alternative tools for inspecting workers before deletion.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
jobd_workersA
Fleet snapshot: every registered worker with state (online/stale/offline), live capacity ad (free_vram_gb, unregistered_vram_gb, free_ram_gb, idle_cpus), capability tags (cuda tiers, arch/os), slot usage (running/max_concurrent), and last_heartbeat — plus an overall health rollup (healthy|degraded|empty). Use before submitting GPU work to see what's free, or to diagnose why a job isn't being dispatched.
| Name | Required | Description | Default |
|---|---|---|---|
No parameters | |||
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description handles transparency itself by labeling it a 'snapshot' (read-only point-in-time view) and naming all returned categories. It doesn't cover auth, staleness, or error behavior, but for a zero-parameter fleet status read the essential behavior is disclosed.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first sentence front-loads 'Fleet snapshot' and then packs the output contract into compact parenthetical lists; the second sentence adds concrete usage guidance. Every clause earns its place and there is no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Despite lacking an output schema, the description enumerates the output fields, their possible values, and the health rollup states, and it supplies two relevant use cases. For a parameterless read-only fleet query, this is sufficient for an agent to invoke it correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The tool takes zero parameters and schema coverage is 100%, so there are no parameter semantics to add. The description's field list primarily documents return content rather than inputs, which fits the baseline for parameterless tools.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with 'Fleet snapshot' and enumerates exactly what it returns: worker state, capacity fields, capability tags, slot usage, heartbeat, and health rollup. This makes the tool's resource and output clear, though it does not explicitly differentiate it from sibling tools like jobd_status or jobd_list.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It provides explicit use cases: use before submitting GPU work to check available capacity, or to diagnose why jobs aren't dispatched. It doesn't spell out when not to use it or name alternative tools, but the context is strong.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
9 tool updates
v0.5.47- Changed
jobd_cancel2 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / reason / descriptionAdded value: +"Why — recorded on the broker's job_cancelled event (see jobd_events)."
- Changed
jobd_events2 fields changed- added
Input schema / additionalPropertiesAdded value: +false - changed
Input schema / properties / event / descriptionPrevious value: -"Filter to one event type. Known types: admission_blocked, auto_preempt, checkpoint_complete, cwd_identity_applied, cwd_refused, cwd_route_warning, dispatch_skip, env_scrubbed, gpu_contention_warning, job_cancelled, job_completed, job_dispatched, job_orphaned, job_resurrected, job_started, job_submitted, job_uncancelled, jobs_pruned, logs_pruned, preflight_warning, reclaim_suppressed, scheduling_timeout, serialization_warning, stale_scope_sweep, submit_warning, sweep_warning, unknown_project, version_drift, watchdog_fired, worker_offline, worker_registered, worker_shutdown, worker_stale. Hook-ingested events may carry custom names beyond these."New value: +"Filter to one event type. Known types: admission_blocked, auto_preempt, checkpoint_complete, cwd_identity_applied, cwd_refused, cwd_route_warning, dispatch_skip, env_scrubbed, gpu_contention_warning, job_adopted, job_cancelled, job_completed, job_dispatched, job_orphaned, job_resurrected, job_retry_scheduled, job_started, job_submitted, job_uncancelled, jobs_pruned, logs_pruned, preflight_warning, reclaim_suppressed, scheduling_timeout, serialization_warning, stale_scope_sweep, submit_warning, sweep_warning, unknown_project, version_drift, watchdog_fired, worker_offline, worker_registered, worker_shutdown, worker_stale. Hook-ingested events may carry custom names beyond these."
- Changed
jobd_list2 fields changed- added
Input schema / additionalPropertiesAdded value: +false - added
Input schema / properties / state / items / enumAdded value: +[ + "queued", + "assigned", + "running", + "completed", + "failed", + "cancelled", + "preempted", + "orphaned", + "scheduling_timeout" +]
- Changed
jobd_logs1 field changed- added
Input schema / additionalPropertiesAdded value: +false
- Changed
jobd_preempt1 field changed- added
Input schema / additionalPropertiesAdded value: +false
- Changed
jobd_status1 field changed- added
Input schema / additionalPropertiesAdded value: +false
- Changed
jobd_submit6 fields changed- added
Input schema / $defsAdded value: +{ + "JobRequires": { + "additionalProperties": false, + "properties": { + "arch": { + "default": "any", + "title": "Arch", + "type": "string" + }, + "gpu": { + "anyOf": [ + { + "type": "boolean" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Gpu" + }, + "idempotent": { + "default": false, + "title": "Idempotent", + "type": "boolean" + }, + "needs": { + "items": { + "type": "string" + }, + "title": "Needs", + "type": "array" + }, + "os": { + "default": "any", + "title": "Os", + "type": "string" + } + }, + "title": "JobRequires", + "type": "object" + }, + "SweepAxis": { + "additionalProperties": false, + "description": "One named axis of a parameter sweep: a key and the values it ranges over.\n\nThe broker takes the cartesian product of all axes to fan out array members;\neach member substitutes `{key}` → its value in the command and env. `{i}`\n(the flat member index) is always available alongside the named keys, so\n`i` is reserved and rejected as an axis key. See jobd.arrays.", + "properties": { + "key": { + "minLength": 1, + "title": "Key", + "type": "string" + }, + "values": { + "items": { + "type": "string" + }, + "minItems": 1, + "title": "Values", + "type": "array" + } + }, + "required": [ + "key", + "values" + ], + "title": "SweepAxis", + "type": "object" + } +} - added
Input schema / additionalPropertiesAdded value: +false - changed
Input schema / properties / extra / additionalPropertiesPrevious value: -trueNew value: +false - changed
Input schema / properties / extra / descriptionPrevious value: -"Escape hatch: idempotent (bool), depends_on (int[]), depends_on_any_exit (bool), priority (int delta), max_wall_s (int), idle_timeout_s (int), scheduling_timeout_s (int 1..604800 — give up and terminate the job as 'scheduling_timeout' if it is still QUEUED after N seconds; omit to wait indefinitely for a capable worker), checkpoint_grace_s (int 1..300), vram_gb (float — explicit GPU VRAM the job needs at dispatch; falls back to cuda-Ngb tier-tag max, then to 2 GB floor for --gpu jobs), count (int 1..1000 — submit a job array of N members, with `{i}` in the command replaced by the 0-based index; response is {array_id, count, job_ids, warnings} instead of a single job), sweep (list of {key, values[]} — parameter-sweep axes; broker fans out the cartesian product, substituting `{key}` per member plus `{i}`; mutually exclusive with count; product capped at 1000), profile (str), env (dict), preemptible (bool), session_id (str), arch (str — pin to a worker CPU arch), os (str — pin to a worker OS)."New value: +"Escape hatch: idempotent (bool), depends_on (int[]), depends_on_any_exit (bool), priority (int delta), max_wall_s (int), idle_timeout_s (int), max_retries (int, default 0: re-run up to N times on a plain non-zero exit), retry_delay_s (int), scheduling_timeout_s (int 1..604800 — give up and terminate the job as 'scheduling_timeout' if it is still QUEUED after N seconds; omit to wait indefinitely for a capable worker), checkpoint_grace_s (int 1..300), vram_gb (float — explicit GPU VRAM the job needs at dispatch; falls back to cuda-Ngb tier-tag max, then to 2 GB floor for --gpu jobs), count (int 1..1000 — submit a job array of N members, with `{i}` in the command replaced by the 0-based index; response is {array_id, count, job_ids, warnings} instead of a single job), sweep (list of {key, values[]} — parameter-sweep axes; broker fans out the cartesian product, substituting `{key}` per member plus `{i}`; mutually exclusive with count; product capped at 1000), profile (str), env (dict), preemptible (bool), session_id (str), arch (str — pin to a worker CPU arch), os (str — pin to a worker OS)." - added
Input schema / properties / extra / propertiesAdded value: +{ + "arch": { + "default": "any", + "title": "Arch", + "type": "string" + }, + "checkpoint_grace_s": { + "anyOf": [ + { + "maximum": 300, + "minimum": 1, + "type": "integer" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Checkpoint Grace S" + }, + "count": { + "default": 1, + "maximum": 1000, + "minimum": 1, + "title": "Count", + "type": "integer" + }, + "depends_on": { + "items": { + "type": "integer" + }, + "title": "Depends On", + "type": "array" + }, + "depends_on_any_exit": { + "default": false, + "title": "Depends On Any Exit", + "type": "boolean" + }, + "env": { + "additionalProperties": { + "type": "string" + }, + "title": "Env", + "type": "object" + }, + "idempotent": { + "default": false, + "title": "Idempotent", + "type": "boolean" + }, + "idle_timeout_s": { + "anyOf": [ + { + "maximum": 86400, + "minimum": 1, + "type": "integer" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Idle Timeout S" + }, + "max_retries": { + "default": 0, + "maximum": 20, + "minimum": 0, + "title": "Max Retries", + "type": "integer" + }, + "max_wall_s": { + "anyOf": [ + { + "maximum": 604800, + "minimum": 1, + "type": "integer" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Max Wall S" + }, + "os": { + "default": "any", + "title": "Os", + "type": "string" + }, + "preemptible": { + "anyOf": [ + { + "type": "boolean" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Preemptible" + }, + "priority": { + "default": 0, + "title": "Priority Delta", + "type": "integer" + }, + "profile": { + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Profile" + }, + "retry_delay_s": { + "default": 0, + "maximum": 86400, + "minimum": 0, + "title": "Retry Delay S", + "type": "integer" + }, + "scheduling_timeout_s": { + "anyOf": [ + { + "maximum": 604800, + "minimum": 1, + "type": "integer" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Scheduling Timeout S" + }, + "session_id": { + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Session Id" + }, + "sweep": { + "items": { + "$ref": "#/$defs/SweepAxis" + }, + "title": "Sweep", + "type": "array" + }, + "vram_gb": { + "default": 0, + "minimum": 0, + "title": "Vram Gb", + "type": "number" + } +} - changed
Input schema / properties / gpu / descriptionPrevious value: -"Pin to GPU-capable worker."New value: +"true = pin to a GPU-capable worker. false or omitted = no GPU preference (any worker, GPU or not) — the same as leaving off the CLI's --gpu. Forbidding GPU workers (CLI --no-gpu) is not offered here."
- Changed
jobd_worker_delete1 field changed- added
Input schema / additionalPropertiesAdded value: +false
- Changed
jobd_workers1 field changed- added
Input schema / additionalPropertiesAdded value: +false
2 tool updates
v0.5.44- Changed
jobd_events1 field changed- changed
Input schema / properties / event / descriptionPrevious value: -"Filter to one event type. Known types: admission_blocked, auto_preempt, checkpoint_complete, cwd_refused, dispatch_skip, env_scrubbed, job_cancelled, job_completed, job_dispatched, job_orphaned, job_resurrected, job_started, job_submitted, job_uncancelled, jobs_pruned, logs_pruned, reclaim_suppressed, scheduling_timeout, stale_scope_sweep, submit_warning, sweep_warning, version_drift, watchdog_fired, worker_offline, worker_registered, worker_shutdown, worker_stale. Hook-ingested events may carry custom names beyond these."New value: +"Filter to one event type. Known types: admission_blocked, auto_preempt, checkpoint_complete, cwd_identity_applied, cwd_refused, cwd_route_warning, dispatch_skip, env_scrubbed, gpu_contention_warning, job_cancelled, job_completed, job_dispatched, job_orphaned, job_resurrected, job_started, job_submitted, job_uncancelled, jobs_pruned, logs_pruned, preflight_warning, reclaim_suppressed, scheduling_timeout, serialization_warning, stale_scope_sweep, submit_warning, sweep_warning, unknown_project, version_drift, watchdog_fired, worker_offline, worker_registered, worker_shutdown, worker_stale. Hook-ingested events may carry custom names beyond these."
- Changed
jobd_submit1 field changed- changed
Input schema / properties / project / descriptionPrevious value: -"Priority lookup key; falls back to _default."New value: +"Scheduling identity. A registered projects.yaml name (matched case- and -/_-insensitively) prices at its priority; an unregistered name is priced by the project whose roots: contain cwd, else by _default. The result's project_label carries the name as typed when the two differ."
5 tool updates
v0.5.36- Added
jobd_events - Added
jobd_list - Added
jobd_preempt - Added
jobd_worker_delete - Added
jobd_workers
6 tool updates
v0.5.35- Removed
jobd_list - Added
jobd_logs - Removed
jobd_preempt - Added
jobd_status - Added
jobd_submit - Removed
jobd_worker_delete
5 tool updates
v0.5.34- Removed
jobd_events - Removed
jobd_logs - Removed
jobd_status - Removed
jobd_submit - Removed
jobd_workers
1 tool update
v0.5.31- Changed
jobd_events1 field changed- changed
Input schema / properties / event / descriptionPrevious value: -"Filter to one event type. Known types: admission_blocked, auto_preempt, checkpoint_complete, cwd_refused, dispatch_skip, job_cancelled, job_completed, job_dispatched, job_orphaned, job_resurrected, job_started, job_submitted, job_uncancelled, jobs_pruned, logs_pruned, scheduling_timeout, stale_scope_sweep, submit_warning, sweep_warning, watchdog_fired, worker_offline, worker_registered, worker_shutdown, worker_stale. Hook-ingested events may carry custom names beyond these."New value: +"Filter to one event type. Known types: admission_blocked, auto_preempt, checkpoint_complete, cwd_refused, dispatch_skip, env_scrubbed, job_cancelled, job_completed, job_dispatched, job_orphaned, job_resurrected, job_started, job_submitted, job_uncancelled, jobs_pruned, logs_pruned, scheduling_timeout, stale_scope_sweep, submit_warning, sweep_warning, version_drift, watchdog_fired, worker_offline, worker_registered, worker_shutdown, worker_stale. Hook-ingested events may carry custom names beyond these."
3 tool updates
v0.5.26- Added
jobd_events - Removed
jobd_job_get - Changed
jobd_submit1 field changed- changed
Input schema / properties / extra / descriptionPrevious value: -"Escape hatch: idempotent (bool), depends_on (int[]), depends_on_any_exit (bool), priority (int delta), max_wall_s (int), idle_timeout_s (int), checkpoint_grace_s (int 1..300), vram_gb (float — explicit GPU VRAM the job needs at dispatch; falls back to cuda-Ngb tier-tag max, then to 2 GB floor for --gpu jobs), count (int 1..1000 — submit a job array of N members, with `{i}` in the command replaced by the 0-based index; response is {array_id, count, job_ids, warnings} instead of a single job), sweep (list of {key, values[]} — parameter-sweep axes; broker fans out the cartesian product, substituting `{key}` per member plus `{i}`; mutually exclusive with count; product capped at 1000), profile (str), env (dict), preemptible (bool)."New value: +"Escape hatch: idempotent (bool), depends_on (int[]), depends_on_any_exit (bool), priority (int delta), max_wall_s (int), idle_timeout_s (int), scheduling_timeout_s (int 1..604800 — give up and terminate the job as 'scheduling_timeout' if it is still QUEUED after N seconds; omit to wait indefinitely for a capable worker), checkpoint_grace_s (int 1..300), vram_gb (float — explicit GPU VRAM the job needs at dispatch; falls back to cuda-Ngb tier-tag max, then to 2 GB floor for --gpu jobs), count (int 1..1000 — submit a job array of N members, with `{i}` in the command replaced by the 0-based index; response is {array_id, count, job_ids, warnings} instead of a single job), sweep (list of {key, values[]} — parameter-sweep axes; broker fans out the cartesian product, substituting `{key}` per member plus `{i}`; mutually exclusive with count; product capped at 1000), profile (str), env (dict), preemptible (bool), session_id (str), arch (str — pin to a worker CPU arch), os (str — pin to a worker OS)."
1 tool update
v0.5.12- Changed
jobd_list2 fields changed- changed
Input schema / properties / limit / descriptionPrevious value: -"Advisory cap on returned jobs (the broker currently returns its default window)."New value: +"Max jobs returned (newest first; clamped to [1,200]). `counts` still covers every job matching the filters; a `truncated` field reports how many were cut." - changed
Input schema / properties / state / descriptionPrevious value: -"States to include — any of: queued, assigned, running, completed, failed, cancelled, preempted, orphaned, scheduling_timeout. Currently only the first is forwarded to the broker (single state_filter)."New value: +"States to include — any of: queued, assigned, running, completed, failed, cancelled, preempted, orphaned, scheduling_timeout. Omit for the active set (queued/assigned/running); pass [] for all states."
3 tool updates
- Changed
jobd_job_get1 field changed- added
Input schema / properties / job_id / descriptionAdded value: +"Numeric job id as returned by jobd_submit or shown in jobd_list."
- Changed
jobd_list3 fields changed- added
Input schema / properties / limit / descriptionAdded value: +"Advisory cap on returned jobs (the broker currently returns its default window)." - added
Input schema / properties / project / descriptionAdded value: +"Restrict to one project's jobs (the --project value used at submit)." - changed
Input schema / properties / state / descriptionPrevious value: -"States to include. Currently only the first is forwarded to the broker (single state_filter)."New value: +"States to include — any of: queued, assigned, running, completed, failed, cancelled, preempted, orphaned, scheduling_timeout. Currently only the first is forwarded to the broker (single state_filter)."
- Changed
jobd_logs2 fields changed- added
Input schema / properties / job_id / descriptionAdded value: +"Numeric job id whose captured output to read." - added
Input schema / properties / tail_bytes / descriptionAdded value: +"How many bytes from the END of the log to return (server caps reads at 1 MiB). Raise for context, lower for a quick liveness peek."
9 tool updates
v0.5.5- First observed
jobd_cancel - First observed
jobd_job_get - First observed
jobd_list - First observed
jobd_logs - First observed
jobd_preempt - First observed
jobd_status - First observed
jobd_submit - First observed
jobd_worker_delete - First observed
jobd_workers
TDQS
Scored across 9 tools
Each tool targets a distinct operation: status/logs/cancel/preempt/list/events/workers/worker_delete/submit. There is no overlap between them; even status vs events are clearly differentiated (state snapshot vs why-events).
All tools follow a consistent jobd_<verb> or jobd_<noun>_<verb> pattern (jobd_status, jobd_logs, jobd_cancel, jobd_preempt, jobd_list, jobd_events, jobd_workers, jobd_worker_delete, jobd_submit). The pattern is uniform and predictable.
9 tools is well-scoped for a job scheduling/broker domain: submit, inspect (status/logs/list/events), control (cancel/preempt), and fleet management (workers/worker_delete). Each tool earns its place.
The surface covers the full job lifecycle: submit, status, logs, cancel, preempt, list, events, and worker management. Minor gaps like resubmit or job deletion are absent, but the core scheduling workflow is complete.
Maintenance
Related MCP Connectors
On-demand GPU nodes for agents: create nodes, run commands, and submit jobs, billed by the minute.
HiveCompute MCP Server — decentralized inference router for AI agents
MCP-first toolbox for agents: KV storage, auth, queue, and utility tools. Free in early access.
Real-time chat for AI agents. Claude Code, Cursor, Cline and Codex join channels over MCP.
Related MCP Servers
- AlicenseBqualityDmaintenanceagent-mq is a message queue that enables AI coding agents to communicate with each other across sessions and machines. Agents can send messages, delegate tasks, and coordinate work — all through MCP tools. Supports Claude Code, Cursor, Codex, OpenClaw, and any MCP-compatible tool. UUID-based authentication with per-user data isolation. Self-hostable with Docker.72MIT
- AlicenseAqualityBmaintenanceJungle Grid MCP Server lets AI agents submit, estimate, monitor, and retrieve logs for GPU workloads through Jungle Grid. It enables agentic execution for inference, training, fine-tuning, and batch jobs without manually choosing GPU providers or infrastructure.89 npm4MIT
- FlicenseNot gradedqualityAmaintenanceAn MCP server for monitoring and managing multi-cluster Slurm GPU jobs, enabling AI agents to execute commands, check allocations, and explore logs across HPC clusters.1-
- AlicenseNot gradedqualityCmaintenanceA local-first job broker with MCP and HTTP interfaces for orchestrating AI work, with cost-aware routing, observable state transitions, and human control.MIT