Skip to main content
Glama

jobd

CI PyPI Python Glama GHCR License: MIT DOI

A self-hostable, GPU-aware job broker for your own machines — with native MCP/agent integration.

Like task-spooler or pueue, but across all your machines — and VRAM-aware.

You have a couple of boxes with GPUs — a workstation, a server, maybe a laptop — wired together over Tailscale or a LAN. You want to fire off training runs, data pipelines, and long batch jobs from anywhere, have them land on whichever machine actually has the VRAM free, survive across sessions, and get preempted cleanly when something more important shows up. You don't have a cloud, a Kubernetes cluster, or a Slurm install, and you don't want one.

jobd is that missing piece: a lightweight, single-process broker that turns a handful of personal machines into a single queue — and an LLM agent can drive it directly.

# from any machine on your tailnet:
job submit --project myproj --gpu --vram-required 16 --wait -- python train.py
# → routed to whichever worker has ≥16 GB VRAM free, streamed back to your terminal

VRAM routing tracks one GPU per host: the worker reports free memory on GPU index 0 only.

Why it exists

Most schedulers assume a datacenter. The lightweight ones that don't (a bare nohup, a tmux session, an ssh-and-pray script) give you nothing: no queue, no VRAM-aware routing, no preemption, no record of what ran where. jobd fills the gap between "ssh in and run it" and "stand up Slurm":

  • VRAM-fit routing. The broker matches each job against live worker capacity (free VRAM / RAM / CPUs, capability tags, arch/OS) and dispatches to a worker that actually fits — instead of you guessing which box is free. One GPU per host is tracked (GPU index 0).

  • Preempt + checkpoint window. A higher-priority job can preempt a running one: the worker sends SIGTERM, the workload gets a grace window, then SIGKILL. jobd gives each job a per-job JOBD_CHECKPOINT_DIR to write into during that window; saving the checkpoint is the workload's job. A preempted job ends in the terminal preempted state and is not re-run automatically — to resume, you submit a new job pointed at the old checkpoint. (See docs/preemption.md.)

  • Survives sessions. Submit, close your laptop, check back tomorrow. Jobs live in the broker, not your shell.

  • Agent-native. Ships a first-class MCP server so an LLM agent (Claude Code, etc.) can submit, monitor, and babysit jobs as tool calls.

  • Yours. One broker process you run on a machine you own. No accounts, no telemetry, no per-GPU-hour billing; the broker and worker talk only to each other and to your clients. The optional self-update scripts do go online to fetch releases: scripts/update-worker.sh (which job fleet add installs on a timer) installs from PyPI, and scripts/deploy-broker.sh queries the GitHub API and pulls the image from GHCR.

  • Loopback by default, tailnet-only beyond it. The broker binds 127.0.0.1 unless you set JOBD_HOST. A request from any address that is neither loopback nor in Tailscale's CGNAT range (100.64.0.0/10) gets a 403 (JOBD_DISABLE_TAILNET_ACL=1 turns that check off).

Related MCP server: jungle-grid-mcp-server

Why not just use…?

Tool

What it gives you

Why jobd instead

nohup / tmux / ssh-and-pray

Runs a command on one box

No queue, no VRAM-aware routing, no preemption, no record of what ran where

task-spooler

A real job queue — on a single machine

jobd queues across all your machines and routes by live VRAM/CPU fit

Pueue

A mature single-machine command queue daemon

Pueue's own README declares distributed execution out of scope — jobd is that missing layer, plus GPU awareness

HyperQueue

Multi-machine task scheduling with HPC roots, single binary

HQ counts GPUs but doesn't track VRAM, and has no preemption/checkpoint contract or agent interface

Slurm

Datacenter-grade scheduling

Heavy to stand up and operate for 2–3 personal boxes; jobd is one process + a poller per host

SkyPilot / dstack

Provision and run on clouds + your own machines

SkyPilot's "existing machines" mode installs a Kubernetes cluster (k3s) on your boxes; dstack wants Docker + passwordless sudo on every host. jobd is one process + a poller — no containers, no sudo, no K8s

Modal

Serverless GPU compute on Modal's cloud

Cloud-only: it runs on Modal's machines, not yours

Ray

A distributed-compute framework; Ray Jobs also runs any shell command on a Ray cluster

You first stand up and run a Ray cluster; jobd is one broker process + a poller per host, with live VRAM-fit routing and a checkpoint window on preemption

Closest in spirit are Pueue and task-spooler (single-machine by design) and HyperQueue (multi-machine, HPC-shaped). jobd's niche is the 2–5-GPU homelab: multi-machine live VRAM-fit routing (one GPU per host tracked) + preempt/checkpoint window + a native MCP interface — a combination we haven't found in the tools above — with nothing heavier than a Python process per host.

Architecture

flowchart TD
    CLI["job CLI"]:::client --> B
    MCP["jobd-mcp<br/>MCP tools"]:::client --> B
    API["HTTP · SSE"]:::client --> B
    B["<b>jobd broker</b> — FastAPI<br/>queue · matcher · priorities · SQLite"]:::broker
    B <-->|poll · dispatch| WA["worker A<br/>24 GB GPU"]:::worker
    B <-->|poll · dispatch| WB["worker B<br/>8 GB GPU"]:::worker
    B <-->|poll · dispatch| WC["worker C<br/>CPU-only"]:::worker
    classDef client fill:#1f2937,stroke:#4b5563,color:#e5e7eb;
    classDef broker fill:#0e7490,stroke:#155e75,color:#ecfeff;
    classDef worker fill:#14532d,stroke:#166534,color:#dcfce7;

Workers poll the broker (pull model — no inbound connection to a worker); the broker matches each job against live capacity and hands it back on the poll. One broker process, one poller per host.

  • Broker — a FastAPI + SQLite service. Holds the queue, runs the matcher, resolves per-project priorities and defaults, exposes a small HTTP API and an SSE stream. Single source of truth.

  • Workers — lightweight polling agents, one per host. Each advertises live capacity via heartbeat, claims jobs it can run, executes them (shell=False: the argv you submit is run as-is, not through a shell — except job submit --stdin and the MCP jobd_submit tool, which take a command string and run it as bash -c <string>), streams logs back, and honors preemption signals.

  • Clients — the job CLI, the jobd-mcp MCP server, or anything that speaks the HTTP API.

Install

pip install jobd               # broker + CLI
pip install "jobd[mcp]"        # adds the MCP server
pip install "jobd[worker]"     # adds the worker daemon (jobd-worker)

Requires Python ≥ 3.11. Everything ships in the one jobd package: the broker (jobd), the CLI (job), the MCP server (jobd-mcp), and the worker (jobd-worker). The worker's extra runtime deps (psutil, nvidia-ml-py) live behind the [worker] extra since they're only needed on machines that actually run jobs. scripts/install-worker.sh sets a worker up under ~/jobd-worker with its own venv and a generated config.

Quickstart (single host)

# 1. start the broker (binds 127.0.0.1:8765 by default)
JOBD_ALLOW_NO_AUTH=1 jobd          # no-auth is fine for a loopback-only broker

# 2. in another shell, install + start a worker pointed at it
pip install "jobd[worker]"
JOBD_URL=http://127.0.0.1:8765 JOBD_WORKER_HOST=local jobd-worker

# 3. submit a job and wait for it
job submit --project demo --wait -- echo hello
job list
job logs <id>

For a real multi-host deployment (Docker broker + systemd workers, Tailscale binding, shared auth token), see docs/security.md and the templates in docker-compose.yml and scripts/. Adding a worker to a running fleet is one command:

job fleet add user@newbox      # ssh in, install pinned to the broker's version,
                               # wire systemd units + the self-update timer,
                               # verify it registers. `job fleet status` shows drift.

Day-2 operations (health, draining a worker, upgrades, token rotation, backups) are in docs/runbook.md.

Supported platforms

Python 3.11+ everywhere.

Component

Linux

macOS

Windows

Broker (jobd)

✅

☑️

☑️ (WSL recommended)

CLI (job) / MCP (jobd-mcp)

✅

☑️

☑️

Worker (jobd-worker)

✅ full

⚠️ degraded

untested

✅ = CI-tested on Linux (Ubuntu), with limits: CI has no GPU runner, so the NVIDIA/VRAM paths are tested only against mocks, and the tests that need a systemd --user scope skip on GitHub's runners. ☑️ = pure-Python and expected to work, but not exercised by CI — please file an issue if something is broken there.

The worker runs its best on Linux with a systemd user instance: memory caps, process reaping, and preemption use systemd-run --user scopes and cgroups. On non-systemd hosts the worker still executes jobs, but silently drops those guarantees — fine for a single trusted box, not for hard resource isolation. GPU features need NVIDIA + nvidia-ml-py. The broker, CLI, and MCP server are pure-Python and portable.

CLI

job submit -p PROJ [--gpu] [--vram-required N] [--needs TAG]... [--count N | --sweep K=v1,v2]... [--wait] -- CMD...
job list [--state STATE] [--project P] [--array A<id>]   # queue + recent jobs
job status ID | A<id> [--watch]             # one job, or an array's aggregate
job logs ID [-n BYTES]                      # tail captured output
job wait ID                                 # block until terminal
job cancel ID  /  job preempt ID            # stop a job
job adopt --pid N -p PROJ [--gpu]           # register a process you already started (Linux)
job workers                                 # fleet snapshot + health
job projects list | set NAME PRI | nudge NAME DELTA
job audit [--project P] [--since 24h]       # event history

job adopt makes a process started outside jobd (nohup, tmux) visible to the broker: it holds a slot and its VRAM until it exits, and nothing is launched. Its exit code is unknowable, so it ends orphaned, never completed — see docs/adoption.md.

job submit --explain dry-runs the resolution (priority, profile, project defaults, host pin) and prints the effective config without enqueuing anything.

Job arrays

Submit N jobs from one template with --count N. Each member is a normal job — it routes, runs, preempts, and checkpoints independently — and {i} in the command is replaced by the member's 0-based index:

job submit -p train --count 8 -- python train.py --fold {i}
# → Submitted array A42: 8 jobs (ids 42..49)

job list --array A42         # the members, with their index annotations
job status A42               # aggregate: state tally + per-member rollup

The array is identified as A<id> (the first member's job id). job status A42 exits non-zero if any member ended in a non-completed terminal state, so it composes with shell &&.

For a grid search, use --sweep KEY=v1,v2,v3 (repeatable) instead of --count. The broker fans out the cartesian product of all axes, substituting {KEY} per member; {i} (the flat member index) is also available:

job submit -p train --sweep lr=0.1,0.01 --sweep seed=1,2,3 \
  -- python train.py --lr {lr} --seed {seed} --out run-{i}
# → Submitted array A50: 6 jobs (ids 50..55)   # 2 × 3 = 6 members

--sweep and --count are mutually exclusive, the product is capped at 1000 members, and i is reserved as an axis key. Substitution is a literal {key} replace (not str.format), so JSON literals and shell braces in the command pass through untouched.

Coming from pueue or task-spooler?

Most verbs map directly — what changes is that the queue spans every machine you own. The last row is only a rough equivalent: jobd has no named groups with their own parallelism limit.

You ran…

With jobd

tsp <cmd> / pueue add -- <cmd>

job submit -p <project> -- <cmd>

tsp -w / pueue follow <id>

job logs -f <id> (or job wait <id>) — streams, exits with the job's own exit code

tsp / pueue status

job list

pueue log <id>

job logs <id>

commands piped to simple_gpu_scheduler

... | job submit -p <project> --stdin — one job per line, fleet-wide

pueue kill <id>

job cancel <id>

pueue group / parallelism limits

approximately: projects + priorities (projects.yaml); per-worker slots via JOBD_WORKER_MAX_CONCURRENT_JOBS

What you gain on top: jobs route to whichever machine actually has the VRAM/CPU free, live in the broker rather than in one machine's shell, can be preempted with a checkpoint window instead of killed, and are drivable by an LLM agent over MCP. What you lose: there is no pause/resume, stash, or edit of a queued job (cancel and resubmit instead), and if a worker dies mid-job, a running job not marked idempotent ends orphaned rather than being re-run. A one-machine deployment (broker + one worker on the same host) otherwise behaves like a network-reachable pueue.

MCP / agent integration

jobd ships an MCP server (jobd-mcp) exposing the queue as nine tools — jobd_submit, jobd_status, jobd_logs, jobd_list, jobd_cancel, jobd_preempt, jobd_events, jobd_workers, jobd_worker_delete. docs/agent-cookbook.md is the worked tour: fire-and-babysit polling, surviving preemption with checkpoints, sweeps, and asking the broker why a job won't schedule.

One-liner for Claude Code:

claude mcp add jobd --env JOBD_URL=http://127.0.0.1:8765 --env JOBD_API_TOKEN=<your-token> -- jobd-mcp

Or point any other MCP client at it:

{
  "mcpServers": {
    "jobd": {
      "command": "jobd-mcp",
      "env": {
        "JOBD_URL": "http://127.0.0.1:8765",
        "JOBD_API_TOKEN": "<your-token>"
      }
    }
  }
}

JOBD_API_TOKEN must match the broker's token, or every call returns 401. Omit it only when the broker runs with JOBD_ALLOW_NO_AUTH=1.

Now an agent can "run this overnight," check on it next session, and route GPU work through the broker instead of colliding on a shared card. The examples/claude-code-hooks/ directory has optional Claude Code hooks that nudge (or hard-block) an agent toward submitting heavy commands through jobd — including a VRAM-aware GPU guard with # NO_GPU / # CONCURRENT_OK / # VRAM=NGB override markers.

Configuration

Three optional YAML files under JOBD_CONFIG_DIR (default /app/config, the Docker image's path — set it when you run from pip). Example files live in this repo's config/ directory; they are not included in the pip package, and docker-compose.yml mounts them into the container:

  • projects.yaml — per-project base priority and submit defaults (preemptibility, wall/idle timeouts, host pins, capability requirements). Entries may also declare roots: so a job typed with an unregistered run label is priced by the project whose directory it runs in. See docs/projects-yaml.md for the full resolution model and docs/events.md for the event catalog.

  • profiles.yaml — named resource bundles (--profile gpu-train-large) the matcher uses to size a job.

  • classifier.yaml — rules that auto-suggest a profile from the command string.

All three are optional; with none present, every job runs at the global default priority.

Everything else is environment variables — the complete JOBD_* catalog (broker, worker, CLI/MCP, and the vars provided to workloads) lives in docs/configuration.md, and a CI test keeps it in lockstep with the source in both directions.

Concurrency (multislotting)

By default each worker runs one job at a time (JOBD_WORKER_MAX_CONCURRENT_JOBS=1). Raise it to let a worker bin-pack several jobs that fit side by side:

JOBD_WORKER_MAX_CONCURRENT_JOBS=3 jobd-worker

The matcher is resource-aware, so this is not blind N-up oversubscription. Each in-flight job reserves its vram_gb / ram_gb / cpus footprint, and the worker's heartbeat advertises only what's left. For VRAM, NVML's free figure already counts memory a running job has allocated, so only the not-yet-allocated part of the reservations is subtracted: free_vram = nvml_free − max(0, Σ in-flight vram_gb − VRAM held by this worker's jobs), where nvml_free is GPU index 0 only. RAM and CPUs subtract the full reservations. The broker won't place a job that doesn't fit the remaining headroom. The practical payoff: a CPU-only job and a GPU job run at the same time — the CPU job reserves 0 VRAM, so it never blocks the GPU slot, and vice-versa. Two GPU jobs co-run only if both fit live VRAM: before starting a job it was handed, the worker re-reads free VRAM and refuses the job if it no longer fits. A job whose command contains the literal marker # CONCURRENT_OK skips that last check — use it when you know the requested VRAM is overstated.

job workers reports each worker's slot usage — running jobs out of max_concurrent — alongside the live resource ad:

// job workers
{ "host": "desktop", "state": "online", "running": 2, "max_concurrent": 3,
  "free_vram_gb": 9.1, "idle_cpus": 6, ... }

Set the limit per worker from its environment (systemd unit, shell, or worker.yaml env) — it's a worker-local knob, not a broker setting.

Retention

By default jobd keeps every job record and .log file forever — history is never lost. On a long-running broker, opt into pruning:

JOBD_JOB_RETENTION_DAYS=30 jobd   # delete terminal jobs + their logs after 30 days

The sweeper deletes jobs in a terminal state whose finished_at is older than the horizon, unlinks their per-job .log, and emits a jobs_pruned event. Freed SQLite pages are reused under WAL, so the DB file stays bounded without a global-locking VACUUM. The default (0) keeps everything; pruning old terminal parents is safe for any still-pending dependents.

Security

The broker has no TCP-layer auth beyond a shared bearer token, so it is meant to run on a trusted network (loopback or a Tailscale tailnet), never on a public interface. Three stacked controls:

  1. Interface binding — set JOBD_HOST to 127.0.0.1 (the default) or a Tailscale CGNAT address (100.64.0.0/10), not 0.0.0.0. The broker itself does not enforce this: it starts on any bind address, and only warns (or refuses) when a non-loopback bind is combined with no-auth. The only check on the bind value is a CI lint (tests/test_deploy_lint.py) on the shipped Docker deployment; control 2 is what holds at runtime.

  2. Source-IP check (runtime) — whatever the bind, the broker answers 403 to any request whose source address is neither loopback nor in 100.64.0.0/10 (src/jobd/auth.py). JOBD_DISABLE_TAILNET_ACL=1 turns this off.

  3. Bearer token — set JOBD_API_TOKEN (≥32 random bytes) on every broker/worker/CLI/MCP host. The broker refuses to start without it unless you explicitly set JOBD_ALLOW_NO_AUTH=1. JOBD_ALLOW_NO_AUTH=1 is for a loopback-only broker (JOBD_HOST=127.0.0.1) — for local dev/tests. Combined with a non-loopback JOBD_HOST it exposes an unauthenticated RCE endpoint to your whole tailnet; the broker logs a startup warning if you do this. Don't.

Three endpoints are exempt from the source-IP check and the token — /livez, /readyz and /metrics answer with no bearer token and no source-IP check, because a generic HTTP monitor cannot send a token. /metrics is the one that matters: it publishes the broker version, job counts by state, and every worker's hostname and version. No commands, cwd, env or project names — but it does fingerprint the fleet. That is why, for these three, the JOBD_HOST bind above is load-bearing rather than defence-in-depth: port-forward the broker and you publish that inventory. Full table: Unauthenticated surface.

Full threat model, env-var reference, and token rotation: docs/security.md.

License

MIT — see LICENSE.

Available Tools

9 tools
jobd_cancelB

Cancel a job (queued → cancelled; running → SIGTERM via worker signal poll, ~2s).

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYes
reasonNoWhy — recorded on the broker's job_cancelled event (see jobd_events).

TDQS

B3.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral transparency burden on its own. It discloses meaningful state-dependent behavior: queued jobs are cancelled, and running jobs receive SIGTERM via a worker signal poll in about 2 seconds. It does not cover idempotency, auth, or failure conditions, but the core cancellation semantics are clearly exposed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single dense sentence that front-loads the primary action and immediately clarifies state-dependent behavior. There is no filler, and the parenthetical adds crucial implementation detail without bloating the text.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The description is adequate for a simple cancellation operation, but gaps remain. It does not explain the response/acknowledgement behavior, what happens if the job is already in a terminal state, or how cancellation relates to the sibling jobd_preempt. Given that there is no output schema to compensate, these missing contextual details keep it below a 4.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 50%, and the description adds little parameter-level meaning. 'Cancel a job' implicitly identifies job_id as the target job, but the description does not elaborate on job_id or reason. The reason parameter is already described in the schema, so the description itself contributes minimal semantic value beyond the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb and resource: 'Cancel a job'. It further clarifies the two main job states and what happens to each, which is more informative than a generic 'cancel job'. However, it does not explicitly distinguish itself from the sibling jobd_preempt, so it does not earn a 5.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives no explicit guidance on when to use this tool versus alternatives such as jobd_preempt or jobd_submit. It implies usage through its name and purpose, but the agent is left to infer the appropriate context, and no exclusions or alternative routing are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jobd_eventsA

The broker's event stream — the surface that explains WHY, not just what. /jobs says a job is queued; only this says it has been skipped 400 times because no worker advertises cuda-32gb, or that its dependency was cancelled, or that a watchdog killed it. Filter by since (2h/3d/1w), event type, job_id, project, or source (broker|worker). Use when a job is not doing what you expect and jobd_status alone does not explain it.

ParametersJSON Schema
NameRequiredDescriptionDefault
eventNoFilter to one event type. Known types: admission_blocked, auto_preempt, checkpoint_complete, cwd_identity_applied, cwd_refused, cwd_route_warning, dispatch_skip, env_scrubbed, gpu_contention_warning, job_adopted, job_cancelled, job_completed, job_dispatched, job_orphaned, job_resurrected, job_retry_scheduled, job_started, job_submitted, job_uncancelled, jobs_pruned, logs_pruned, preflight_warning, reclaim_suppressed, scheduling_timeout, serialization_warning, stale_scope_sweep, submit_warning, sweep_warning, unknown_project, version_drift, watchdog_fired, worker_offline, worker_registered, worker_shutdown, worker_stale. Hook-ingested events may carry custom names beyond these.
limitNoMax rows, newest-last. Broker clamps to 10000.
sinceNoRelative window (2h, 3d, 1w) or an ISO-8601 timestamp. Default: all retained.
job_idNoOnly events for this job.
sourceNoWhich side emitted the event.
projectNoOnly events for this project.

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

There are no annotations, so the description carries the full burden of behavioral disclosure. It does explain the semantic role of the stream and its diagnostic value, but it never explicitly states that this is a read-only operation, nor does it describe response shape, ordering, retention, or error behavior. The schema supplies some behavioral details (limit default and broker clamp), so this is not a complete gap, but the description alone is only minimally transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with a strong purpose statement and every sentence contributes something: reason to use it, concrete illustrative examples, filter options, and a selection rule. It is slightly more expansive than necessary, but not bloated or repetitive.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a filtered event stream with six optional parameters, the description covers the key selection context and when to use it, while the schema fills in defaults, enums, and the known event-type list. The lack of annotations and an output schema means returned payload shape is not fully specified, and the source mismatch is an inconsistency, but the tool is still reasonably invokable and well contextualized.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline for parameter semantics is 3. However, the description's filter summary says source (broker|worker, while the schema enumerates four sources: broker, worker, hook, and mcp. This narrowing contradicts the schema and could lead an agent to ignore hook or mcp sources; it also omits limit from the summary. The description adds conceptual examples but actually detracts from parameter clarity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource—the broker's event stream—and states its distinguishing purpose: explaining WHY, not just what. It differentiates itself from /jobs and jobd_status with concrete examples (skipped 400 times, dependency cancelled, watchdog killed), so an agent can tell it apart from sibling tools without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives an explicit when-to-use rule: 'Use when a job is not doing what you expect and jobd_status alone does not explain it.' It also contrasts with /jobs and mentions the alternative surface (jobd_status), providing an agent with a clear decision condition.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jobd_listA

List jobs on the broker with per-state counts. Defaults to the active set (queued/assigned/running); filter by state (e.g. ['failed']) or project to find past runs. Each row is a compact summary: job_id, project, state, host, exit_code, queued_at, started_at — call jobd_status for a job's full record. Use to answer 'what is running / queued right now?' or to locate a job id you've lost.

ParametersJSON Schema
NameRequiredDescriptionDefault
limitNoMax jobs returned (newest first; clamped to [1,200]). `counts` still covers every job matching the filters; a `truncated` field reports how many were cut.
stateNoStates to include — any of: queued, assigned, running, completed, failed, cancelled, preempted, orphaned, scheduling_timeout. Omit for the active set (queued/assigned/running); pass [] for all states.
projectNoRestrict to one project's jobs (the --project value used at submit).

TDQS

A4.3/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the behavioral burden. It discloses that the tool defaults to the active set, returns per-state counts, and emits compact summary rows with specified fields. It also clarifies that jobd_status is needed for full records. It doesn't explicitly state that the operation is read-only or side-effect-free, but the verb 'list' implies it, and the behavior described is accurate and useful.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three sentences, each earning its place: purpose plus default scope, filtering guidance, and output format with a pointer to jobd_status. The most important operational details are front-loaded, with no filler or repetition of schema content.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

No output schema exists, so the description compensates by enumerating the row fields and indicating that counts accompany the list. It also explains defaults, filter options, and how to get full records. This is complete for a list-style tool of moderate complexity, and the schema handles the remaining input details.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description adds a little usage context ('filter by state (e.g. ["failed"]) or project to find past runs') but does not materially extend what the schema already documents. It earns the baseline but no more.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'List jobs on the broker with per-state counts.' It clearly distinguishes itself from sibling jobd_status by stating that each row is a compact summary and that jobd_status provides the full record. An agent can immediately tell what this tool does and how it differs from related tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It states when to use the tool: answer 'what is running / queued right now?' or locate a lost job id, and it explains the default active set and how to filter by state or project for past runs. It doesn't explicitly say 'use this instead of X' beyond pointing to jobd_status for full records, but the use case is clearly scoped.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jobd_logsA

Tail the captured stdout/stderr of a job (workers stream output to the broker's per-job log as it runs). Returns log_tail (last tail_bytes, default 8 KiB, max 1 MiB) plus size_bytes/returned_bytes/truncated — works for running AND finished jobs. Use to check progress mid-run, diagnose a failure's traceback, or grab a job's final output.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYesNumeric job id whose captured output to read.
tail_bytesNoHow many bytes from the END of the log to return (server caps reads at 1 MiB). Raise for context, lower for a quick liveness peek.

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden. It discloses stream mechanics ('workers stream output to the broker's per-job log as it runs'), lifecycle support ('running AND finished'), and exact return semantics including tail_bytes default/max and truncation metadata. It does not explicitly state permissions, rate limits, or side-effect absence, but those are less critical for a tail/read operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three dense sentences front-load the core action, follow with return format and lifecycle nuance, and end with concrete use cases. No filler, no repetition of schema fields that isn't functional.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite lacking an output schema, the description fully specifies the return shape (log_tail plus size_bytes/returned_bytes/truncated), the parameter semantics, and the supported job states. For a 2-parameter read tool, nothing essential is missing for an agent to select and invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds practical guidance beyond the schema: 'Raise for context, lower for a quick liveness peek' for tail_bytes, and it reiterates the default and max. This is modest but genuine added semantic value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb+resource: 'Tail the captured stdout/stderr of a job.' The description clearly distinguishes from siblings like jobd_status by focusing on output logs, and clarifies it returns a tail of the log. It names the return payload and available metadata, making the tool's function unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit use cases are given: 'check progress mid-run, diagnose a failure's traceback, or grab a job's final output.' It also notes the tool works for 'running AND finished jobs,' which helps an agent decide when to call it. It does not name alternatives or when not to use it, so it stops shy of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jobd_preemptA

Preempt a running/assigned preemptible job (worker SIGTERMs with grace; final state 'preempted'). Refused if not preemptible or not running.

ParametersJSON Schema
NameRequiredDescriptionDefault
job_idYes

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, this description carries the behavioral burden and does well: it discloses the worker side effect (SIGTERM with grace), the terminal state ('preempted'), and refusal conditions. There is a minor ambiguity between 'assigned' being preemptable and 'refused if not running', but overall the key behaviors are transparent.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The entire description is one dense sentence with no filler. It front-loads the core action and immediately follows with behavioral consequences and refusal criteria; every phrase earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter tool with no output schema, the description supplies the essential invocation context: what happens, what the final state is, and when the call fails. It does not describe the return payload, but given no output schema and simple side effects, the description is reasonably complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 0%, but the sole parameter job_id is a simple integer whose meaning is obvious from the description. The description adds semantic constraints on which job_ids are valid (preemptible, running/assigned), which is useful context beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Preempt') and a precise resource ('a running/assigned preemptible job'), and it names the distinguishing final state ('preempted'). It also implicitly differentiates from jobd_cancel by restricting preemption to preemptible jobs, so an agent can tell which sibling applies.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It clearly states when the tool is valid ('running/assigned preemptible job') and when it will be refused ('if not preemptible or not running'). It does not explicitly name an alternative such as jobd_cancel for non-preemptible jobs, which keeps it just short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jobd_statusA

The full JobInfo record for one job_id: state, exit_code, timings, host, and the scheduling internals — depends_on + cascade policy (depends_on_any_exit), pending cancel/preempt signal, resolved profile, requires (gpu/tags/idempotent), host pin, fast_path, timeouts, termination_reason. Use for a quick state check AND for debugging why a job routed/failed/stalled. Pass wait=true to block until terminal or wait_timeout_s.

ParametersJSON Schema
NameRequiredDescriptionDefault
waitNo
job_idYes
wait_timeout_sNoServer clamps to 270.

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations present, the description carries the behavioral disclosure burden. It discloses the full return surface, the wait=true blocking behavior, and the timeout parameter, which is substantial context. It does not mention error behavior or side effects, but for a status tool this is reasonably complete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but every clause earns its place: return contents first, use cases second, wait behavior last. It avoids fluff and is well-structured for scanning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the lack of an output schema, the description compensates by enumerating the returned fields and giving practical use cases. It could additionally explain how this tool differs from jobd_events or jobd_logs, and clarify whether a timeout returns partial data or errors, but it is largely sufficient.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 33%, but the description explains wait semantics directly ('block until terminal or wait_timeout_s') and makes job_id's role implicitly clear. The timeout clamp is already in the schema. This adds meaningful meaning beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description precisely identifies the resource ('full JobInfo record for one job_id') and enumerates the returned fields, from state and exit_code to scheduling internals. It also distinguishes itself from sibling tools like jobd_list by emphasizing a single job and deep debugging detail.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description clearly states when to use the tool: for a quick state check and for debugging why a job routed, failed, or stalled. It does not explicitly name alternatives or exclusions, so it stops short of full guidance, but the usage context is unambiguous.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jobd_submitA

Submit a job to the jobd broker. Default async; pass wait=true to block up to wait_timeout_s (server clamps to 270).

ParametersJSON Schema
NameRequiredDescriptionDefault
cwdYesAbsolute path; broker validates against worker mount_roots.
gpuNotrue = pin to a GPU-capable worker. false or omitted = no GPU preference (any worker, GPU or not) — the same as leaving off the CLI's --gpu. Forbidding GPU workers (CLI --no-gpu) is not offered here.
hostNoHost alias pin (laptop, desktop-vm).
waitNoSync mode: block until terminal or timeout. For an array submit (count/sweep), waits on every member under one shared deadline and returns an aggregate {array_id, count, job_ids, states, all_completed, members:[{job_id, state, exit_code}]}.
extraNoEscape hatch: idempotent (bool), depends_on (int[]), depends_on_any_exit (bool), priority (int delta), max_wall_s (int), idle_timeout_s (int), max_retries (int, default 0: re-run up to N times on a plain non-zero exit), retry_delay_s (int), scheduling_timeout_s (int 1..604800 — give up and terminate the job as 'scheduling_timeout' if it is still QUEUED after N seconds; omit to wait indefinitely for a capable worker), checkpoint_grace_s (int 1..300), vram_gb (float — explicit GPU VRAM the job needs at dispatch; falls back to cuda-Ngb tier-tag max, then to 2 GB floor for --gpu jobs), count (int 1..1000 — submit a job array of N members, with `{i}` in the command replaced by the 0-based index; response is {array_id, count, job_ids, warnings} instead of a single job), sweep (list of {key, values[]} — parameter-sweep axes; broker fans out the cartesian product, substituting `{key}` per member plus `{i}`; mutually exclusive with count; product capped at 1000), profile (str), env (dict), preemptible (bool), session_id (str), arch (str — pin to a worker CPU arch), os (str — pin to a worker OS).
needsNoTool tags (R, python3, cuda).
commandYesShell command run by the worker shell.
dry_runNoPreview mode: run full validation + routing decision (profile, project defaults, cwd, depends_on, preflight, gpu_contention) and return the would-be plan WITHOUT queueing. Response has state='dry-run', would_route_to (list[host]), would_use_worker (host or null), validation (resolved fields + warnings). Per dry-run convention 2026-05-18.
projectYesScheduling identity. A registered projects.yaml name (matched case- and -/_-insensitively) prices at its priority; an unregistered name is priced by the project whose roots: contain cwd, else by _default. The result's project_label carries the name as typed when the two differ.
wait_timeout_sNoSeconds; permissive — server clamps to 270.

TDQS

A4.1/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden, and it does disclose key behavioral traits: submission is asynchronous by default, wait=true switches to blocking, and wait_timeout_s is clamped to 270. It does not describe success responses or error conventions, though the side effect of submission is clear from the verb.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with no filler: the primary behavior is front-loaded, and the timeout caveat is stated compactly. Nothing in the description repeats schema text unnecessarily.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter submission tool with no output schema, the description is relatively thin: it explains sync versus async but never states what a successful response contains (e.g., job_id or array_id), which the agent needs to hand off to jobd_status or jobd_logs. The rich parameter schema compensates for parameter details, but response expectations are left to inference.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the parameter baseline is 3. The description adds a useful restatement of wait/wait_timeout_s semantics, but the remaining parameters are already fully documented in the schema, so no further compensation is needed.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a direct verb–resource pair ('Submit a job to the jobd broker') that clearly states the operation. It is set apart from the sibling tools, which are all post-submission queries or mutations such as jobd_status, jobd_logs, and jobd_cancel.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It gives an explicit usage default ('Default async') and the exact condition for the blocking alternative ('pass wait=true'), so an agent knows how to invoke it. It does not explicitly name sibling tools as alternatives, but the sibling list is dominated by post-submission operations, making the correct context fairly clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jobd_worker_deleteA

Remove a worker from the broker registry. The broker refuses (409) if the worker is still online — caller stops the worker process or waits for the heartbeat sweeper first.

ParametersJSON Schema
NameRequiredDescriptionDefault
hostYesWorker host identifier.

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden of behavioral disclosure. It reveals the 409 conflict behavior, the online/offline prerequisite, and the need to stop the worker or wait for the sweeper. This is substantive context beyond what the schema alone provides.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences, each earning its place: the first states the action, the second explains the key precondition and failure mode. No filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter, no-output-schema tool, the description covers the action, prerequisite, and error condition. It could optionally mention the success response shape or idempotency, but nothing essential to invoking it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the parameter 'host' is already documented as 'Worker host identifier.' The tool description adds no further meaning about the host parameter, so a baseline score of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('Remove') with a specific resource ('worker from the broker registry'), making the action unmistakable. It clearly differentiates from siblings like jobd_workers or jobd_submit without needing to compare schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context: the worker must not be online, and the broker will return 409 otherwise. It also tells the caller how to proceed ('stops the worker process or waits for the heartbeat sweeper first'), though it does not explicitly mention alternative tools for inspecting workers before deletion.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

jobd_workersA

Fleet snapshot: every registered worker with state (online/stale/offline), live capacity ad (free_vram_gb, unregistered_vram_gb, free_ram_gb, idle_cpus), capability tags (cuda tiers, arch/os), slot usage (running/max_concurrent), and last_heartbeat — plus an overall health rollup (healthy|degraded|empty). Use before submitting GPU work to see what's free, or to diagnose why a job isn't being dispatched.

ParametersJSON Schema
NameRequiredDescriptionDefault

No parameters

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description handles transparency itself by labeling it a 'snapshot' (read-only point-in-time view) and naming all returned categories. It doesn't cover auth, staleness, or error behavior, but for a zero-parameter fleet status read the essential behavior is disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The first sentence front-loads 'Fleet snapshot' and then packs the output contract into compact parenthetical lists; the second sentence adds concrete usage guidance. Every clause earns its place and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite lacking an output schema, the description enumerates the output fields, their possible values, and the health rollup states, and it supplies two relevant use cases. For a parameterless read-only fleet query, this is sufficient for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The tool takes zero parameters and schema coverage is 100%, so there are no parameter semantics to add. The description's field list primarily documents return content rather than inputs, which fits the baseline for parameterless tools.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Fleet snapshot' and enumerates exactly what it returns: worker state, capacity fields, capability tags, slot usage, heartbeat, and health rollup. This makes the tool's resource and output clear, though it does not explicitly differentiate it from sibling tools like jobd_status or jobd_list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It provides explicit use cases: use before submitting GPU work to check available capacity, or to diagnose why jobs aren't dispatched. It doesn't spell out when not to use it or name alternative tools, but the context is strong.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 9 tool updatesv0.5.47
    • Changedjobd_cancel2 fields changed
      • addedInput schema / additionalProperties
        Added value: +false
      • addedInput schema / properties / reason / description
        Added value: +"Why — recorded on the broker's job_cancelled event (see jobd_events)."
    • Changedjobd_events2 fields changed
      • addedInput schema / additionalProperties
        Added value: +false
      • changedInput schema / properties / event / description
        Previous value: -"Filter to one event type. Known types: admission_blocked, auto_preempt, checkpoint_complete, cwd_identity_applied, cwd_refused, cwd_route_warning, dispatch_skip, env_scrubbed, gpu_contention_warning, job_cancelled, job_completed, job_dispatched, job_orphaned, job_resurrected, job_started, job_submitted, job_uncancelled, jobs_pruned, logs_pruned, preflight_warning, reclaim_suppressed, scheduling_timeout, serialization_warning, stale_scope_sweep, submit_warning, sweep_warning, unknown_project, version_drift, watchdog_fired, worker_offline, worker_registered, worker_shutdown, worker_stale. Hook-ingested events may carry custom names beyond these."New value: +"Filter to one event type. Known types: admission_blocked, auto_preempt, checkpoint_complete, cwd_identity_applied, cwd_refused, cwd_route_warning, dispatch_skip, env_scrubbed, gpu_contention_warning, job_adopted, job_cancelled, job_completed, job_dispatched, job_orphaned, job_resurrected, job_retry_scheduled, job_started, job_submitted, job_uncancelled, jobs_pruned, logs_pruned, preflight_warning, reclaim_suppressed, scheduling_timeout, serialization_warning, stale_scope_sweep, submit_warning, sweep_warning, unknown_project, version_drift, watchdog_fired, worker_offline, worker_registered, worker_shutdown, worker_stale. Hook-ingested events may carry custom names beyond these."
    • Changedjobd_list2 fields changed
      • addedInput schema / additionalProperties
        Added value: +false
      • addedInput schema / properties / state / items / enum
        Added value: +[
        +  "queued",
        +  "assigned",
        +  "running",
        +  "completed",
        +  "failed",
        +  "cancelled",
        +  "preempted",
        +  "orphaned",
        +  "scheduling_timeout"
        +]
    • Changedjobd_logs1 field changed
      • addedInput schema / additionalProperties
        Added value: +false
    • Changedjobd_preempt1 field changed
      • addedInput schema / additionalProperties
        Added value: +false
    • Changedjobd_status1 field changed
      • addedInput schema / additionalProperties
        Added value: +false
    • Changedjobd_submit6 fields changed
      • addedInput schema / $defs
        Added value: +{
        +  "JobRequires": {
        +    "additionalProperties": false,
        +    "properties": {
        +      "arch": {
        +        "default": "any",
        +        "title": "Arch",
        +        "type": "string"
        +      },
        +      "gpu": {
        +        "anyOf": [
        +          {
        +            "type": "boolean"
        +          },
        +          {
        +            "type": "null"
        +          }
        +        ],
        +        "default": null,
        +        "title": "Gpu"
        +      },
        +      "idempotent": {
        +        "default": false,
        +        "title": "Idempotent",
        +        "type": "boolean"
        +      },
        +      "needs": {
        +        "items": {
        +          "type": "string"
        +        },
        +        "title": "Needs",
        +        "type": "array"
        +      },
        +      "os": {
        +        "default": "any",
        +        "title": "Os",
        +        "type": "string"
        +      }
        +    },
        +    "title": "JobRequires",
        +    "type": "object"
        +  },
        +  "SweepAxis": {
        +    "additionalProperties": false,
        +    "description": "One named axis of a parameter sweep: a key and the values it ranges over.\n\nThe broker takes the cartesian product of all axes to fan out array members;\neach member substitutes `{key}` → its value in the command and env. `{i}`\n(the flat member index) is always available alongside the named keys, so\n`i` is reserved and rejected as an axis key. See jobd.arrays.",
        +    "properties": {
        +      "key": {
        +        "minLength": 1,
        +        "title": "Key",
        +        "type": "string"
        +      },
        +      "values": {
        +        "items": {
        +          "type": "string"
        +        },
        +        "minItems": 1,
        +        "title": "Values",
        +        "type": "array"
        +      }
        +    },
        +    "required": [
        +      "key",
        +      "values"
        +    ],
        +    "title": "SweepAxis",
        +    "type": "object"
        +  }
        +}
      • addedInput schema / additionalProperties
        Added value: +false
      • changedInput schema / properties / extra / additionalProperties
        Previous value: -trueNew value: +false
      • changedInput schema / properties / extra / description
        Previous value: -"Escape hatch: idempotent (bool), depends_on (int[]), depends_on_any_exit (bool), priority (int delta), max_wall_s (int), idle_timeout_s (int), scheduling_timeout_s (int 1..604800 — give up and terminate the job as 'scheduling_timeout' if it is still QUEUED after N seconds; omit to wait indefinitely for a capable worker), checkpoint_grace_s (int 1..300), vram_gb (float — explicit GPU VRAM the job needs at dispatch; falls back to cuda-Ngb tier-tag max, then to 2 GB floor for --gpu jobs), count (int 1..1000 — submit a job array of N members, with `{i}` in the command replaced by the 0-based index; response is {array_id, count, job_ids, warnings} instead of a single job), sweep (list of {key, values[]} — parameter-sweep axes; broker fans out the cartesian product, substituting `{key}` per member plus `{i}`; mutually exclusive with count; product capped at 1000), profile (str), env (dict), preemptible (bool), session_id (str), arch (str — pin to a worker CPU arch), os (str — pin to a worker OS)."New value: +"Escape hatch: idempotent (bool), depends_on (int[]), depends_on_any_exit (bool), priority (int delta), max_wall_s (int), idle_timeout_s (int), max_retries (int, default 0: re-run up to N times on a plain non-zero exit), retry_delay_s (int), scheduling_timeout_s (int 1..604800 — give up and terminate the job as 'scheduling_timeout' if it is still QUEUED after N seconds; omit to wait indefinitely for a capable worker), checkpoint_grace_s (int 1..300), vram_gb (float — explicit GPU VRAM the job needs at dispatch; falls back to cuda-Ngb tier-tag max, then to 2 GB floor for --gpu jobs), count (int 1..1000 — submit a job array of N members, with `{i}` in the command replaced by the 0-based index; response is {array_id, count, job_ids, warnings} instead of a single job), sweep (list of {key, values[]} — parameter-sweep axes; broker fans out the cartesian product, substituting `{key}` per member plus `{i}`; mutually exclusive with count; product capped at 1000), profile (str), env (dict), preemptible (bool), session_id (str), arch (str — pin to a worker CPU arch), os (str — pin to a worker OS)."
      • addedInput schema / properties / extra / properties
        Added value: +{
        +  "arch": {
        +    "default": "any",
        +    "title": "Arch",
        +    "type": "string"
        +  },
        +  "checkpoint_grace_s": {
        +    "anyOf": [
        +      {
        +        "maximum": 300,
        +        "minimum": 1,
        +        "type": "integer"
        +      },
        +      {
        +        "type": "null"
        +      }
        +    ],
        +    "default": null,
        +    "title": "Checkpoint Grace S"
        +  },
        +  "count": {
        +    "default": 1,
        +    "maximum": 1000,
        +    "minimum": 1,
        +    "title": "Count",
        +    "type": "integer"
        +  },
        +  "depends_on": {
        +    "items": {
        +      "type": "integer"
        +    },
        +    "title": "Depends On",
        +    "type": "array"
        +  },
        +  "depends_on_any_exit": {
        +    "default": false,
        +    "title": "Depends On Any Exit",
        +    "type": "boolean"
        +  },
        +  "env": {
        +    "additionalProperties": {
        +      "type": "string"
        +    },
        +    "title": "Env",
        +    "type": "object"
        +  },
        +  "idempotent": {
        +    "default": false,
        +    "title": "Idempotent",
        +    "type": "boolean"
        +  },
        +  "idle_timeout_s": {
        +    "anyOf": [
        +      {
        +        "maximum": 86400,
        +        "minimum": 1,
        +        "type": "integer"
        +      },
        +      {
        +        "type": "null"
        +      }
        +    ],
        +    "default": null,
        +    "title": "Idle Timeout S"
        +  },
        +  "max_retries": {
        +    "default": 0,
        +    "maximum": 20,
        +    "minimum": 0,
        +    "title": "Max Retries",
        +    "type": "integer"
        +  },
        +  "max_wall_s": {
        +    "anyOf": [
        +      {
        +        "maximum": 604800,
        +        "minimum": 1,
        +        "type": "integer"
        +      },
        +      {
        +        "type": "null"
        +      }
        +    ],
        +    "default": null,
        +    "title": "Max Wall S"
        +  },
        +  "os": {
        +    "default": "any",
        +    "title": "Os",
        +    "type": "string"
        +  },
        +  "preemptible": {
        +    "anyOf": [
        +      {
        +        "type": "boolean"
        +      },
        +      {
        +        "type": "null"
        +      }
        +    ],
        +    "default": null,
        +    "title": "Preemptible"
        +  },
        +  "priority": {
        +    "default": 0,
        +    "title": "Priority Delta",
        +    "type": "integer"
        +  },
        +  "profile": {
        +    "anyOf": [
        +      {
        +        "type": "string"
        +      },
        +      {
        +        "type": "null"
        +      }
        +    ],
        +    "default": null,
        +    "title": "Profile"
        +  },
        +  "retry_delay_s": {
        +    "default": 0,
        +    "maximum": 86400,
        +    "minimum": 0,
        +    "title": "Retry Delay S",
        +    "type": "integer"
        +  },
        +  "scheduling_timeout_s": {
        +    "anyOf": [
        +      {
        +        "maximum": 604800,
        +        "minimum": 1,
        +        "type": "integer"
        +      },
        +      {
        +        "type": "null"
        +      }
        +    ],
        +    "default": null,
        +    "title": "Scheduling Timeout S"
        +  },
        +  "session_id": {
        +    "anyOf": [
        +      {
        +        "type": "string"
        +      },
        +      {
        +        "type": "null"
        +      }
        +    ],
        +    "default": null,
        +    "title": "Session Id"
        +  },
        +  "sweep": {
        +    "items": {
        +      "$ref": "#/$defs/SweepAxis"
        +    },
        +    "title": "Sweep",
        +    "type": "array"
        +  },
        +  "vram_gb": {
        +    "default": 0,
        +    "minimum": 0,
        +    "title": "Vram Gb",
        +    "type": "number"
        +  }
        +}
      • changedInput schema / properties / gpu / description
        Previous value: -"Pin to GPU-capable worker."New value: +"true = pin to a GPU-capable worker. false or omitted = no GPU preference (any worker, GPU or not) — the same as leaving off the CLI's --gpu. Forbidding GPU workers (CLI --no-gpu) is not offered here."
    • Changedjobd_worker_delete1 field changed
      • addedInput schema / additionalProperties
        Added value: +false
    • Changedjobd_workers1 field changed
      • addedInput schema / additionalProperties
        Added value: +false
  2. 2 tool updatesv0.5.44
    • Changedjobd_events1 field changed
      • changedInput schema / properties / event / description
        Previous value: -"Filter to one event type. Known types: admission_blocked, auto_preempt, checkpoint_complete, cwd_refused, dispatch_skip, env_scrubbed, job_cancelled, job_completed, job_dispatched, job_orphaned, job_resurrected, job_started, job_submitted, job_uncancelled, jobs_pruned, logs_pruned, reclaim_suppressed, scheduling_timeout, stale_scope_sweep, submit_warning, sweep_warning, version_drift, watchdog_fired, worker_offline, worker_registered, worker_shutdown, worker_stale. Hook-ingested events may carry custom names beyond these."New value: +"Filter to one event type. Known types: admission_blocked, auto_preempt, checkpoint_complete, cwd_identity_applied, cwd_refused, cwd_route_warning, dispatch_skip, env_scrubbed, gpu_contention_warning, job_cancelled, job_completed, job_dispatched, job_orphaned, job_resurrected, job_started, job_submitted, job_uncancelled, jobs_pruned, logs_pruned, preflight_warning, reclaim_suppressed, scheduling_timeout, serialization_warning, stale_scope_sweep, submit_warning, sweep_warning, unknown_project, version_drift, watchdog_fired, worker_offline, worker_registered, worker_shutdown, worker_stale. Hook-ingested events may carry custom names beyond these."
    • Changedjobd_submit1 field changed
      • changedInput schema / properties / project / description
        Previous value: -"Priority lookup key; falls back to _default."New value: +"Scheduling identity. A registered projects.yaml name (matched case- and -/_-insensitively) prices at its priority; an unregistered name is priced by the project whose roots: contain cwd, else by _default. The result's project_label carries the name as typed when the two differ."
  3. 5 tool updatesv0.5.36
    • Addedjobd_events
    • Addedjobd_list
    • Addedjobd_preempt
    • Addedjobd_worker_delete
    • Addedjobd_workers
  4. 6 tool updatesv0.5.35
    • Removedjobd_list
    • Addedjobd_logs
    • Removedjobd_preempt
    • Addedjobd_status
    • Addedjobd_submit
    • Removedjobd_worker_delete
  5. 5 tool updatesv0.5.34
    • Removedjobd_events
    • Removedjobd_logs
    • Removedjobd_status
    • Removedjobd_submit
    • Removedjobd_workers
  6. 1 tool updatev0.5.31
    • Changedjobd_events1 field changed
      • changedInput schema / properties / event / description
        Previous value: -"Filter to one event type. Known types: admission_blocked, auto_preempt, checkpoint_complete, cwd_refused, dispatch_skip, job_cancelled, job_completed, job_dispatched, job_orphaned, job_resurrected, job_started, job_submitted, job_uncancelled, jobs_pruned, logs_pruned, scheduling_timeout, stale_scope_sweep, submit_warning, sweep_warning, watchdog_fired, worker_offline, worker_registered, worker_shutdown, worker_stale. Hook-ingested events may carry custom names beyond these."New value: +"Filter to one event type. Known types: admission_blocked, auto_preempt, checkpoint_complete, cwd_refused, dispatch_skip, env_scrubbed, job_cancelled, job_completed, job_dispatched, job_orphaned, job_resurrected, job_started, job_submitted, job_uncancelled, jobs_pruned, logs_pruned, scheduling_timeout, stale_scope_sweep, submit_warning, sweep_warning, version_drift, watchdog_fired, worker_offline, worker_registered, worker_shutdown, worker_stale. Hook-ingested events may carry custom names beyond these."
  7. 3 tool updatesv0.5.26
    • Addedjobd_events
    • Removedjobd_job_get
    • Changedjobd_submit1 field changed
      • changedInput schema / properties / extra / description
        Previous value: -"Escape hatch: idempotent (bool), depends_on (int[]), depends_on_any_exit (bool), priority (int delta), max_wall_s (int), idle_timeout_s (int), checkpoint_grace_s (int 1..300), vram_gb (float — explicit GPU VRAM the job needs at dispatch; falls back to cuda-Ngb tier-tag max, then to 2 GB floor for --gpu jobs), count (int 1..1000 — submit a job array of N members, with `{i}` in the command replaced by the 0-based index; response is {array_id, count, job_ids, warnings} instead of a single job), sweep (list of {key, values[]} — parameter-sweep axes; broker fans out the cartesian product, substituting `{key}` per member plus `{i}`; mutually exclusive with count; product capped at 1000), profile (str), env (dict), preemptible (bool)."New value: +"Escape hatch: idempotent (bool), depends_on (int[]), depends_on_any_exit (bool), priority (int delta), max_wall_s (int), idle_timeout_s (int), scheduling_timeout_s (int 1..604800 — give up and terminate the job as 'scheduling_timeout' if it is still QUEUED after N seconds; omit to wait indefinitely for a capable worker), checkpoint_grace_s (int 1..300), vram_gb (float — explicit GPU VRAM the job needs at dispatch; falls back to cuda-Ngb tier-tag max, then to 2 GB floor for --gpu jobs), count (int 1..1000 — submit a job array of N members, with `{i}` in the command replaced by the 0-based index; response is {array_id, count, job_ids, warnings} instead of a single job), sweep (list of {key, values[]} — parameter-sweep axes; broker fans out the cartesian product, substituting `{key}` per member plus `{i}`; mutually exclusive with count; product capped at 1000), profile (str), env (dict), preemptible (bool), session_id (str), arch (str — pin to a worker CPU arch), os (str — pin to a worker OS)."
  8. 1 tool updatev0.5.12
    • Changedjobd_list2 fields changed
      • changedInput schema / properties / limit / description
        Previous value: -"Advisory cap on returned jobs (the broker currently returns its default window)."New value: +"Max jobs returned (newest first; clamped to [1,200]). `counts` still covers every job matching the filters; a `truncated` field reports how many were cut."
      • changedInput schema / properties / state / description
        Previous value: -"States to include — any of: queued, assigned, running, completed, failed, cancelled, preempted, orphaned, scheduling_timeout. Currently only the first is forwarded to the broker (single state_filter)."New value: +"States to include — any of: queued, assigned, running, completed, failed, cancelled, preempted, orphaned, scheduling_timeout. Omit for the active set (queued/assigned/running); pass [] for all states."
  9. 3 tool updates
    • Changedjobd_job_get1 field changed
      • addedInput schema / properties / job_id / description
        Added value: +"Numeric job id as returned by jobd_submit or shown in jobd_list."
    • Changedjobd_list3 fields changed
      • addedInput schema / properties / limit / description
        Added value: +"Advisory cap on returned jobs (the broker currently returns its default window)."
      • addedInput schema / properties / project / description
        Added value: +"Restrict to one project's jobs (the --project value used at submit)."
      • changedInput schema / properties / state / description
        Previous value: -"States to include. Currently only the first is forwarded to the broker (single state_filter)."New value: +"States to include — any of: queued, assigned, running, completed, failed, cancelled, preempted, orphaned, scheduling_timeout. Currently only the first is forwarded to the broker (single state_filter)."
    • Changedjobd_logs2 fields changed
      • addedInput schema / properties / job_id / description
        Added value: +"Numeric job id whose captured output to read."
      • addedInput schema / properties / tail_bytes / description
        Added value: +"How many bytes from the END of the log to return (server caps reads at 1 MiB). Raise for context, lower for a quick liveness peek."
  10. 9 tool updatesv0.5.5
    • First observedjobd_cancel
    • First observedjobd_job_get
    • First observedjobd_list
    • First observedjobd_logs
    • First observedjobd_preempt
    • First observedjobd_status
    • First observedjobd_submit
    • First observedjobd_worker_delete
    • First observedjobd_workers

TDQS

A4.1/5.0

Scored across 9 tools

Disambiguation5/5

Each tool targets a distinct operation: status/logs/cancel/preempt/list/events/workers/worker_delete/submit. There is no overlap between them; even status vs events are clearly differentiated (state snapshot vs why-events).

Naming Consistency5/5

All tools follow a consistent jobd_<verb> or jobd_<noun>_<verb> pattern (jobd_status, jobd_logs, jobd_cancel, jobd_preempt, jobd_list, jobd_events, jobd_workers, jobd_worker_delete, jobd_submit). The pattern is uniform and predictable.

Tool Count5/5

9 tools is well-scoped for a job scheduling/broker domain: submit, inspect (status/logs/list/events), control (cancel/preempt), and fleet management (workers/worker_delete). Each tool earns its place.

Completeness4/5

The surface covers the full job lifecycle: submit, status, logs, cancel, preempt, list, events, and worker management. Minor gaps like resubmit or job deletion are absent, but the core scheduling workflow is complete.

Maintenance

ActivityActive
ResponsivenessWithin a week

Related MCP Connectors

Related MCP Servers

  • A
    license
    B
    quality
    D
    maintenance
    agent-mq is a message queue that enables AI coding agents to communicate with each other across sessions and machines. Agents can send messages, delegate tasks, and coordinate work — all through MCP tools. Supports Claude Code, Cursor, Codex, OpenClaw, and any MCP-compatible tool. UUID-based authentication with per-user data isolation. Self-hostable with Docker.
    7
    2
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    Jungle Grid MCP Server lets AI agents submit, estimate, monitor, and retrieve logs for GPU workloads through Jungle Grid. It enables agentic execution for inference, training, fine-tuning, and batch jobs without manually choosing GPU providers or infrastructure.
    8
    9 npm
    4
    MIT
  • F
    license
    Not graded
    quality
    A
    maintenance
    An MCP server for monitoring and managing multi-cluster Slurm GPU jobs, enabling AI agents to execute commands, check allocations, and explore logs across HPC clusters.
    1
    -