Skip to main content
Glama
office233
by office233

job_start

Destructive

Run long commands like tests, builds, or installs as background jobs; get pass/fail summaries or poll status if still running.

Instructions

Run a long command (test suite, build, install, benchmark) as a background job. Waits up to waitMs (default 25s): if it finishes you get a smart summary (pass/fail counts, failing tests with context, output tail) instead of the raw log; otherwise poll job_status.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
cwdNo
titleNo
waitMsNo
commandYes
timeoutMsNoKill after this long (default 1h)
resourceClassNoAdmission class; defaults to automatic command classification
resourceWaitMsNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv3.0.1

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare destructiveHint=true, openWorldHint=true, and idempotentHint=false, so the safety profile is covered. The description adds genuinely new behavior: the wait window, the smart-summary return shape, and the polling fallback when the job outlives the wait. It doesn't disclose that the job keeps running after waitMs elapses, which is implied but not stated.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two dense sentences: the capability first, then the wait/poll contract and what you get back. Every clause earns its place with no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Return behavior is well explained despite no output schema, and the wait/poll loop is clear. But for a 7-parameter destructive mutation tool, five parameters are undocumented in both the description and the schema, leaving an agent guessing about timeout, resource class, and working directory.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is only 29%, so the description needs to compensate, but it only restates waitMs and its default (already in the schema). cwd, title, timeoutMs, resourceClass, and resourceWaitMs receive no explanation, and the 50s cap on waitMs is left to the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (run a long command as a background job) and enumerates concrete use cases (test suite, build, install, benchmark). It doesn't explicitly contrast with the similarly named process_start/run_command/pty_start siblings, so an agent must infer the distinction.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear usage context: use for long-running commands, then wait up to waitMs, otherwise poll job_status. The fallback path to a named sibling is explicit. It lacks guidance on when to prefer run_command or process_start over a background job.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools