Skip to main content
Glama
CodeMonk6

RISBridge MCP

by CodeMonk6

Multi-GPU torch job

ris_submit_multigpu_torch_job

Submit single-node multi-GPU PyTorch training jobs to Slurm using torchrun, with H100 defaults and dry-run validation before confirmation.

Instructions

Single-node multi-GPU PyTorch training via torchrun (typed H100 by default).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
argsNo
cpusNo
memGbNo
dryRunNoIf true (default) build+validate and return the confirmation summary; DO NOT submit.
scriptYes
accountNo
confirmNoMust be true (with dryRun=false) to actually submit.
gpuTypeNoH100
jobNameNoJob name; defaults to the project + job type.
modulesNo
profileNoNamed profile from ~/.risbridge-mcp/config.json. Omit for the default.
projectYesProject name; becomes one directory under the storage workspace.
condaEnvNo
gpuCountNo
walltimeNo01:00:00
partitionNo
nprocPerNodeNo
allowUntypedGpuNoUse untyped --gres=gpu:N (required/auto on the preempt partition).

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations declare destructiveHint=false and openWorldHint=true, so the agent knows this is a non-destructive write that hits external systems. But the description omits the critical safety workflow: dryRun defaults to true and confirm must be true with dryRun=false to actually submit. That two-step confirmation gate is a major behavioral trait conveyed only by inline param descriptions, not the tool description. No mention of auth requirements, scheduling semantics, or consequences of submission.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

A single tightly-packed sentence with no filler. Front-loads the operation, scope, and default in a scannable form.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For an 18-parameter submission tool with no output schema and a non-trivial two-step confirm workflow, the description is far too thin. It doesn't warn about the dryRun/confirm contract, doesn't clarify that the default only permits validation, and provides no scheduling guidance. The agent is left to discover the submission gate from individual parameter descriptions.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 33% (6 of 18 params documented inline), so baseline 3 applies. The description adds one default value (H100) but explains nothing else about the many parameters (gpuCount, partition, walltime, modules, condaEnv, etc.). It does not compensate for the low coverage; the schema is the primary documentation here.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb (submit), resource (multi-GPU PyTorch training job), and mechanism (torchrun, typed H100 by default). It distinguishes itself from sibling ris_submit_python_job, ris_submit_r_job, etc. via the multi-GPU/torchrun framing, though it doesn't name a sibling directly.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No when-to-use guidance, no exclusions, no alternatives named. An agent must infer from the name that this is for multi-GPU PyTorch specifically versus ris_submit_python_job or ris_submit_gpu_smoke_test. No mention of the dryRun/confirm submission workflow that the schema implies is mandatory.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.