Skip to main content
Glama

submit_resilient_job

Submit long computations as chained, resumable Slurm segments that checkpoint before timeout and restart from checkpoint_dir, overcoming partition time limits.

Instructions

Soumet un calcul long sous forme de chaine de segments reprenables, pour depasser la limite de temps d'une partition rapide. Chaque segment recoit SIGUSR1 avant son expiration pour sauvegarder, et le suivant demarre apres lui via une dependance, en reprenant du dernier point de sauvegarde. Un marqueur de fin fait sauter les segments restants si le calcul se termine avant terme. Ton code doit savoir reprendre depuis checkpoint_dir et, idealement, traiter SIGUSR1.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
archNo
nameYes
mem_gbNo
commandYes
confirmNo
workdirNo
segment_timeNo1h
cpus_per_taskNo
gpus_per_nodeNo
signal_beforeNo
stage_archiveNo
checkpoint_dirNo
max_total_timeNo6h
spack_packagesNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv1.4.0

TDQS

A3.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only cover non-readOnly and non-destructive status, while the description adds rich runtime behavior: segment chain execution, SIGUSR1 before expiration, dependency-based continuation, resume from checkpoint_dir, end-marker skipping, and the requirement that user code handle resumes and ideally SIGUSR1. This is exactly the operational detail annotations cannot convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Four sentences, front-loaded with purpose followed by the execution mechanism, with no obvious filler. It is appropriately sized for the complex execution model, though it could better organize the parameter implications.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 14-parameter submission tool with no schema descriptions, the description omits most input parameters and any alternative-tool guidance. Output schema exists so return values need not be explained, but invocation-critical parameter context is largely absent.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across 14 parameters, so the description must carry parameter meaning. It only mentions checkpoint_dir and links SIGUSR1 to the signal mechanism, while omitting name, command, segment_time, max_total_time, cpus_per_task, gpus_per_node, confirm, workdir, arch, mem_gb, stage_archive, and spack_packages.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource: submits a long computation as a chain of resumable segments to exceed a fast partition's time limit. It does not name or contrast with sibling tools like submit_job or submit_array_job, so sibling differentiation is left to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear context for use: long computations that need to exceed a fast partition's time limit via checkpointed segments. No explicit when-not condition or alternative tool is named, but the intended scenario is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.