Skip to main content
Glama

job_system_health

Read-only

Diagnose a running Slurm job by comparing CPU load to reserved cores, I/O wait, and memory use to reservation. Flags single-core parallel code, filesystem waits, and oversized memory reservations.

Instructions

Sante systeme d'un job EN COURS, GPU ou non : charge processeur face aux coeurs reserves, attente d'entrees-sorties, memoire utilisee face a la reservation. Detecte les trois gaspillages classiques : un code cense etre parallele qui tourne sur un seul coeur, un calcul qui passe son temps a attendre le systeme de fichiers, et une reservation memoire massivement surdimensionnee.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
job_idYes

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv1.4.0

TDQS

A3.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, so the safety profile is covered. The description adds substantive behavioral context beyond that: it discloses the exact comparison metrics and the three waste patterns it detects, which tells the agent what kind of analysis to expect. It does not discuss permissions, cost, or interpretation, but with annotations present the extra specificity is valuable.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is front-loaded with the core purpose and then efficiently lists the diagnostic signals and waste patterns in one dense sentence. There is no filler, though the colon-and-list structure is slightly long for the amount of routing information provided.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so the description need not explain return values. It covers the tool's purpose, scope, and the specific diagnostics it performs, which is sufficient for a one-parameter read-only health check. The only missing piece is any parameter-level detail, but that is a small gap given the schema.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% and there is one required parameter (job_id). The description never mentions the parameter, its expected format, or its role. While the name 'job_id' is largely self-explanatory, the description does not compensate for the missing schema documentation as required when coverage is low.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific resource ('santé système d'un job EN COURS') and enumerates the exact signals it inspects (CPU load vs reserved cores, I/O wait, memory used vs reserved) plus the three classic wastes it detects. It clearly distinguishes a running-job health check from a general status or efficiency report, though it does not explicitly contrast itself with sibling tools like job_efficiency or job_live_metrics.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly scopes usage to a running job ('EN COURS'), which gives a clear context for when to call it. It adds that the job may be GPU or non-GPU. However, it offers no explicit when-not guidance or alternative routing against siblings such as job_efficiency or diagnose_job.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.