Skip to main content
Glama

run_cluster_sanity_check

Read-only

Check GPU health across cluster nodes for thermal/power throttling, uncorrected memory errors, and low clocks. Get a ready-to-use --exclude list for suspect nodes to prevent silent job slowdowns.

Instructions

Verifie la sante des GPU sur un ou plusieurs noeuds : bridage thermique ou de puissance, erreurs memoire non corrigees, frequence anormalement basse. Un noeud degrade ne plante pas, il ralentit tout un job reparti sans erreur visible. Rend une clause --exclude prete a l'emploi pour les noeuds suspects.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
nodesNo
minutesNo
max_nodesNo
check_typeNogpu

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv1.4.0

TDQS

A4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

readOnlyHint=true already establishes the safe-read profile, so the description is not obligated to restate safety. It adds genuine behavioral context: the failure mode is silent slowdown rather than a crash, and the output is a ready-to-use --exclude clause for suspect nodes. It does not mention runtime cost or how long a sanity sweep takes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three tight sentences with no filler: what is checked, why it matters, and what comes back. The important scoping information is front-loaded before the rationale.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return-value detail is not required, though the description helpfully previews the --exclude output. The gap is on the input side: with zero schema description coverage and undocumented parameters like check_type and minutes, an agent cannot confidently set non-default values.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across four parameters, so the description must compensate and largely does not. 'Un ou plusieurs noeuds' loosely maps to the nodes parameter, but minutes, max_nodes, and check_type (whose enumeration of allowed values is undisclosed) are never explained.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (verifie/verify) and a precisely scoped resource: GPU health on one or more nodes, enumerating the exact conditions checked (thermal/power throttling, uncorrected memory errors, abnormal frequency). This distinguishes it from general health siblings like job_system_health or romeo_selfcheck, and signals its Slurm-oriented output (--exclude clause).

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explains the motivating scenario clearly: a degraded node does not crash but silently slows a distributed job, which tells the agent to run this before submitting spread work. No explicit 'use X instead of Y' routing versus siblings such as diagnose_job or job_system_health, so it stops short of a 5.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.