Skip to main content
Glama

job_stack_trace

Read-only

Captures call stacks from a running Slurm job's processes to diagnose hangs like MPI deadlocks, frozen CUDA kernels, or lock waits without interrupting the computation.

Instructions

Capture la pile d'appels des processus d'un job EN COURS, pour diagnostiquer un blocage : interblocage MPI, noyau CUDA fige, attente sur verrou. Utilise pstack, avec repli sur gdb puis eu-stack. Ne s'attache pas de maniere interactive : il preleve une trace et rend la main, ce qui n'interrompt pas le calcul.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
job_idYes
process_nameNo
max_processesNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv1.4.0

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the readOnlyHint annotation, the description discloses the fallback chain (pstack -> gdb -> eu-stack), that attachment is non-interactive, and crucially that the sampling does not interrupt the running computation. That is exactly the behavioral context an agent needs before touching a live job.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loads the purpose, then the mechanism, then the non-intrusive guarantee in a few dense sentences with no filler. Slightly information-heavy for its length but every clause earns its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present and readOnlyHint covering the safety profile, the description is largely complete on purpose and behavior. The one real gap is parameter semantics, which nothing else fills since schema coverage is 0%.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% for all three parameters (job_id, process_name, max_processes), so the description must carry the burden and does not. It never mentions that a running job_id is required, what process_name filters, or what max_processes limits (default 3).

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Capture la pile d'appels des processus d'un job EN COURS') and scopes it to in-progress jobs, immediately separating it from the many job-status siblings. The stated diagnostic intent (interblocage MPI, noyau CUDA fige, attente sur verrou) makes the tool's niche unmistakable.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives a clear triggering context: use it when a job is blocked/hung, with three concrete hang categories enumerated. It does not name an alternative sibling (e.g., diagnose_job) or state when NOT to reach for it, so it stops short of explicit routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.