Skip to main content
Glama

Triage a Problem

diagnose
Read-onlyIdempotent

Triage a "something is wrong" report on Cycle. Give it the environment, container, and/or server the user suspects; it fans out the platform's high-signal diagnostic reads concurrently and returns findings ranked critical > warning > info, each naming the follow-up tool (with arguments) to run next. Always start here when the problem is vague, then drill down with the suggested tools.

Problem-class guide — pass as focus, or omit to run every check the scope allows:

  • Container or VM unreachable from its URL → focus "unreachable" with environment (+ container). Covers virtual machines as fully as containers — a LINKED record can point at either. Cross-checks three independent layers so a pass at one never masks a failure at another: (1) ingress CONFIG — LINKED records vs the target's public network and port mappings against the LB's actual controller config, DMZ records included; (2) DNS RESOLUTION — each linked domain against its zone's authoritative nameservers AND public resolvers, which is the only way to see an unpropagated or address-less record; (3) an END-TO-END synthetic HTTP GET at each linked domain (max 5, from the MCP server's network) — probe.http.ok means VERIFIED SERVING, and 502-504 usually means the app is bound to localhost or IPv4-only instead of :: . Also covers LB/gateway/discovery service health, instance readiness, and LB destination errors. A target can be fully healthy yet unreachable — config mismatches are the most common cause; every finding carries a machine-readable code and explains itself.

  • Container stopped, crashing, or restarting → focus "crashloop" with container. Checks state drift, broken instances, restart/healthcheck events, and error-pattern logs.

  • Out of disk / containers can't write → focus "storage" with server (or container — a "no space left" log traces to its host). Checks storage pool and mount utilization, storage-full events.

  • CPU/RAM running low → focus "resources" with server or container. Checks load vs cores, RAM headroom, allocation pressure, instance OOM/throttling.

  • Containers not talking to each other → focus "networking" with environment (+ server for mesh problems). Checks discovery/VPN services, mesh and neighbor events; suggests the neighbor_latency metrics preset.

  • Stack build stuck or failed, deploy never produced containers → focus "build" with stack (+ build id), or with a container deployed from the stack. Reads the build's state and error and every image build it contains, attaching the tail of each failed build log as evidence.

  • LB weirdness / 502s → focus "load_balancer" with environment. Checks LB service state, DNS-record-to-container port routing, disconnect reasons, per-destination response codes, plus the synthetic HTTP probe above. LB telemetry lags several minutes and 404s when absent — that never means the LB itself is gone.

A finding of category "platform" means a diagnostic read itself errored or could not be attempted — run the suggested tool manually. healthy=true means no warning-or-worse findings in the window (default: the last hour). Anything that could NOT be verified (an unprobed ingress path, unreachable DNS) is reported as a warning so healthy is never true by omission. Read-only: this never changes anything. For in-container investigation afterwards, use run_instance_command; for a VM guest, run_vm_command; for live output, capture_stream.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
buildNoSpecific stack build ID to inspect instead of the latest. Requires stack.
focusNoProblem class to prioritize; omit to run every check applicable to the scope.
stackNoStack to diagnose: name, identifier, or ID. Checks its latest build and that build's images; a container deployed from a stack implies this.
serverNoServer to diagnose: hostname, nickname, or ID. Use for storage-full, resource exhaustion, or host-down suspicions.
contextNoWhy are you calling this tool? Briefly describe the user's goal.
containerNoContainer to diagnose. Best for crash loops and won't-start problems; its environment is diagnosed too.
environmentNoEnvironment to diagnose. Use for app-level problems: unreachable URLs, 502s, containers not talking.
conversation_idNoConversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation.
lookback_minutesNoHow far back to scan events, logs, and telemetry, 5-1440 minutes.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover the safety profile (readOnly, idempotent, non-destructive), but the description adds substantial behavior beyond them: findings are ranked, each names its follow-up tool with arguments, 'platform' category means a diagnostic read itself failed, healthy=true only means no warning-or-worse in the window, unverified paths are downgraded to warnings so healthy is never true by omission, and the default lookback is the last hour.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The critical guidance is front-loaded in the first two sentences, and the bulleted problem-class guide is structured and scannable. It is long and dense, but for a 9-parameter tool with seven focus classes almost every clause carries decision-relevant content, so the length is largely earned rather than padding.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description explains the return shape (ranked findings with machine-readable codes, a healthy flag, a conversation_id line to thread back), the read-only guarantee, and how to interpret failure/omission cases. For a broad triage tool with 9 optional params, nothing an agent needs to call it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% (baseline 3), but the description adds real meaning: which scope parameter to pair with each focus class, that build requires stack, that a container implies its stack, when to use server vs container vs environment, and that telemetry probes are capped at 5 domains. This is genuinely beyond the schema text.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening sentence states a precise verb+resource ('Triage a "something is wrong" report on Cycle') and immediately describes the mechanism: concurrent high-signal diagnostic reads with findings ranked critical > warning > info. It also names the sibling tools it defers to for drill-down (run_instance_command, run_vm_command, capture_stream), so an agent can place it against alternatives without opening a schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicit when-to-use ('Always start here when the problem is vague, then drill down with the suggested tools') plus a full problem-class guide that maps each focus value to the triggering symptom and the required scope parameter. It even notes the omission case ('omit to run every check the scope allows'), leaving nothing to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources