Triage a Problem
diagnoseTriage a "something is wrong" report on Cycle. Give it the environment, container, and/or server the user suspects; it fans out the platform's high-signal diagnostic reads concurrently and returns findings ranked critical > warning > info, each naming the follow-up tool (with arguments) to run next. Always start here when the problem is vague, then drill down with the suggested tools.
Problem-class guide — pass as focus, or omit to run every check the scope allows:
Container or VM unreachable from its URL → focus "unreachable" with environment (+ container). Covers virtual machines as fully as containers — a LINKED record can point at either. Cross-checks three independent layers so a pass at one never masks a failure at another: (1) ingress CONFIG — LINKED records vs the target's public network and port mappings against the LB's actual controller config, DMZ records included; (2) DNS RESOLUTION — each linked domain against its zone's authoritative nameservers AND public resolvers, which is the only way to see an unpropagated or address-less record; (3) an END-TO-END synthetic HTTP GET at each linked domain (max 5, from the MCP server's network) — probe.http.ok means VERIFIED SERVING, and 502-504 usually means the app is bound to localhost or IPv4-only instead of :: . Also covers LB/gateway/discovery service health, instance readiness, and LB destination errors. A target can be fully healthy yet unreachable — config mismatches are the most common cause; every finding carries a machine-readable code and explains itself.
Container stopped, crashing, or restarting → focus "crashloop" with container. Checks state drift, broken instances, restart/healthcheck events, and error-pattern logs.
Out of disk / containers can't write → focus "storage" with server (or container — a "no space left" log traces to its host). Checks storage pool and mount utilization, storage-full events.
CPU/RAM running low → focus "resources" with server or container. Checks load vs cores, RAM headroom, allocation pressure, instance OOM/throttling.
Containers not talking to each other → focus "networking" with environment (+ server for mesh problems). Checks discovery/VPN services, mesh and neighbor events; suggests the neighbor_latency metrics preset.
Stack build stuck or failed, deploy never produced containers → focus "build" with stack (+ build id), or with a container deployed from the stack. Reads the build's state and error and every image build it contains, attaching the tail of each failed build log as evidence.
LB weirdness / 502s → focus "load_balancer" with environment. Checks LB service state, DNS-record-to-container port routing, disconnect reasons, per-destination response codes, plus the synthetic HTTP probe above. LB telemetry lags several minutes and 404s when absent — that never means the LB itself is gone.
A finding of category "platform" means a diagnostic read itself errored or could not be attempted — run the suggested tool manually. healthy=true means no warning-or-worse findings in the window (default: the last hour). Anything that could NOT be verified (an unprobed ingress path, unreachable DNS) is reported as a warning so healthy is never true by omission. Read-only: this never changes anything. For in-container investigation afterwards, use run_instance_command; for a VM guest, run_vm_command; for live output, capture_stream.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| build | No | Specific stack build ID to inspect instead of the latest. Requires stack. | |
| focus | No | Problem class to prioritize; omit to run every check applicable to the scope. | |
| stack | No | Stack to diagnose: name, identifier, or ID. Checks its latest build and that build's images; a container deployed from a stack implies this. | |
| server | No | Server to diagnose: hostname, nickname, or ID. Use for storage-full, resource exhaustion, or host-down suspicions. | |
| context | No | Why are you calling this tool? Briefly describe the user's goal. | |
| container | No | Container to diagnose. Best for crash loops and won't-start problems; its environment is diagnosed too. | |
| environment | No | Environment to diagnose. Use for app-level problems: unreachable URLs, 502s, containers not talking. | |
| conversation_id | No | Conversation tracking id. Omit on your first tool call; every result then includes a conversation_id line — pass that exact value on all later calls in this conversation. | |
| lookback_minutes | No | How far back to scan events, logs, and telemetry, 5-1440 minutes. |