Root cause
root_causeFind which service broke first and why during an incident by comparing errors, traffic and latency to a baseline, then ranking causes with timeline evidence.
Instructions
Find which service broke first and why. Compares every service's errors, traffic and latency with a baseline, pins the first error of each to the millisecond, detects deploys/restarts/host rollouts from the logs, infers the call graph from traces, and returns a ranked verdict with a timeline and evidence. Start here for 'what is causing this incident?'.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| query | No | Optional Lucene filter for scope, e.g. 'env:prod' | |
| range | No | Window to analyse (the incident), e.g. '1h', '30m' | 1h |
| streams | No | Stream titles or ids to search in; all streams when omitted | |
| to_time | No | Absolute end, same formats as from_time; default now | |
| baseline | No | Length of the normal period right before the window; same as the window by default | |
| instance | No | Graylog instance (environment) from list_instances, e.g. 'staging' or 'prod'; the default instance when omitted | |
| from_time | No | Absolute start: ISO 8601 or 'YYYY-MM-DD HH:MM' in the instance timezone; overrides range |