engine_health_rca
Identify engine-level problems by checking health, version, clock, backups, certificate expiry, and HA reservations, then rank correlated warning events worst-first for root-cause analysis.
Instructions
[READ] Problems with the engine itself, ranked worst first, in one call.
Reads the engine's health check, its version, its clock against this machine and
its summary counts, plus warning-or-worse events that name no host, VM or storage
domain: engine or CA certificate expiry, missing or failed engine backups, a
cluster failing its HA reservation. Those alerts appear in no other diagnosis, so
call this first for "is anything wrong". Event 10803 (a storage-pool vdsm command
failed) names no host and lands here: its cause names the command, which is the
subject — not the engine — and its severity is the event's own. Event 2024 is a record
that someone ran unlock_entity.sh by hand: it is low, and whatever the unlocked
entity then did wrong is reported by its own event. Report findings in rank order and
quote their signal.
Args: events_limit: Recent warning-or-worse events to correlate, 1-1000 (default 200). events_window_hours: Ignore events older than this many hours, 1-720 (default 24). target: Engine target name from config; omit to use the default.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| target | No | ||
| events_limit | No | ||
| events_window_hours | No |