Skip to main content
Glama
AIops-tools

io.github.AIops-tools/olvm-aiops

Official

host_health_rca

Identify and prioritize KVM host issues by combining status, update flags, and recent warning events. Get ranked findings with signal, cause, and action to guide remediation.

Instructions

[READ] What needs attention on KVM hosts, ranked worst first, in one call.

Combines host status and status detail, reinstall/update flags and recent warning-or-worse events that name a host. Each finding has signal (what was measured), cause, action and rank. Hosts that the engine is installing or rebooting are reported as in progress, not failed; alert 9000 (power management not verifiable) is informational on hosts without fencing hardware. A failed guest-agent call (event 10802, a command starting with VmLogon or VmLogoff whose message names the guest agent) is a problem of a guest, not of the host that ran the call: it is reported low and carries vmCandidates, the VMs the engine reports on that host. The same command failing for another reason stays a host finding. The event names no VM, so present those as candidates to check, never as the affected VM; if that envelope has an error, or scanTruncated is true, say the candidate list is incomplete. Every event-10802 finding reports one vdsm command rather than the host's own state, so its cause names the command; do not restate it as "this host is failing" unless the host's status, external status or flags say so. It also carries relatedVmEvents: the VM-level failures the engine logged within windowSeconds on this host, which is where the operation's own failure is recorded (vm_health_rca reports that half). They are paired on time alone — say "possibly the same operation", never that the host finding is about a named VM; an empty list means none was logged, not that none was looked for. Report findings in rank order and quote their signal.

Args: events_limit: Recent warning-or-worse events to correlate, 1-1000 (default 200). events_window_hours: Ignore events older than this many hours, 1-720 (default 24); older ones are counted in eventsOutsideWindow. target: Engine target name from config; omit to use the default.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
targetNo
events_limitNo
events_window_hoursNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full behavioral burden and does so thoroughly. It discloses result fields (signal, cause, action, rank), ordering requirements, edge cases (in-progress vs failed, alert 9000, event 10802), incomplete-candidate behavior, and the time-only pairing semantics of relatedVmEvents. This goes far beyond a basic 'gets host health' statement.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Although long, the description is densely informative and every sentence earns its place. It front-loads the primary purpose, then systematically covers result shape, interpretation rules, edge cases, and parameter semantics without filler. The length is justified by the tool's complexity and lack of an output schema.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's complexity, the absence of annotations, and the absence of an output schema, the description is complete enough for correct invocation and interpretation. It defines finding fields, ranking, special-case semantics, parameter effects, and how to report candidate VMs, leaving no major gap for an agent deciding whether and how to call it.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema provides only names, types, defaults, and no descriptions, so the description's Args section is essential. It explains the meaning and ranges of events_limit and events_window_hours, notes side effects like eventsOutsideWindow, and clarifies that target refers to the engine target from config. All three parameters receive meaningful semantic context.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a clear purpose: it identifies what needs attention on KVM hosts, ranked worst first, in one call. It specifies the resource (KVM hosts), the output orientation (ranked findings), and the combined data sources (status, reinstall/update flags, recent events), which distinguishes it from single-purpose siblings like host_get and event_list.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit interpretation rules and routes to a sibling: VM-level failures are said to be covered by vm_health_rca, and event-10802 guest-agent failures are explicitly excluded from being treated as host problems. It also tells the agent when not to restate a finding as 'this host is failing' and how to treat in-progress hosts and informational alert 9000, giving clear when/when-not guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.