Skip to main content
Glama
AIops-tools

io.github.AIops-tools/olvm-aiops

Official

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault
OLVM_AIOPS_HOMENoOptional directory for audit.db and other runtime data, overriding default ~/.olvm-aiops.
OLVM_RUNAWAY_MAXNoOptional tight‑loop circuit breaker threshold for runaway guard.
OLVM_AIOPS_CONFIGNoOptional path to config.yaml, overriding default ~/.olvm-aiops/config.yaml.
OLVM_MAX_TOOL_CALLSNoOptional cumulative call cap for runaway guard.
OLVM_MAX_TOOL_SECONDSNoOptional cumulative wall-time cap for runaway guard.
OLVM_AUDIT_APPROVED_BYNoOptional operator-supplied approver/rationale recorded in audit entries.
OLVM_AIOPS_MASTER_PASSWORDYesMaster password to unlock the encrypted secrets store. Required for the MCP server to access credentials; without it tools return a teaching error.

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": false
}
prompts
{
  "listChanged": false
}
resources
{
  "subscribe": false,
  "listChanged": false
}
experimental
{}

Tools

Functions exposed to the LLM to take actions

NameDescription
datacenter_listA

[READ] Data centers with status (up, uninitialized, maintenance…) and compat version.

Args: limit: Rows to return, 1-1000 (default 100); truncated says when more exist. target: Engine target name from config; omit to use the default.

cluster_listA

[READ] Clusters with compatibility version, CPU type and memory over-commit policy.

Args: limit: Rows to return, 1-1000 (default 100). target: Engine target name from config; omit to use the default.

host_listA

[READ] KVM hosts with status, SPM role, memory and VM counts.

A field the engine did not report is null, never 0. For "what is wrong with my hosts" use host_health_rca instead of reading this list yourself.

Args: limit: Rows to return, 1-1000 (default 100); truncated says when more exist. search: Engine query language, e.g. "status=up" or "status!=up". target: Engine target name from config; omit to use the default.

host_getB

[READ] One KVM host by id (see host_list).

Args: host_id: Host id. target: Engine target name from config; omit to use the default.

event_listA

[READ] Engine events at or above a severity, newest first.

To follow new events, pass the highest index you have already seen as after_index: events after it come back OLDEST first (order), so repeat with the highest index returned while truncated is true and nothing is skipped. since_minutes keeps events from the last N minutes and sets scanTruncated when older ones in the window may be missing. SSO session ids in login events are redacted.

Args: limit: Rows to return, 1-1000 (default 100); truncated says when more exist. min_severity: One of normal, warning, error, alert (default normal). page: Older pages of events (not combinable with since_minutes or after_index). after_index: Only events with a higher index than this, oldest first. since_minutes: Only events from the last N minutes (1-129600). target: Engine target name from config; omit to use the default.

job_listA

[READ] Engine jobs (long-running operations) with status and duration.

Args: limit: Rows to return, 1-1000 (default 100). status: Only one status: started, finished, failed, aborted, unknown. target: Engine target name from config; omit to use the default.

host_health_rcaA

[READ] What needs attention on KVM hosts, ranked worst first, in one call.

Combines host status and status detail, reinstall/update flags and recent warning-or-worse events that name a host. Each finding has signal (what was measured), cause, action and rank. Hosts that the engine is installing or rebooting are reported as in progress, not failed; alert 9000 (power management not verifiable) is informational on hosts without fencing hardware. A failed guest-agent call (event 10802, a command starting with VmLogon or VmLogoff whose message names the guest agent) is a problem of a guest, not of the host that ran the call: it is reported low and carries vmCandidates, the VMs the engine reports on that host. The same command failing for another reason stays a host finding. The event names no VM, so present those as candidates to check, never as the affected VM; if that envelope has an error, or scanTruncated is true, say the candidate list is incomplete. Every event-10802 finding reports one vdsm command rather than the host's own state, so its cause names the command; do not restate it as "this host is failing" unless the host's status, external status or flags say so. It also carries relatedVmEvents: the VM-level failures the engine logged within windowSeconds on this host, which is where the operation's own failure is recorded (vm_health_rca reports that half). They are paired on time alone — say "possibly the same operation", never that the host finding is about a named VM; an empty list means none was logged, not that none was looked for. Report findings in rank order and quote their signal.

Args: events_limit: Recent warning-or-worse events to correlate, 1-1000 (default 200). events_window_hours: Ignore events older than this many hours, 1-720 (default 24); older ones are counted in eventsOutsideWindow. target: Engine target name from config; omit to use the default.

engine_health_rcaA

[READ] Problems with the engine itself, ranked worst first, in one call.

Reads the engine's health check, its version, its clock against this machine and its summary counts, plus warning-or-worse events that name no host, VM or storage domain: engine or CA certificate expiry, missing or failed engine backups, a cluster failing its HA reservation. Those alerts appear in no other diagnosis, so call this first for "is anything wrong". Event 10803 (a storage-pool vdsm command failed) names no host and lands here: its cause names the command, which is the subject — not the engine — and its severity is the event's own. Event 2024 is a record that someone ran unlock_entity.sh by hand: it is low, and whatever the unlocked entity then did wrong is reported by its own event. Report findings in rank order and quote their signal.

Args: events_limit: Recent warning-or-worse events to correlate, 1-1000 (default 200). events_window_hours: Ignore events older than this many hours, 1-720 (default 24). target: Engine target name from config; omit to use the default.

storage_domain_listA

[READ] Storage domains with status and capacity (free, used, committed bytes and %).

The status of an attached domain comes from its data center (statusSource); it is null — not guessed — when that could not be read (statusErrors).

Args: limit: Rows to return, 1-1000 (default 100); truncated says when more exist. target: Engine target name from config; omit to use the default.

storage_domain_getA

[READ] One storage domain by id, with its data-center-scoped status.

Args: domain_id: Storage domain id (see storage_domain_list). target: Engine target name from config; omit to use the default.

vm_listA

[READ] VMs with status, host, vCPUs, memory and high-availability flag.

Args: limit: Rows to return, 1-1000 (default 100); truncated says when more exist. search: Engine query language, e.g. "status=up" or "cluster=Default". target: Engine target name from config; omit to use the default.

vm_getA

[READ] One VM by id (see vm_list).

Args: vm_id: VM id. target: Engine target name from config; omit to use the default.

vm_statsC

[READ] A VM's current statistics: memory, CPU %, network and disk usage, with units.

Args: vm_id: VM id. target: Engine target name from config; omit to use the default.

storage_capacity_rcaA

[READ] Storage-domain problems ranked worst first, in one call.

Critical when free space is below the engine's critical blocker (the engine then refuses new disks and snapshots); medium below the low-space warning, or when thin disks are committed beyond capacity while space is low (low on its own); high when an attached domain is inactive, unknown or mixed. Warning-or-worse events naming a storage domain (for example "deactivated by system") are attached to it. Unattached domains such as the default image repository are skipped. A domain whose thresholds are unset (the engine reports 0) is reported low with its measured free space: nothing — not the engine, not this tool — will warn before it fills, and how full is too full is not recorded anywhere to read. Over-commit on its own is a planning limit, and its signal carries actual use next to it: do not report an over-committed domain as short of space unless a low-space or blocker finding says so. Report findings in rank order and quote their signal.

Args: events_limit: Recent warning-or-worse events to correlate, 1-1000 (default 200). events_window_hours: Ignore events older than this many hours, 1-720 (default 24). target: Engine target name from config; omit to use the default.

vm_health_rcaA

[READ] VM problems ranked worst first, in one call.

High for VMs stuck not_responding/unknown, paused (often storage I/O errors or a full domain — check storage_capacity_rca), or down with high availability; medium for image_locked; info for VMs mid-transition; low for config changes waiting for a restart. Recent warning-or-worse events naming a VM are attached to it, one finding per VM and code; an event is superseded (info) when the VM started after it, or when the engine reported it back up and it is up now. Down VMs without HA are normal. Report findings in rank order and quote their signal.

Args: events_limit: Recent warning-or-worse events to correlate, 1-1000 (default 200). events_window_hours: Ignore events older than this many hours, 1-720 (default 24); older ones are counted in eventsOutsideWindow. target: Engine target name from config; omit to use the default.

undo_listA

[READ] List recorded, not-yet-applied undo tokens (most recent first).

Each entry names the original tool, the inverse tool that undo_apply would run, and a human note. Use the undoId with undo_apply.

One extra row is read so truncated is measured rather than guessed from the row count matching the limit; when it is true, older tokens exist.

Each entry carries effectVerified. False means the original write lost its response, so the change it reverses is PROBABLE, not confirmed — check the live state before applying, and do not report the result as a restore of a state that may never have been reached.

Args: limit: Max rows to return (default 50, capped at 500). target: Unused (undo state is host-local); accepted for CLI uniformity.

undo_applyA

[WRITE][risk=medium] Apply a recorded undo by dispatching its inverse tool.

The inverse runs through its own governed tool, so it is audited under its own risk tier. Pass dry_run=True to preview the inverse call without executing it. A token can only be applied once.

Args: undo_id: The undoId from undo_list (or an _undo_id in a write result). dry_run: If True, preview the inverse tool + params without running it. target: Passed through to the inverse tool when it accepts a target.

Prompts

Interactive templates invoked by user choice

NameDescription

No prompts

Resources

Contextual data attached and managed by the client

NameDescription

No resources

TDQS

A3.9/5.0

Scored across 17 tools

Disambiguation5/5

Each tool targets a distinct resource and action: list/get for inventory, *_rca for ranked health findings, event/job lists for raw operations. The rca tools are explicitly cross-referenced (host_list points to host_health_rca), so an agent can choose without guessing.

Naming Consistency4/5

Names follow a consistent snake_case resource_action pattern with list/get/rca/stats suffixes, and the rca suffix clearly marks health diagnostics. Minor deviations like vm_stats and storage_capacity_rca vs storage_domain_* are easy to predict but break the pure resource_verb symmetry.

Tool Count4/5

17 tools is slightly above the ideal 3-15 range, but each covers a meaningful read or diagnostic purpose with no obvious duplication. The set is a bit heavy for a read-only server, but not bloated enough to be confusing.

Completeness4/5

The toolset covers the main oVirt resources (datacenters, clusters, hosts, storage, VMs) plus raw events/jobs and targeted health RCAs, so common troubleshooting paths are supported. Notable minor gaps are no cluster-level health RCA and no datacenter/cluster get, but list tools and engine RCA fill most practical needs.

Maintenance

ActivityMaintained
ResponsivenessResponsive