Skip to main content
Glama
vmware-skills

VMware-Monitor

cluster_health_summary

Read-onlyIdempotent

Get an at-a-glance health summary for every cluster, scoring CPU, memory, VM power state, and alarms as ok/warn/critical and ranking top issues for quick triage.

Instructions

[READ] One-glance health rollup for every cluster — "is anything on fire?".

Start here for single-vCenter triage. Batches hosts, VM power state, live CPU/memory pressure and alarms per cluster, scores each "ok"/"warn"/ "critical", and ranks the anomalies into top_issues. Use this instead of stitching list_all_clusters + list_esxi_hosts + get_alarms yourself.

Returns {totals, top_issues, issues_total, clusters, snapshot, customization_hint} — not the list envelope. Lead with top_issues (worst first), show clusters as context, always echo customization_hint last. Point-in-time — no trending. top_issues includes datastores thin-provisioned past 100% of capacity (kind capacity, scope datastore), attributed to the cluster of a host that mounts them; datastore_capacity has the full table. Alarm issues carry condition_now and acknowledged_days. Only a cleared alarm (its condition is read to be false now) ranks after live issues; an unknown one is not re-checked and keeps its severity rank, however long ago it was acknowledged — report it as possibly still live.

Then drill into what top_issues names with vm_investigation_bundle, host_investigation_bundle or datastore_investigation_bundle; use cross_vcenter_attention to cover every target at once. Acting on a finding belongs to vmware-aiops.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
top_nNoCap ``top_issues`` (default 10; 0 omits it). ``issues_total`` is the pre-cap count.
targetNovCenter/ESXi target from config (default if omitted).
include_vmsNoRoll up VM counts (default True); False skips that pass.
cluster_filterNoCase-insensitive substring; only matching clusters show (None = all, plus a standalone-hosts row).

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed5 schema fields changedv1.9.2
    • addedInput schema / additionalProperties
      Added value: +false
    • addedInput schema / properties / cluster_filter / description
      Added value: +"Case-insensitive substring; only matching clusters show (None = all, plus a standalone-hosts row)."
    • addedInput schema / properties / include_vms / description
      Added value: +"Roll up VM counts (default True); False skips that pass."
    • addedInput schema / properties / target / description
      Added value: +"vCenter/ESXi target from config (default if omitted)."
    • addedInput schema / properties / top_n / description
      Added value: +"Cap ``top_issues`` (default 10; 0 omits it). ``issues_total`` is the pre-cap count."
  2. Addedv1.7.6

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, idempotentHint=true, and destructiveHint=false, and the description is consistent ('[READ]'). It adds substantial behavioral context beyond annotations: the return is 'not the list envelope', it is 'Point-in-time — no trending', and it details the ranking semantics (cleared alarms rank after live issues; unknown alarms keep severity rank and are reported as possibly still live). No contradiction with annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence earns its place — purpose is front-loaded in the first line, followed by usage, return format, ranking behavior, and downstream tool routing. The density is justified by the tool's complexity, though it could be trimmed slightly without losing essential information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description correctly carries the full burden of explaining return values: it lists the fields ({totals, top_issues, issues_total, clusters, snapshot, customization_hint}) and how to present them, plus the inclusion rule for thin-provisioned datastores and alarm condition fields. Nothing an agent needs to call it correctly or interpret results is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds some value by explaining how top_n caps top_issues and that issues_total is the pre-cap count, and how include_vms affects the rollup. However, most parameter meaning is already fully covered by the schema descriptions, so the description doesn't dramatically compensate beyond the baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a specific verb-resource pair ('health rollup for every cluster') and the goal ('is anything on fire?'). It explicitly differentiates from siblings by saying 'Use this instead of stitching list_all_clusters + list_esxi_hosts + get_alarms yourself,' so an agent can tell it apart from the list/detail tools without opening schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use guidance ('Start here for single-vCenter triage') and names alternatives ('Then drill into what top_issues names with vm_investigation_bundle, host_investigation_bundle or datastore_investigation_bundle; use cross_vcenter_attention to cover every target at once'). Also gives exclusion guidance ('Acting on a finding belongs to vmware-aiops'), covering when NOT to act with this tool.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.